(Self Promotion) Kimi vs. Claude vs. GPT vs. Gemini as teammates. Who actually coordinates?
Benchmarks test models alone. I wanted to know how they do with a partner.
We paired four models in every combination in a co-op game where players are tied by a rope. Top with top won most. A third or fourth agent hurt every model. Human pairs still beat all of them.
I work at Skillprint. We build games like this to capture how people and models coordinate, because AI that works alongside people needs that context.
Pairing matrix and GIFs: https://experiments.skillprint.co/posts/signal/
Play it yourself: https://experiments.skillprint.co/play
Which matchups should we run next?
scorecomments6 sightings
first seen 2026-10-07 01:49 UTClast seen 2026-10-07 21:24 UTCscore then 0score now 0gained 0sightings 6
open on reddit ↗
💬 1 (+1)