← r/LocalLLaMA
▲
218
+151
12👁
r/LocalLLaMA · u/facethef · 17h ago

jevman: AI decision models play Pac-Man

post image

The other week I posted about Jev vs. Kev compared and since then, OpenAI released the decisions endpoint, Cloudflare released Clef and many here asked about Laya as well.

This time we compared six popular decision models by making them play Pac-Man: kev 1.13, Kev 4B, Clef, Clef Flash, GPT-6 Luna and Laya.

Since they respond within ms it works for them to play the game in real time.

We published a leaderboard and the repo is open-source, so anyone can run their own decision model, like your own fine-tuned one run locally or hosted somewhere, and join the leaderboard.

|Model|Avg score|High score|Avg latency|
|:-|:-|:-|:-|
|jev 1.13|2,750|6,380|290 ms|
|GPT-6 Luna|2,568|5,920|179 ms|
|Clef Flash|2,538|4,260|256 ms|
|Clef|2,476|4,820|398 ms|
|Kev 4B|1,506|5,280|231 ms|
|Laya|639|1,200|104 ms|

For the leaderboard we let each model run 100 times and took mean score with a 95% margin of error (±2 standard errors), so some models tie on top spot.

You can also play yourself as Pac-Man, and the ghosts are the decision models, either all jev, clef, Luna or Laya, or a mix of models taking over each ghost.

218 0 218 10/8 15:32 10/9 07:36 UTC
scorecomments12 sightings
first seen 2026-10-08 15:32 UTClast seen 2026-10-09 07:36 UTCscore then 67score now 218gained +151sightings 12
open on reddit ↗ 💬 70 (+46)