← r/LocalLLaMA
▲
31
 
1👁
r/LocalLLaMA · u/SoAp9035 · 5h ago

Tested Mellum2.1-12B-A2.5B on PI Coding Agent - surprisingly usable, but not great at one-shot projects

I tested Mellum2.1-12B-A2.5B (Q8) locally using Pi and llama-server. All five tests were one-shot.

Results were pretty mixed:

\- Pelican SVG, Browser OS, Minecraft: Poor results.
\- Bouncing Hexagon: Physics were okay, but surprisingly it made it run in the terminal.
\- Flappy Bird: Completed it, but the visuals were very basic and the game was way too difficult.

Pelican SVG

Browser OS

Minecraft

Flappy Bird

Bouncing Hexagon

Okay, these might not look great, but hear me out. This model is actually pretty good.

its agentic behavior is goood. Tool calling is genuinely good, and it often catches and fixes its own editing mistakes. One time it failed to call the tool and stopped completely.

I also tried it on an existing game project. It explored the codebase and successfully changed the dogs' jump height to 3x.

Another nice surprise: after I recorded the tests, I asked the model to convert the screen recordings into GIFs and organize them into a specific folder. It handled the whole thing without any issues.

It does sometimes struggle with unfamiliar workflows, though.

Running it on a laptop with 8GB VRAM and 32GB RAM at 131K context. I'm getting around 40 t/s.

You might ask why I don't use Qwen3.6 35B-A3B instead. It gets slower with larger contexts, and my laptop gets so hot that it actually burned out the lid sensor. My laptop can barely handle it.

Overall, not a model that blows me away with one-shot creations, but definitely one I can see myself using regularly for smaller development tasks. Thanks for the model jetBrains.

posted Fri, 09 Oct 2026 15:50:15 GMTseen 1 time
open on reddit ↗ 💬 10