2.3x faster Qwen3.8 27B on a 5090: ninfer vs llama.cpp, 4 setups, same prompt - speed and quality tested
Hi guys I keep seeing people talk about ninfer, so I wanted to know if switching from llama.cpp is actually worth it. This was the prompt that I was using (physics, spin, full rules, the works) Setup: RTX 5090, Qwen3.8 27B, thinking on (xhigh), 120k context, default sampling settings, one run oneshot Speed |Setup|Output tokens|Time|Decode tokens/s| |:-|:-|:-|:-| |ninfer, \[precision of the non-NVFP4 build\], MTP|67,539|7m 40s|\~147| |llama.cpp Q4\_K\_M + MTP (draft-n-max 3)|69,749|8m 13s|\~141| |ninfer, NVFP4, MTP|93,663|10m 3s|\~155| |llama.cpp Q4\_K\_M, no MTP|78,976|19m 31s|\~68| A few things stood out. Stock llama.cpp without MTP is less than half as fast as ninfer. But once you turn on MTP in llama.cpp it jumps from 68 to 141 t/s and lands very close to ninfer, so a big part of the "ninfer is fast" story is really "MTP is fast". NVFP4 had the highest t/s, but it also wrote the most tokens (mostly thinking), so it only finished third on wall-clock time. For reasoning models I'd look at time-to-result, not just t/s. Quality I checked all four games with a script that fires about 2,400 random shots (random angle, power and spin) at each one, plus a few scripted rule scenarios. Good news: none of them crashed, produced NaNs or got stuck, so all four run. The differences are in the rules: ||ninfer NVFP4|ninfer \[non NVFP4\]|llama Q4\_K\_M|llama Q4\_K\_M + MTP| |:-|:-|:-|:-|:-| |Can you legally win by potting the 8?|yes|no|stripes only|no| |8-ball on the break|respotted|re-rack|counts as a loss|counts as a loss| |Starting rack OK?|yes|yes|balls overlap|yes| |Sound|no|yes|no|yes| |Lines of code|933|1185|1127|965| All four run fine, but only the NVFP4 game can actually be won. The other three have small logic bugs in the win condition (and one has a broken starting rack), so none of them is quite finished. You can try them yourself: Qwen 3.8 27B ninfer NVFP4: https://claude.ai/artifact/1oj8LJkBRrLmHhQLSe9kCu Qwen 3.8 27B [ninfer \[non NVFP4\]](https://huggingface.co/neroued/Qwen3.8-27B-NInfer): https://claude.ai/artifact/WbW2XSmKDfxgAArKGEeihC Qwen 3.8 27B llama.cpp Q4\_K\_M: https://claude.ai/artifact/DqYMjSR7unZ8kicr3JKnSg Qwen 3.8 27B llama.cpp Q4\_K\_M + MTP: https://claude.ai/artifact/2sunpBgBJBRvbaJAJnkM3t Keep this in mind before you trust my numbers: One run per setup, so some of the bugs could just be bad luck. Everything ran on the default reasoning effort (xhigh), which inflates the token counts. NVFP4 and Q4\_K\_M are different quant schemes, so don't treat them as equivalent. My take: I'm sticking with the non-NVFP4 ninfer build for my next round of prompts. Of the four games, that one was my favorite to actually play. It had the most polish: sound, realistic ball size, the break rules, a proper kitchen for ball-in-hand. The only thing that bugged me is that you can't win a game legally, because potting the 8 after clearing your group counts as a foul. Funny enough, it turned out to be a one-line bug (an inverted check), so it was really close to being the best of the bunch.