← r/LocalLLaMA
▲
4
+1
11👁
r/LocalLLaMA · u/SignatureMoney6648 · 8d ago

FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.

I've run the benchmark on a RTX 3090, 1024 tokens in / 256 out, concurrency 1–32.

If the model fits on vRAM (Gemma-4-26B-A4B, byte-identical GGUF on both engines): llama.cpp has 2.2–3.2× the throughput and 5–6× faster TTFT. FreeToken 0.1.2 can't keep 4-bit experts in VRAM at all, and it OOM'd at 8 concurrent.

If the model doesn't fit (gpt-oss-120b, 63 GB): FreeToken's TTFT stays at \~9 s from 2 to 8 users while llama.cpp's goes 17 → 58 s. At 32 users it's 19 s vs 139 s. Throughput is basically a tie (10–17 tok/s for both).

Spilling to system RAM costs \~10× in generation speed whichever engine you use.

FreeToken's PCIe link sits at its ceiling the whole time, so PCIe 4.0 should help it a lot (I've run this on a gen3 motherboard).

So from this test FreeToken only makes sense with many concurrent users in models that cannot be hold inside vRAM. But I am not sure if that is always the case or an artifact of the gen3 bottleneck on my PC.

Has anyone run a benchmark like that with a gen4 Motherboard?

Full details on the link. BTW: I used AI to generate the charts and correct my spelling and grammar.

5 0 4 10/3 06:29 10/7 11:52 UTC
scorecomments11 sightings
first seen 2026-10-03 06:29 UTClast seen 2026-10-07 11:52 UTCscore then 3score now 4gained +1sightings 11
open on reddit ↗ 💬 18