FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.
I've run the benchmark on a RTX 3090, 1024 tokens in / 256 out, concurrency 1–32.
If the model fits on vRAM (Gemma-4-26B-A4B, byte-identical GGUF on both engines): llama.cpp has 2.2–3.2× the throughput and 5–6× faster TTFT. FreeToken 0.1.2 can't keep 4-bit experts in VRAM at all, and it OOM'd at 8 concurrent.
If the model doesn't fit (gpt-oss-120b, 63 GB): FreeToken's TTFT stays at \~9 s from 2 to 8 users while llama.cpp's goes 17 → 58 s. At 32 users it's 19 s vs 139 s. Throughput is basically a tie (10–17 tok/s for both).
Spilling to system RAM costs \~10× in generation speed whichever engine you use.
FreeToken's PCIe link sits at its ceiling the whole time, so PCIe 4.0 should help it a lot (I've run this on a gen3 motherboard).
So from this test FreeToken only makes sense with many concurrent users in models that cannot be hold inside vRAM. But I am not sure if that is always the case or an artifact of the gen3 bottleneck on my PC.
Has anyone run a benchmark like that with a gen4 Motherboard?
Full details on the link. BTW: I used AI to generate the charts and correct my spelling and grammar.