← r/LocalLLaMA
▲
222
+165
7👁
r/LocalLLaMA · u/Fun-Meaning-6474 · 15h ago

Running decision model locally on an RTX 4090 to find out which one is the fastest

post image

recently saw a bunch of open decision models pop out of nowhere in the last two weeks (laya, liquid's d1, cloudflare's clef-flash, interfaze's lev), so I wanted to see how far apart they actually are on the same GPU(yes, model size is a huge factor, but still isn't the only factor). all four had the same task of reading nine wikipedia articles about centipedes (9,534 words) word by word and flag every word that names a centipede. one /v1/systemone call per word, the next word goes out the second the answer comes back

request for every word:

{"state": "Word: \"Scolopendra\".", "questions": {"centipede": {"type": "noul", "instructions": "Does this word name a kind of centipede?"}}}

|model|weights|engine|per word (p50)|words in 32s|accuracy|centipede names caught|wrong picks|
|:-|:-|:-|:-|:-|:-|:-|:-|
|Laya|Laya-BF16.gguf|llama.cpp b11495|3.9 ms|7,980|97.4%|70%|98|
|d1 3B|d1-3B-AD-Q4\_K\_M.gguf|llama.cpp b11495|6.0 ms|5,306|96.5%|51%|51|
|Clef-Flash 9B|Clef-Flash-Q8\_0.gguf|llama.cpp b11495|24.4 ms|1,292|97.2%|36%|2|
|Lev 4B|interfaze-ai/lev, bf16|lev serve (PyTorch)|51.0 ms|626|98.9%|83%|4|

laya and d1 gap the other models in speed, though not so much on accuracy (yes, it does say 95%, but even saying "no" counts as a correct answer, so that's where the high acc comes from). what everyone might care about more is how well each one did their respective task and lev catches the most while being 13x slower than laya, partly because it runs in its own pytorch server instead of llama.cpp (it measured 68 ms on a different 4090, so it's CPU-sensitive too). but in the end Laya is the fastest model overall, and considering how easily it can be fine-tuned for any use case I'd say that be my go to pick

setup:

  • GPU: rented RTX 4090 (driver 580.119.02, 32 vCPU)
  • engine: llama.cpp b11495 (commit 37ac63456, CUDA 12.8 release build), -ngl 99, everything else default
  • Laya, Clef-Flash: the ggml-org GGUFs
  • d1: our own AD-Q4\_K\_M quant (atomic.chat), runs natively on /v1/systemone since the lfm2-d1 support landed in #30110
  • Lev: interfaze's LoRA on Qwen3.5-4B in its own lev serve, default settings (--compile never finished warming up)
  • latency: end to end from a Python client on the same box over localhost
222 0 222 10/8 19:30 10/9 07:36 UTC
scorecomments7 sightings
first seen 2026-10-08 19:30 UTClast seen 2026-10-09 07:36 UTCscore then 57score now 222gained +165sightings 7
open on reddit ↗ 💬 46 (+29)