X account claims high t/s setup, but thin on details
According to this post it is possible to reach very high numbers using mtp, however I have failed to reproduce the 50+ tps for 5060ti.
Am I just ignorant or how exactly do this? Or is this a fake post?
Freshly built llama fork for sm120 (Blackwell):
https://github.com/Anbeeld/beellama.cpp
"C:\\llama\\build\\bin\\llama-server.exe" \^
\-m "%MODEL%" \^
\--port 8090 --host 127.0.0.1 \^
\-ngl 99 \^
\--cache-type-k kvarn3 --cache-type-v kvarn3 \^
\--flash-attn on \^
\--load-mode mlock \^
\--jinja \^
\-c 98304 --parallel 1 \^
\--fit off \^
\--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-ubatch-size 128 \^
\-ctkd q8\_0 -ctvd q4\_0 \^
\--kv-tail-tokens auto
Runs at 20-35 t/s.
scorecomments13 sightings
first seen 2026-10-03 06:30 UTClast seen 2026-10-07 08:02 UTCscore then 0score now 0gained 0sightings 13
open on reddit ↗
💬 15