Follow up: Qwen 3.8 27B at ~96t/s decode with NInfer on a 16GB RTX 5080, 110k context
Hi all,
I previously posted about getting Qwen 3.8 27B running at around 75t/s with llama.cpp. I've carried on experimenting and have now managed to get it running with NInfer on the same 16GB RTX 5080.
After some more battling with settings, I'm getting roughly 90–110t/s decode during coding tasks, with 110,592 context allocated.
Looking through 32 completed requests from a Zoo Code session:
- Median decode: 96.45t/s
- Lowest: 84.6t/s
- Highest: 131.7t/s
- Median time to first token: 1.4 seconds, with prompt caching working on most turns
These were requests with tool calls and conversation history, with prompts growing to around 77–79k tokens. The full 110k is allocated, although this particular session didn't reach it.
I'm running NInfer v1.5 in Ubuntu 24.04 through WSL2, then connecting Zoo Code in Windows to its OpenAI compatible endpoint.
These are the settings I've ended up using:
~/ninfer-5080/build/apps/ninfer-serve \
~/models/qwen3_8_27b.ninfer \
--host 0.0.0.0 \
--port 8080 \
--model-id qwen3.8-27b \
--max-context 110592 \
--kv-capacity 110592 \
--prefill-chunk 896 \
--kv-dtype q4 \
--spec mtp \
--draft-tokens 3 \
--embedding-host \
--max-concurrency 1 \
--default-thinking-budget 2048 \
--prefix-checkpoint-policy rolling-tool
Getting everything into 16GB was the fiddly bit. The weights take about 11.86 GiB according to the startup log. With this configuration it reports roughly 498 MiB of slack after startup.
I settled on 110,592 context to leave a bit of breathing room. Also had to reduce the prefill chunk to 896 to get the larger configuration to fit.
Here's an example from a turn with almost 50k context:
prompt=49907 gen=409 reasoning=126 cache=49309
ttft=558ms prefill=1177.7tok/s decode=104.6tok/s
wall=4.46s speculative=mtp 3.00tok/round (66.7%)
And further into the conversation:
prompt=77003 gen=2688 reasoning=2048 cache=73877
ttft=2753ms prefill=1175.9tok/s decode=90.3tok/s
wall=32.52s speculative=mtp 2.74tok/round (58.0%)
It can still take a while to finish a turn. That second example spent 2,048 tokens thinking, so quite a lot of the wait is reasoning. Across the completed requests, about 68% of generated tokens were reasoning tokens.
Losing the prompt cache also makes a big difference. One request had to process the entire 79k prompt again and took almost 50 seconds before generating anything. Once it started generating, it was still doing about 94t/s.
A couple of things caught me out connecting Zoo Code:
- The base URL needs to be
http://127.0.0.1:8080/v1. Leaving off/v1gave me a 404. - Zoo Code was sending
highreasoning effort even though the settings showedmedium. NInfer rejected it. Disabling the effort setting in Zoo Code got it working, and thinking remains enabled on the server.
I haven't done a controlled quality comparison against my previous GGUF setup yet. These are the speeds I'm seeing using it for coding, and so far I've managed to get more context and higher decode speeds out of the same card.
Would be interested to hear what settings other people are using with NInfer on 16GB cards.