← r/LocalLLaMA
▲
57
+22
47👁
r/LocalLLaMA · u/roofkid · 6d ago

I built Ninfer 4080 for 16GB class GPUs

Hi everyone,

TL/DR

I created NInfer 4080 to run ISTA-DASLab-Qwen-3.8-27B-GSQ at 100k context on an RTX 4080 16GB GPU using way more of the hardware capabilities (max overall: 2720 tok/s prefill, 262 tok/s generation) and sharing it with the community now so others can also have the benefit.

https://github.com/roofkid/ninfer-4080

Full Version

After seeing all the amazing work done in the community creating Ninfer 5090, 4090 and 3090 I admit I was a little sad to not being able to use any of it on my RTX 4080 with only 16GB of memory. I still had about $13 of credits sitting idle on the DeepSeek platform as I never expected how much usage I would get out of it.

For context I have over 20 years of experience in Software Engineering and Architecture, but have no experience whatsoever in GPU Kernel development, so this was a very interesting pet project also from a professional experience for me. Mainly because I can read and understand C++ but could not judge the actual Kernel code. So I approached it from a product owner and requirements perspective only, made sure good software engineering practices are followed and only made "business decisions".

I've been actively following the local LLM community for the last 2-3 years, probably have tried out all models I could over that time and followed the progress with amazement like many of you.

Guiding principles

  • Fit into RTX 4080 16GB GPU
  • Use ISTA-DASLab-Qwen-3.8-27B-GSQ -> Reasoning can be seen in the ByteShape article, really good for the size and they claim even better accuracy than much larger Unsloth UD quants: https://byteshape.com/blogs/Qwen3.8-27B/#96-gb-rtx-pro-6000 I also have very good personal experience with it, it is my daily driver
  • Use DFlash2 speculative decoding
  • Reach 100k+ context
  • Significantly improve prefill and token generation speeds to utilize the hardware better than general purpose inference engines like llama.cpp or vllm
  • Measure after changes to also ensure accuracy remains, I also have a M4 48GB available to test higher quants for comparisons, though of course that is much lower speed
  • Use DeepSeek V4.1 Flash for the work for cost efficiency
  • Use Pi as the harness (only non-cosmectic extensions: hashline edit pro, internet search with ketch through local SearXNG with a self-written skill)
  • Runtime also available as a Docker image so it's easy for folks to run

Results

|Depth|Prefill t/s (DFlash2)|MTP3 decode t/s|DFlash2 K=7 decode t/s|
|:-|:-|:-|:-|
|8K|2719.9|151.2 (100%)|166.7 (54.0%)|
|32K|2424.9|141.7 (100%)|262.3 (100%)|
|64K|2125.5|130.7 (100%)|239.1 (100%)|
|98K|1895.1|122.3 (100%)|212.7 (98.2%)|

In real work I really do see the high prefill numbers (2k+) if the prompt is long enough and about 150-200 decode speed on coding and 100ish on prose. It subjectively feels significantly faster than beellama (my previous daily driver) at the same benchmark results. I mainly used MBPP and HumanEval as I needed something that I can run reasonably fast (\~30min). MBPP stays in 90-92% territory and HumanEval at 95-96%. Please be realistic and do expect tiny degradations that are within measurement noise. They are mainly coming from KV quantization according to my measurements so you can always trade context for accuracy if needed by switching.

What I learned

  • It is absolutely mental how much performance is left on the table by using the general purpose engines. From a bird's eye view it's totally understandable as we trade the wide support for performance, I just didn't expect how much that would be. When I saw the first memory throughput measurements being in the 200 GB/s range and having a theoretical maximum of 720 GB/s in the device my jaw dropped because of the low efficiency back when I started
  • I think in the community we've all seen more specialized inference engines making significant performance improvements possible. vllm-radiance for R9700, NInfer variants for CUDA, Splash for Metal - with software creation becoming cheaper and cheaper I expect more of this for and from our "tinkerer" group here
  • Spending about 2 billion tokens for this work for only $13 is just crazy (only off-hours). Low cache read tokens costs on agentic work are so much more important than even I expected. It's the classic difference between cognitively fully understanding how LLM turns work and seeing big data results. The reality is that with THAT kind of pricing I think I pay more for electricity to get the same amount of tokens out
  • I went back to xhigh thinking on Qwen 3.8 27B as the speed is so high, that I don't really care/notice. I've also hidden the thinking blocks again as I cannot follow any more anyway
  • The prefill speed really caught me of guard. I was really floored when I tried it in Pi after the first big improvements were done and it IMMEDIATELY answered with token streaming. I was so used to waiting 5-10s without a cached system prompt. I significantly underestimated how important that is for the user experience. Feels like a cloud endpoint to me now.
  • At these high prefill speeds your context window is full in 40 seconds, definite "oh my god" moment for me when that happened the first time
  • Reaching 100k context means significant KV compression as full 256k context F16 needs exactly 16GB of VRAM on Qwen 3.8 27B. I was too afraid of "high" (4bit style) KV compressions. So many advances have been made here. Originally I never went below Q8\_0. I then used kvarn5/kvarn5 previously on beellama after benchmarking and cannot measure a noticeable difference to the now used rk4v4-e8 variant used here. I think good software engineering practices are way more important and catch problems that might come from it. Also subjectively I do not experience a "fast garbage" phenomenon here

Conclusion

For me this is a good version 1 and I don't intend to spend significant effort on this for Qwen 3.8 27B. It's at the pareto 80% state. I just want to be happily using it now and reap the rewards. I hope you are too! Of course when Qwen 4 27B comes around soon I will check it out again.

If you have another 16GB RTX 4xxx card I would be interested in knowing if that works on them too and what speeds you're seeing. I honestly can't judge how tied to the RTX 4080 hardware it is. If you have a 4080, enjoy :)

Shoutouts

  • Every person who worked on NInfer before me, you guys rock and provided a stable base for me to fork from
  • Special hats off to sergiuszm who created NInfer-4090, I think you did all the heavy lifting for SM\_89 already
  • ISTA-DASlab for their work on GSQ and providing the safetensor checkpoint for it! Cheers to Austria from Germany :) Love seeing important contributions to the community from the EU
62 0 57 10/4 01:29 10/9 06:22 UTC
scorecomments47 sightings
first seen 2026-10-04 01:29 UTClast seen 2026-10-09 06:22 UTCscore then 35score now 57gained +22sightings 47
open on reddit ↗ 💬 62 (+51)