Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s
HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format.
Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does:
- 262K context (the model's full native window)
- \~130 tok/s decode with MTP3, 70.8% acceptance
- 3,591 tok/s prefill on a 9K prompt
- Perplexity within 1.3% of the official artifact, so the conversion is clean
- Vision and tool calls still work
Converter + writeup here: https://github.com/T-Crypt/ninfer-4090/pull/4
Questions welcome.
Repo: ninfer-uncensored
scorecomments2 sightings
first seen 2026-10-09 05:35 UTClast seen 2026-10-09 07:36 UTCscore then 2score now 6gained +4sightings 2
open on reddit ↗
💬 11 (+11)