← r/LocalLLaMA
▲
6
+4
2👁
r/LocalLLaMA · u/Distinct-Pie2389 · 3h ago

Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s

HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format.

Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does:

  • 262K context (the model's full native window)
  • \~130 tok/s decode with MTP3, 70.8% acceptance
  • 3,591 tok/s prefill on a 9K prompt
  • Perplexity within 1.3% of the official artifact, so the conversion is clean
  • Vision and tool calls still work

Converter + writeup here: https://github.com/T-Crypt/ninfer-4090/pull/4

Questions welcome.

Repo: ninfer-uncensored

6 0 6 10/9 05:35 10/9 07:36 UTC
scorecomments2 sightings
first seen 2026-10-09 05:35 UTClast seen 2026-10-09 07:36 UTCscore then 2score now 6gained +4sightings 2
open on reddit ↗ 💬 11 (+11)