What are you using for NVFP4 and Do you like it?
What are people using to run nvfp4 on multi-gpu?
the only thing i can get to run is unsloth and vllm is too slow and takes FOREVER to fucking load. TensorRT-LLM has too many issues, so what are people using?
And do you like nvfp4 vs q4 qguf for qwen 3.8? apples to oranges is nvfp4 closer to Q6 gguf than q4 gguf is?
well i tired the model here https://huggingface.co/neroued/Qwen3.8-27B-nvfp4-NInfer
which works with https://github.com/Neroued/ninfer/tree/master
and initial testing has not been great for code but vision and tool calling is very impressive. Brief testing on complex scenes showed slightly better than what I've seen on q6 gguf
but im going to revisit surely im missing something, i had to spend a lot of time wiring ninfer into my frontend so ill look at it again with fresh eyes.
The speed is fantastic on 2x5060ti with 197k ctx and model loading in 4-6 seconds is wild
scorecomments7 sightings
first seen 2026-10-08 19:30 UTClast seen 2026-10-09 07:36 UTCscore then 3score now 1gained -2sightings 7
open on reddit ↗
💬 40 (+25)