← r/LocalLLaMA
▲
7
-2
11👁
r/LocalLLaMA · u/knob-0u812 · 13d ago

vLLM Recipe for Qwen38 Flash Next NVFP4 TP=2 for RTX Pro 5000 72g

I couldn't find a recipe for this model on my hardware, so I used Hermes and Unsloth's 4-bit quant of the same model to cook up a vLLM recipe for the NVFP4 quant with PLE offloading. I've been running the model for about a week and it's taken everything I've thrown at it. Very happy with how it's performing. Here's the Git repo Feedback welcome. https://preview.redd.it/1fzvnt961rrh1.png?width=768&format=png&auto=w…

9 0 7 10/3 06:32 10/4 17:48 UTC
scorecomments11 sightings
first seen 2026-10-03 06:32 UTClast seen 2026-10-04 17:48 UTCscore then 9score now 7gained -2sightings 11
open on reddit ↗ 💬 0