← r/LocalLLaMA
▲
299
-4
28👁
r/LocalLLaMA · u/buttplugs4life4me · 19d ago

Please stop with the FP4 inference engines for the love of god

Every day there's a new post of some optimized config or new inference engine that is just super good at one specific thing and their claims make sense.

And then at the bottom of the post or maybe after someone asked it says "NVFP4/MXFP4 only".

Okay dude, good job! You made the fastest possible option a little slightly faster, and most likely your output is completely cooked and you get hallucinations left right and center.

Just saw another one in r/ROCm again.

It's fine if you run 4-bit for large models, they've got lots of shit in them so a little bit of loss just means they won't remember that super good spaghetti Bolognese recipe. But running small dense models at FP4 just kills them. Like, completely. Good luck doing something productive when your model suddenly decides 1+1=3.

Just...stop.

Edit: Just going to put this here since some seem confused. a standard Q4 quantisation usually leaves more sensitive tensors in BF16, Q8 or Q6/5. Also, usually the K/V cache is quantized max to FP8/Q8.

What these "inference engines" do is usually fork an existing one (llama.cpp, SGlang, vLLM) and then just quantise \*everything\* down to FP4. Which is great for speed, especially without online dequantisation, but fucks the quality up \*a lot\*.

Your standard Q4\_K\_M/XL quant from unsloth is fine.

309 0 299 10/2 17:30 10/9 04:52 UTC
scorecomments28 sightings
first seen 2026-10-02 17:30 UTClast seen 2026-10-09 04:52 UTCscore then 303score now 299gained -4sightings 28
open on reddit ↗ 💬 249