← r/LocalLLaMA
▲
13
+6
20👁
r/LocalLLaMA · u/BraceletGrolf · 4d ago

Ok how to actually learn vLLM ?

Said in title, I find the ecosystem difficult to understand, and RTFMing doesn't help me as it's never clear what is the server vs their client library ? I'm using it for voxtral 3B on one GPU, but it's because I can run that with no quantization, I'm lost on learning to run with quantization / more advanced features.

I think it makes sense, because I'm running Qwen 3.8 27B quantized on llama.cpp but with everything on the GPU (RX 7900 XTX).

13 0 13 10/5 11:38 10/9 08:12 UTC
scorecomments20 sightings
first seen 2026-10-05 11:38 UTClast seen 2026-10-09 08:12 UTCscore then 7score now 13gained +6sightings 20
open on reddit ↗ 💬 24 (+11)