Ok how to actually learn vLLM ?
Said in title, I find the ecosystem difficult to understand, and RTFMing doesn't help me as it's never clear what is the server vs their client library ? I'm using it for voxtral 3B on one GPU, but it's because I can run that with no quantization, I'm lost on learning to run with quantization / more advanced features.
I think it makes sense, because I'm running Qwen 3.8 27B quantized on llama.cpp but with everything on the GPU (RX 7900 XTX).
scorecomments20 sightings
first seen 2026-10-05 11:38 UTClast seen 2026-10-09 08:12 UTCscore then 7score now 13gained +6sightings 20
open on reddit ↗
💬 24 (+11)