← r/LocalLLaMA
▲
28
+2
21👁
r/LocalLLaMA · u/Exciting-Engine882 · 12d ago

is switching from llama cpp to vllm worth it

I have hp z8 g4 with 512 ram and 1x3090 1x5060 16gb. has anyone made the transition from llama cpp to vllm recently? is it worth it? docker under windows or full linux install? I am mainly interested in the model support, it seems that many new local models are supported day 0 in official vllm, while for llama cpp it takes months sometimes. LE: I want to use it for big'ish moe models, that would have to offload some tensors to system ram. I will use it just for myself. I don' t need it to be faster than llama cpp, if it runs at about the same speed it is fine , as long as it works.

28 0 28 10/3 06:31 10/6 00:05 UTC
scorecomments21 sightings
first seen 2026-10-03 06:31 UTClast seen 2026-10-06 00:05 UTCscore then 26score now 28gained +2sightings 21
open on reddit ↗ 💬 68 (+1)