← r/LocalLLaMA
▲
10
+1
11👁
r/LocalLLaMA · u/Biomass23 · 8d ago

tp=6 can work on vLLM, with padding

vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).

I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.

So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.

I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.

GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x

11 0 10 10/3 06:29 10/7 13:54 UTC
scorecomments11 sightings
first seen 2026-10-03 06:29 UTClast seen 2026-10-07 13:54 UTCscore then 9score now 10gained +1sightings 11
open on reddit ↗ 💬 24