← r/LocalLLaMA
▲
83
+1
19👁
r/LocalLLaMA · u/ciprianveg · 19d ago

Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.

post image

​

​I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster.

​Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but throughput and concurrency on coding tasks have been impressive.

​Performance Benchmarks

​Coding Generation / Decode: Sustaining \~30 tok/s (peaking around 38 tok/s) during heavy code generation and agentic tasks.

​Prefill Throughput: \~750–910 tok/s (optimized via modified NCCL topology and dual-switch setup).

​Concurrency & Stress Testing: Handling multiple concurrent user requests smoothly without dropping token generation rates or starving KV cache memory.

​Context / Tool Bench: Stable multi-hundred-thousand token context runs agentic workfows with multiple 500k compaction.

​Compute: 16x GB10 Cluster Nodes

​Connectivity: Dual MikroTik Switch (CRS804-4DDQ) using 4x 400G-to-4x100G breakout cables.

​Runtime: Customized gb10-vllm stack using dspark / Inferact/Kimi-K3-DSpark wrappers with custom MLA/KV kernels.

​I attached a short clip showing real-time token streaming, coding output.

​GitHub & Setup Files:

All runtime patches, config files, and build scripts are on my GitHub:

👉 https://github.com/ciprianveg/gb10-vllm

84 0 83 10/3 04:47 10/8 13:12 UTC
scorecomments19 sightings
first seen 2026-10-03 04:47 UTClast seen 2026-10-08 13:12 UTCscore then 82score now 83gained +1sightings 19
open on reddit ↗ 💬 36