Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak.
​
I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster.
Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but throughput and concurrency on coding tasks have been impressive.
Performance Benchmarks
Coding Generation / Decode: Sustaining \~30 tok/s (peaking around 38 tok/s) during heavy code generation and agentic tasks.
Prefill Throughput: \~750–910 tok/s (optimized via modified NCCL topology and dual-switch setup).
Concurrency & Stress Testing: Handling multiple concurrent user requests smoothly without dropping token generation rates or starving KV cache memory.
Context / Tool Bench: Stable multi-hundred-thousand token context runs agentic workfows with multiple 500k compaction.
Compute: 16x GB10 Cluster Nodes
Connectivity: Dual MikroTik Switch (CRS804-4DDQ) using 4x 400G-to-4x100G breakout cables.
Runtime: Customized gb10-vllm stack using dspark / Inferact/Kimi-K3-DSpark wrappers with custom MLA/KV kernels.
I attached a short clip showing real-time token streaming, coding output.
GitHub & Setup Files:
All runtime patches, config files, and build scripts are on my GitHub: