DeepSeek V4.1 Flash on a single DGX Spark: 113.6 GB VQ base + 40 MB domain sidecars, 74–82% top-1 agreement vs original
I’ve been working on YoungAi, a native C/CUDA inference engine that runs DeepSeek V4.1 Flash on a single NVIDIA DGX Spark (GB10, 128 GB unified memory). The original weights are \~510 GB. I deploy it as three files:
- ① Base GGUF — 113.6 GB, universal, zero-corpus. Quantized once from official weights.
- ② Domain sidecar — \~40 MB per domain. Solved once per domain, then frozen.
- ③ Post-training file — experimental, re-solved nightly. Delete it to roll back.
Each routed expert row is scaled by g_base × s_sidecar × s_posttrain, and the router gets a bias Δb_sidecar. The base alone is a complete model; sidecars just add tiny scaling/bias without changing kernels.
TL;DR
- Single DGX Spark, 113.6 GB resident + \~40 MB sidecar.
- 5 domains: finance, code, law, medicine, science.
- Top-1 agreement vs original improves +2.8 to +3.7 points with a domain sidecar.
- Speculative decode: 43 tok/s on a real 14.1k-token Agent request (greedy).
- Prefill: 1,055 tok/s on 12.5k prompt; 671 tok/s on 106.7k prompt.
- English WikiText-2 does not regress when any domain sidecar is attached (it actually goes up).
Core implementation ideas
Base (VQ-8 + per-layer shared codebook). Every 8 consecutive weights in an expert row become one 12-bit (or 13-bit) codebook index, multiplied by a single per-row gain. Codebooks are trained per layer and shared across all 384 experts and three matrices. Codebooks are stored in FP8 (E4M3). 13-bit layers use a “12+1” bit-plane layout for 128-byte cache line alignment. The base is zero-corpus: it never sees domain data.
Domain sidecar (“anti-solver”). For each domain, I solve a multiplicative gain per output channel of every expert’s down projection, plus a router bias per expert. Objective: reproduce the original model’s MoE block output on domain text, layer by layer, using the engine’s own prefill hooks. Gains are stored in FP4 with lattice-aware Gauss-Seidel. The sidecar is \~40 MB and adds \~0.6 MB read per decoded token.
Post-training file (experimental). Turn “the model should write a, not b” into a linear equation on last-layer expert gains, then solve with conjugate gradient. On the training request, decision points flip from 55% to 88%, but it does not generalize across trading days yet.
Multi-domain real metrics
All metrics are teacher-forced against the original DeepSeek V4.1 Flash (official PyTorch code, full precision). Higher top-1 / Σmin is better; lower KL / PPL ratio is better.
|Domain (judgment slice)|Base only top-1|\+ domain sidecar top-1|Σmin (median / p5)|Avg KL|PPL ratio|
|:-|:-|:-|:-|:-|:-|
||
|Finance (8,192 tok)|71.73%|74.57%|0.745 (0.810 / 0.283)|0.512|1.267|
|Code (15,360 tok)|78.61%|82.26%|0.803 (0.872 / 0.402)|0.303|1.229|
|Law (15,360 tok)|72.90%|76.36%|0.760 (0.834 / 0.272)|0.454|1.251|
|Medicine (15,360 tok)|69.08%|72.82%|0.739 (0.767 / 0.325)|0.465|1.280|
|Science (15,360 tok)|71.65%|74.93%|0.753 (0.788 / 0.347)|0.422|1.173|
|English WikiText-2 (512 tok)|78.52%|80.66–82.81% (any sidecar)|0.787–0.801|0.566–0.619|1.564–1.658|
English row shows that domain sidecars don’t hurt general ability; all five sidecars actually improve it slightly.
Speed on one DGX Spark
|Scenario|Prefill|Decode|
|:-|:-|:-|
||
|12.5k-token prompt|1,055 tok/s|—|
|Real 14.1k-token Agent request|940 tok/s|—|
|106.7k-token prompt via server|671 tok/s (159 s TTFT)|—|
|Short prompt, pure greedy|—|30.5–30.7 tok/s|
|14.1k-token Agent request, pure decode|—|28.9–29.4 tok/s|
|Same request, speculative (default)|—|43.0 tok/s (3.04 tok/round)|
|Unseen 9.2k prompt, speculative|—|40.0 tok/s|
|51k context, pure decode|—|27.5 tok/s|
Decode is memory-bound: \~6.3 GB read per token. GB10 measured bandwidth is \~235 GB/s, so the wall is \~37 tok/s; we hit \~32.5 ms, or 82% of the wall.
Honest limitations
- Post-training (③) is a working mechanism, not a product yet. It flips specified decisions on the solving request but does not transfer to held-out days (55% → 55%).
- Five domains only. Sidecars are evaluated teacher-forced on held-out text, not yet end-to-end.
- CUDA only, validated only on DGX Spark. No Metal.
- Speculative decoding only kicks in for greedy; sampling requests fall back to pure decode.
- Source code (engine, quantizer, solver) is not public yet.
Feedback welcome
- Are these top-1 agreement / Σmin numbers useful for real workloads?
- Is the VQ-8 + per-layer codebook + sidecar gain approach reasonable?
- What benchmarks or integration points would you want to see next?
Model card and weights: https://huggingface.co/wenzhouwu/YoungAi-DeepSeek-V4.1-Flash
This is not an official DeepSeek release. If this kind of post isn’t appropriate here, let me know and I’ll move or remove it.
Thanks!