I've been tinkering with GLM-5.3-Flash on two 64GB CMP 170HX cards for a while, and the setup is finally stable enough that I figured I'd share it.
I also compared it against the Qwen3.8-Flash-Next setup I've been using on the same machine: AWQ INT4 + FP8 PLE on vLLM.
Besides PP/TG benchmarks, I hooked both models up to DSH and gave them the same small coding/agent tasks to see how raw inference speed translated into actual task completion time.
A few caveats up front:
- this is not an apples-to-apples quant comparison
- GLM and Qwen are using different engines and different speculative decoding setups
- speculative decode speed depends heavily on acceptance rate and generated text
- the coding tasks are just a few practical examples, not a serious benchmark suite
- when I mention “Strata-style” below, I mean the HBM-first/full-residency approach I previously used with Strata, not that this is running Strata itself
Repo and playable demos:
GitHub:
https://github.com/Flun/glm53-flash-cmp170hx-exl3
Hardware
|CPU|Ryzen 5 5600X|
|:-|:-|
|RAM|80GB DDR4|
|GPU|2x CMP 170HX 64GB|
|GPU arch|SM80|
|PCIe|Gen2 x8|
|GPU P2P|unavailable|
|OS|Ubuntu 24.04|
Both models were tested on the same machine.
For the agent tests I used DSH as the harness.
GLM-5.3-Flash setup
Target model:
turboderp/GLM-5.3-Flash-exl3 3.05bpw
Engine:
ExLlamaV3 1.5.4
Current setup:
- GLM-5.3-Flash EXL3 3.05bpw
- \~125.2GB / 116.6GiB target weights
- target fully resident across the two 64GB cards
k_hcfuse- DFlash2 EXL3 6bpw
- DFlash2 K7
- Q8 KV cache
- 384K context in actual use
- max request budget around 392,960 tokens
GLM-5.3-Flash itself is a 320B-total / \~18B-active MoE model.
The important part here is that the target weights stay resident in HBM. I'm not continuously streaming experts from system RAM during decode.
About the 3.05bpw quality
This was probably the part I cared about most.
At first glance, “3.05bpw” sounds like a pretty aggressive quant, especially compared to the UD Q4 variants people commonly use.
But EXL3 isn't simply “make every tensor 3-bit”.
It uses a trellis-based quantization scheme with different bit allocation depending on the tensor. The 3.05bpw number is an average target bitrate.
The published quant configs for this family also keep more sensitive parts at higher precision. For example, lm_head remains at 6-bit in the 3.05bpw branch.
Looking at the same model family, the 4.05bpw build has been inspected with something roughly like:
- routed experts: K4
- attention: K6
- shared experts: K6
- dense MLP: K5
- lm\_head: K6
- embedding / norms / router: native
So the general idea is to compress the huge routed-expert portion more aggressively while spending more bits on the smaller/more sensitive paths.
That makes quite a bit of sense for a MoE model like this, because most of the storage is in the expert weights.
Rough comparison with the common UD quants
|Quant|Size|Top-1 agreement vs BF16|Mean KLD|
|:-|:-|:-|:-|
|UD-IQ3\_XXS|120.37GB|81.63%|0.28377|
|EXL3 3.05bpw (my target)|125.18GB|\~93.05% (estimated from published 3.0bpw results)|\~0.050 (estimated from published 3.0bpw results)|
|UD-IQ4\_XS|156.82GB|88.18%|0.11665|
|UD-Q4\_K\_XL|199.71GB|92.22%|0.04929|
|UD-Q5\_K\_XL|240.31GB|94.35%|0.02705|
For reference, public GLM-5.3-Flash GGUF fidelity numbers look roughly like this:
One thing worth pointing out is that something named Q4_K_XL is not literally “4 bits per parameter across the entire model”.
At \~200GB for a 320B model, it's a mixed-precision quant with an effective average bitrate much higher than 4bpw.
There is also a published GLM-5.3-Flash EXL3 3.0bpw fidelity test using 51,175 held-out next-token positions that reported:
- Top-1 agreement: \~93.0%
- Mean KLD: \~0.0505
Numerically, that's in roughly the same neighborhood as the published UD-Q4\_K\_XL result.
That said, I would not claim that “EXL3 3bpw is better than UD-Q4\_K\_XL” from those numbers alone.
They were not measured through the exact same evaluation pipeline/corpus, and the public 3.0bpw artifact isn't the exact same quant I'm running either.
My takeaway is simply that 3bpw-class EXL3 can preserve a surprising amount of fidelity for its size, and it doesn't behave like a naive 3-bit quant.
For my use case, getting the target down to \~116.6GiB while still retaining usable coding/agent quality was the main reason this setup was interesting.
DFlash2 6bpw is only the drafter
Just to avoid confusion:
the 3.05bpw model is the actual GLM target.
The 6bpw DFlash2 model is only the speculative drafter.
The drafter proposes tokens, and the GLM target verifies them.
So this is not some kind of “3.05bpw + 6bpw averaged quality” setup.
The draft quant mostly affects draft speed, VRAM use, and acceptance efficiency.
HBM-first / “Strata-style” part
This is where I borrowed an idea from a Strata setup I had used previously.
Again, this does not run Strata.
What I mean by “Strata-style” is simply:
keep as much of the model permanently resident in HBM as possible, and avoid runtime CPU↔GPU weight traffic
The GLM target itself fits across the two cards, so I leave the target fully resident.
Instead of offloading experts, I focused on reducing the memory used by the parts that scale with context: KV cache and the speculative drafter.
For this particular machine that made more sense to me than constantly moving weights over PCIe.
DFlash2 + 384K context
I originally used GLM's MTP d2 path.
Later I switched to DFlash2.
Instead of keeping the original BF16 incoai/GLM-5.3-Flash-DFlash2 drafter, I converted it to an ExLlamaV3-compatible EXL3 6bpw build.
The resulting draft weights are about 0.96GiB.
The bigger problem at long context was actually the draft KV cache.
If the target is running 384K and the drafter also grows a 384K KV cache, VRAM disappears quickly.
So I changed the drafter side to use a fixed SWA window plus a GPU ring cache.
The target still sees the full 384K context and keeps its full target KV.
Only the drafter's KV storage is kept inside a bounded ring.
That's what lets the current setup run:
DFlash2 K7 + Q8 KV + 384K target context
without growing the draft cache to the full target length.
The implementation and validation tests are in the repo.
Cold start
I also measured from a cold compile/start until the API was actually ready.
|Model|Ready time|
|:-|:-|
|GLM-5.3-Flash EXL3|\~1m 04s|
|Qwen3.8 Flash Next / vLLM|\~3m 50s|
This isn't really a model-size comparison.
The Qwen vLLM setup has quite a bit more startup work:
- PP workers
- distributed runtime
- model placement
- MTP
- PLE
- GDN
- Triton compilation
- memory profiling
- KV allocation
The ExLlamaV3 GLM path is comparatively static.
Inference benchmarks
These are the numbers from my dashboard workload.
Again, especially for speculative decode, I wouldn't treat these as universal model speeds.
Acceptance rate and generated text matter a lot.
GLM-5.3-Flash / DFlash2 K7 / Q8
|Input|PP|Decode|
|:-|:-|:-|
|8K|1,529 tok/s|95.6 tok/s|
|40K|1,624|90.7|
|73K|1,647|90.6|
|106K|1,650|91.0|
|131K|1,611|94.4|
|385K|1,535|90.1|
DFlash acceptance on this particular workload was mostly around 87%.
With the older MTP d2 path, the same dashboard workload was generally in the \~60 tok/s range.
Switching to DFlash2 K7 brought it to around \~90 tok/s here.
Qwen3.8 Flash Next / vLLM
The Qwen setup is:
AWQ INT4 + FP8 PLE / PP2 / MTP3
|Input|PP|Decode|
|:-|:-|:-|
|8K|5,513 tok/s|129.5 tok/s|
|40K|5,628|123.6|
|73K|5,466|141.9|
|106K|5,307|141.5|
|131K|5,183|159.9|
|252K|4,706|149.4|
So on raw throughput, Qwen is clearly faster.
At roughly 131K:
- PP: \~5.18K vs \~1.61K
- decode: \~160 vs \~94 tok/s
No argument there.
The interesting part for me was what happened once I actually let both models do multi-step coding work.
Why I stopped at 384K for now
I tested roughly 385K input and PP was still around 1.5K tok/s.
The problem wasn't PP collapsing.
It was simply wall-clock time.
Prefilling \~385K from scratch already takes about 4 minutes.
Even if I can make 1M fit, doing a full 1M cold prefill at this speed isn't particularly attractive for normal use.
So I'm currently leaving the service at 384K Q8.
I still want to see if I can get 1M working eventually, mostly for the technical exercise.
Small agent tests
Originally I was only going to make both models build Tetris and stop there.
Both were connected to DSH and got the same request.
1. Tetris
Prompt:
Build a playable Tetris game for the web.
|Model|Completion time|
|:-|:-|
|GLM-5.3|4m 40s|
|Qwen3.8|5m 30s|
Both produced working versions, and honestly the difference wasn't dramatic enough to be very interesting.
So I added two more tasks.
https://reddit.com/link/1x0b1ws/video/1sn8netal4uh1/player
2. AI mini PC landing page
Exact prompt given to both:
Build a polished single-file HTML landing page for an AI mini PC with a dark theme, specs, performance charts, pricing, FAQ, and smooth scroll animations, using no external libraries.
|Model|Completion time|
|:-|:-|
|GLM-5.3|10m 05s|
|Qwen3.8|14m 40s|
https://reddit.com/link/1x0b1ws/video/pdglw7kbl4uh1/player
3. Vampire-Survivors-style game
Exact prompt:
Build a single-file HTML vampire-survivors-style game with WASD movement, auto-attacks, enemy waves, XP, 3-choice level-up upgrades, HP, game over, and restart, using no external libraries.
|Model|Completion time|
|:-|:-|
|GLM-5.3|4m 40s|
|Qwen3.8|24m 10s|
This one had a much larger difference than I expected.
There is one obvious caveat:
the Qwen version added sound, while the GLM version did not.
The prompt didn't ask for sound, so I didn't go back and ask GLM to add it afterward. I wanted to leave both runs as the result of the same one-shot prompt.
https://reddit.com/link/1x0b1ws/video/4hl6g2bcl4uh1/player
Task completion times
|Task|GLM-5.3|Qwen3.8|
|:-|:-|:-|
|Tetris|4m 40s|5m 30s|
|Landing page|10m 05s|14m 40s|
|Vampire-style game|4m 40s|24m 10s|
I wouldn't read too much into three examples.
This definitely isn't evidence that GLM is “5x better at coding” or anything like that.
What I found interesting is simply that raw tok/s and end-to-end agent completion time didn't track each other very well.
Qwen has much higher PP and decode throughput, but on these particular tasks GLM often finished sooner.
For agent work, planning, number of retries, file rereads, edits, and how close the first implementation is to working all matter too.
So I think raw inference speed and actual task completion time are worth looking at separately.
Current state
The GLM service I'm using now is:
GLM-5.3-Flash EXL3 3.05bpw
- DFlash2 EXL3 6bpw K7
- Q8 KV
- 384K context\*\*
The main thing I like about this configuration is the memory/quality tradeoff.
The target fits in \~116.6GiB of HBM, stays resident, and the public 3bpw-class EXL3 fidelity results suggest the quant is holding up much better than I would have expected from the bitrate alone.
The runtime side is basically an HBM-first setup: keep target weights resident, then save memory on the drafter/KV side rather than moving experts back and forth during decode.
Full config, conversion scripts, ring-cache changes, benchmark code and raw results are here:
GitHub:
https://github.com/Flun/glm53-flash-cmp170hx-exl3
If anyone is running GLM-5.3-Flash on other weird 128GB-class GPU setups, I'd be interested in seeing what numbers you're getting too.
Next thing I want to try is 1M context, although at that point prefill time is probably the bigger problem than just making it fit.