← r/LocalLLaMA
▲
0
-1
10👁
r/LocalLLaMA · u/tabletuser_blogspot · 5d ago

GLM-4.7 benchmark compared MXFP4 vs Q4_K_M vs Q4_K_XL using Radeon 6800H iGPU 680M

Using llama.cpp Ubuntu Vulkan prebuilt binary and the Acemagic miniPC S3A using an AMD Ryzen 7 6800H is a high-performance 8-core, 16-thread mobile processor launched on January 4, 2022, built on the 6nm Zen 3+ architecture loaded with 64GB of DDR5 RAM. It features a 3.2 GHz base clock, a 4.7 GHz boost clock, a 45W TDP, and powerful integrated (iGPU) Radeon 680M graphics. # Tested Models Based on the benchmark commands and llama-bench output labels: 1. GLM-4.7-Flash-MXFP4_MOE.gguf (Reported: deepseek2 30B.A3B MXFP4 MoE | 15.79 GiB) 2. GLM-4.7-Flash-UD-Q4_K_XL.gguf (Reported: deepseek2 30B.A3B Q4\_K - Medium | 16.31 GiB) 3. GLM-4.7-Flash-Q4_K_M.gguf (Reported: deepseek2 30B.A3B Q4\_K - Medium | 17.05 GiB) >Note: The filename contains GLM-4.7, but llama-bench reads the internal GGUF header and reports deepseek2 30B.A3B. The benchmark data corresponds to a \~30B parameter MoE architecture. # Average Performance Results |Model Filename|Reported Name|Size|Avg Prompt Processing (pp512) t/s|Avg Token Gen (tg128) t/s| |:-|:-|:-|:-|:-| |GLM-4.7-Flash-MXFP4_MOE.gguf|deepseek2 30B.A3B MXFP4 MoE|15.79 GiB|258.36 t/s|11.66 t/s| |GLM-4.7-Flash-Q4_K_M.gguf|deepseek2 30B.A3B Q4\_K - Medium|17.05 GiB|218.22 t/s|12.09 t/s| |GLM-4.7-Flash-UD-Q4_K_XL.gguf|deepseek2 30B.A3B Q4\_K - Medium|16.31 GiB|160.31 t/s|13.13 t/s| (Values are arithmetic means of 3 runs. fa on = Flash Attention enabled) # Summary Analysis # 🔹 Hardware & Memory Context Device: AMD Radeon Graphics (RADV REMBRANDT) Integrated GPU Architecture: UMA (Unified Memory Access) with fp16: 1, bf16: 0, fp4: 0 Implication: The models (\~16–17 GB) exceed typical iGPU VRAM, forcing offloading to system RAM. Performance is heavily bound by system memory bandwidth (\~50–65 GB/s DDR5) and PCIe/NB link latency. The fp4: 0 flag confirms native FP4 compute is unsupported, so MXFP4 is emulated or converted at runtime. # 🔹 Prompt Processing (pp512) vs Generation (tg128) Trade-off |Format|PP Speed|TG Speed|Best Use Case| |:-|:-|:-|:-| |MXFP4 MoE|🥇 Fastest (258 t/s)|🥉 Slowest (11.66 t/s)|Long context windows, RAG, document processing| |Q4\_K\_M|🥈 Balanced (218 t/s)|🥈 Balanced (12.09 t/s)|General-purpose chat, mixed workloads| |Q4\_K\_XL|🥔 Slowest (160 t/s)|🥇 Fastest (13.13 t/s)|Fast response generation, streaming UIs| Why MXFP4 excels in PP: Despite lacking native FP4 support, the MoE structure and extreme quantization drastically reduce active compute and memory reads during attention scoring. Flash Attention further optimizes cache locality for prompt parsing. Why Q4\_K\_XL leads in TG: Generation is purely memory-bandwidth bound. The Q4\_K\_XL quantization layout appears better optimized for the RADV driver's memory prefetching, yielding \~13% faster token streaming than Q4\_K\_M and \~12% over MXFP4. # 🔹 Consistency & Stability All runs show extremely tight standard deviations (±0.02–0.06 t/s for TG), indicating stable thermal/power delivery and no background interference. * Outlier: Q4\_K\_XL's first run showed high PP variance (±17.76 t/s), likely due to cold cache/memory allocation overhead. Subsequent runs stabilized (±1.11 and ±1.54), typical of VM/page cache warmup. # 🔹 Recommendations 1. For Chat/Streaming: Use Q4_K_XL. Slightly slower prompt processing is negligible in typical conversational turns, but faster TG improves perceived latency. 2. For RAG/Long Context: Use MXFP4_MOE. The \~60% PP speed boost dramatically reduces wait times for context loading, with minor TG impact being acceptable for batched or paused workflows.

2 0 0 10/4 23:34 10/6 15:54 UTC
scorecomments10 sightings
first seen 2026-10-04 23:34 UTClast seen 2026-10-06 15:54 UTCscore then 1score now 0gained -1sightings 10
open on reddit ↗ 💬 2 (+2)