← r/LocalLLaMA
▲
68
+57
28👁
r/LocalLLaMA · u/Yaniss916 · 6d ago

Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395.

|Model|GLM-5.3-Flash|MiMo-V2.6-Flash-MOPD|
|:-|:-|:-|
|Size|99.7 GB|105 GB|
|Prefill|580 tok/s at 3.5K, 546 at 64K|about 650 tok/s at 4K|
|Decode|26 to 30 tok/s (MTP)|32 prose / 35 chat / 44 code (speculative), 29 plain|
|KLD vs official FP8|0.151|0.0713|
|Top-1 agreement with FP8|89.3 %|92.0 %|

Where the weights come from. MiMo is our own quantisation. The GLM pack mixes turboderp's public 2.05 and 3.05 bpw EXL3 tensors, with our layer mix and a small tuning stage. On the same 129 rows, his 2.05 bpw pack (85 GB) gets KLD 0.275; our mix (100 GB) gets 0.190. His is smaller and decodes about 10 % faster.

Uncensored variants. Separate -Uncensored repos: same weights plus one small file the engine applies at load, one switch turns it off.

Not measured yet. Task-suite scores for MiMo, GLM at 128K context, any GPU other than gfx1151. The conversion pipeline stays private.

Quickstart. Clone, ./build.sh, hf download yamz-labs/GLM-5.3-Flash-EXL3-Yamz, python tools/glm/serve.py --model ./glm-pack -c 131072 --num-draft 2. You get an OpenAI-style API.

Models: https://huggingface.co/yamz-labs

Engine: https://github.com/Yamz-Labs/kyojin

Built on turboderp's ExLlamaV3, with ROCm work from sdougbrown and vcruz305.

If you own a Strix Halo machine, we'd love to see your tok/s. Issues, benchmarks and PRs are all welcome. Which model should we do next?

68 0 68 10/3 15:28 10/8 19:23 UTC
scorecomments28 sightings
first seen 2026-10-03 15:28 UTClast seen 2026-10-08 19:23 UTCscore then 11score now 68gained +57sightings 28
open on reddit ↗ 💬 58 (+44)