← r/LocalLLaMA
▲
23
+22
26👁
r/LocalLLaMA · u/3dluvr · 3d ago

Anyone working on a custom inference engine for GLM-5.3-Flash?

Seeing how Strata brings avg. 2x the performance over llama.cpp using Qwen3.8-Flash-Next, is anyone working on something similar for GLM-5.3-Flash?

After trying the GLM-5.3-Flash online couple of times and it delivering clear solutions for my use case (compared to Claude or ChatGPT), I'd love to be able to run it locally (if at all possible)...7J13/256GB/3x3090.

25 0 23 10/6 15:47 10/9 08:00 UTC
scorecomments26 sightings
first seen 2026-10-06 15:47 UTClast seen 2026-10-09 08:00 UTCscore then 1score now 23gained +22sightings 26
open on reddit ↗ 💬 15 (+15)