← r/LocalLLaMA
▲
4
 
1👁
r/LocalLLaMA · u/yibie · 4h ago

Sharing my latest project: strata-mlx

Strata is impressive: it runs Qwen3.8-Flash-Next, a 125B-parameter mixture-of-experts model, on a 12 GB graphics card. But its authors have said macOS is out of scope.

I still wanted it on the Mac.

So I spent a day and a half and wrote strata-mlx. It is an unofficial MLX engine, modelled on Strata's design, that reads Strata's GGUF files directly, without changing any weights.

I developed it on a 128 GB M4 Max MacBook Pro. After one experiment after another, I finally had some results.

All numbers below are tokens/second, higher is better; Q2\_0 and IQ3\_S are two files of the model, 66 GB and 84 GB.

https://preview.redd.it/yyn57wt5kluh1.png?width=526&format=png&auto=w…

The draft head is a module that comes with the model: it guesses the next few tokens and the model verifies them in one pass. Neither of the other two engines uses it.

What I set out to do: not every Mac has 128 GB of RAM. On a Mac that can't hold the whole model, strata-mlx keeps as many experts in memory as fit and reads the rest from the SSD as tokens need them. By the engine's own estimate, a 48 GB Mac holds the Q2\_0 file whole and never touches the disk; with less memory it reads from the SSD as it goes.

I tested this by capping my Mac to act as a smaller one: the 66 GB file runs inside a 16 GB machine's memory at 6–7 tokens/second, a 24 GB machine's at 19–23, and a 32 GB machine's at 28–47. In that mode, reading a 1,537-token prompt runs at 273–287 tokens/second.

So a small Mac can run a model much bigger than its RAM, but there is no guarantee about speed.

I hope people will join in and test how well it runs LLMs on Macs with different amounts of memory. That's how I can keep improving it.

Code, experiment logs, and the open problems:

https://github.com/yibie/strata-mlx

posted Sat, 10 Oct 2026 08:03:30 GMTseen 1 time
open on reddit ↗ 💬 1