Sharing my latest project: strata-mlx
Strata is impressive: it runs Qwen3.8-Flash-Next, a 125B-parameter mixture-of-experts model, on a 12 GB graphics card. But its authors have said macOS is out of scope.
I still wanted it on the Mac.
So I spent a day and a half and wrote strata-mlx. It is an unofficial MLX engine, modelled on Strata's design, that reads Strata's GGUF files directly, without changing any weights.
I developed it on a 128 GB M4 Max MacBook Pro. After one experiment after another, I finally had some results.
All numbers below are tokens/second, higher is better; Q2\_0 and IQ3\_S are two files of the model, 66 GB and 84 GB.
https://preview.redd.it/yyn57wt5kluh1.png?width=526&format=png&auto=w…
The draft head is a module that comes with the model: it guesses the next few tokens and the model verifies them in one pass. Neither of the other two engines uses it.
What I set out to do: not every Mac has 128 GB of RAM. On a Mac that can't hold the whole model, strata-mlx keeps as many experts in memory as fit and reads the rest from the SSD as tokens need them. By the engine's own estimate, a 48 GB Mac holds the Q2\_0 file whole and never touches the disk; with less memory it reads from the SSD as it goes.
I tested this by capping my Mac to act as a smaller one: the 66 GB file runs inside a 16 GB machine's memory at 6–7 tokens/second, a 24 GB machine's at 19–23, and a 32 GB machine's at 28–47. In that mode, reading a 1,537-token prompt runs at 273–287 tokens/second.
So a small Mac can run a model much bigger than its RAM, but there is no guarantee about speed.
I hope people will join in and test how well it runs LLMs on Macs with different amounts of memory. That's how I can keep improving it.
Code, experiment logs, and the open problems: