Is Strix Halo (GMKtec EVO-X2, etc.) the closest thing we have to a "dream" local LLM box?
I've been looking at the <32B model space and keep coming back to an interesting question.
A few years ago, projects like Hummingbird+ suggested that cheap custom accelerators (FPGA-based) might become the future of local inference. But today it seems like memory capacity is still the real bottleneck rather than raw TOPS.
For someone who wants to run modern 20B-32B models at reasonable quants (Q5/Q6 rather than INT4), the options all seem compromised:
- Consumer GPUs have great bandwidth but limited VRAM.
- NPUs and AI accelerators often have lots of compute but not enough memory.
- FPGA solutions are fascinating but still bandwidth-constrained.
- Strix Halo systems (GMKtec EVO-X2, Framework Desktop, etc.) offer huge unified memory pools, but they're expensive.
The "dream" accelerator would be something like:
48+ GB memory
500+ GB/s bandwidth
under $1000
reasonable power consumption
...but I don't think anything like that actually exists yet.
For those who have used Strix Halo systems for local inference:
How do they feel with current 20B-32B models?
Do you regret not buying a used 3090/4090-based machine instead?
Is unified memory a bigger advantage in practice than benchmarks make it seem?
Curious what people who own both types of systems think.