← r/LocalLLaMA
▲
6
 
1👁
r/LocalLLaMA · u/Edenar · 5h ago

Strix halo + CMP 170HX setup

A bit of backstory : i got lucky in August and was able to grab a CMP 170 HX 8GB for around 500$ on alibaba. I tried later to get a second one for around 1k$ but my order got cancelled and prices went to the moon...
I unlocked it with https://github.com/amoghmunikote/cmpunlocker without any issue to 64GB and restored compute capabilities.
I also got a framework desktop 128GB last year (was 2.2k$ at that time...) that i use mainly for linux at home and ai workload (i went from gpt-oss-120b to qwen 3.5 122b/qwen 3.6 35B and later 3.8 FN as my main models).
The GPU sits on an USB4 eGPU dock and is limited to pcie2x4
I also got my hands on an optane SSD (P5800X 1.6 TB),>! i tried to stream the ngram table from it, doesn't really change anything compared to streaming from a sn850x (pcie 4x4 nvme ssd)!<

At first i wanted to run bigger/better models by adding 120GB-ish iGPU to the 64GB eGPU like ds v4 flash, qwen 3.8 FN at Q8\_K\_XL or even DS flash 4.1. But using 2 differents arch and vendor together (amd gfx1151 and CUDA with SM 80) is not that easy, even with llama.cpp.

So i settled on another setup : On the strix Halo i run qwen 3.8 flash next (halogen or q4 k xl with gufo, tried both they are mostly the same quality. Halogen is a bit faster in agentic turns). I get 1600-1800 tok/s pp and 60 tok/s tg for usual agentic work. And then i use the CMP 170 HX to run 2 concurrent qwen 3.8 27b (lued's w8a16 quant) with around 200k bf16 context each. Aggregate perf at low context are 2000tok/s pp and 150-300 tok/s tg (Dflash 2 acceptance rate varies a lot). For a more realistic usage, my last session with 7.2M token in and 1.1M token out averaged at 1747 tok/s pp and 100 tok/s tg aggregate in \~4h (those stat are only the 2 27b sessions).
I use Pi agent (main agent is qwen 3.8 FN on the strix halo) and the 2 qwen 27b sessions as subagent to delegate some tasks.

This final setup (at least before qwen 4 lineup drops) made me drop cloud usage entirely. When fully loaded the whole thing is probably around 400W (strix halo + 250W eGPU). Entire cost with the dock, import fees and a fan+shroud for the card is around 3k$. Right now i guess cheapest 128GB strix halo are around 3k$ and cmp 170hx 2k$+ so you would need 5-6k$. It could still be an alternative to the overpriced spark if you want to trade power usage and compacity for throuput while still keeping the whole thing reasonable in power and volume.

On the intelligence side, it's better than my experiences with sonnet (i havent tested last 5.5 in the best modes so maybe they are better but sonnet 4.7 and free 5.5 are far less reliable than my local setup) but below opus ofc. I usually ask the main agent to act as an architect and to review what the subagents produce. With websearch, python and a few other tools available the stack feels kinda clever (i mostly do devops stuff : ci/cd pipeline, container deployment , sometime a bit of dev but mostly to fork something existing). I still need to sometime steer the main agent manually and like 1/100 toolcall can fail and needs a retry but overall i need 20 min of work to do what would have taken me a whole day 3 years ago.

I'm curious if some of you run similar setup and what did you achieve (i'm still thinking about trying to fit DS 4.1) !

disclaimer : no AI involved in writing this post and my english sucks, i know.

posted Fri, 09 Oct 2026 22:32:32 GMTseen 1 time
open on reddit ↗ 💬 7