← r/LocalLLaMA
▲
3
 
6👁
r/LocalLLaMA · u/harrythunder · 13d ago

DeepSeek-V4.1-Flash split across M5 Ultra and 2× RTX PRO 6000

DeepSeek-V4.1-Flash's prompt state is only 0.9 KB/token, so you can split it at layer 20 across CUDA/Metal. All you need is 1/10GbE network. FYI. https://tacos8me.github.io/m5-ultra/split/

4 0 3 10/3 06:31 10/5 07:56 UTC
scorecomments6 sightings
first seen 2026-10-03 06:31 UTClast seen 2026-10-05 07:56 UTCscore then 3score now 3gained 0sightings 6
open on reddit ↗ 💬 2