DeepSeek-V4.1-Flash split across M5 Ultra and 2× RTX PRO 6000
DeepSeek-V4.1-Flash's prompt state is only 0.9 KB/token, so you can split it at layer 20 across CUDA/Metal. All you need is 1/10GbE network. FYI. https://tacos8me.github.io/m5-ultra/split/
scorecomments6 sightings
first seen 2026-10-03 06:31 UTClast seen 2026-10-05 07:56 UTCscore then 3score now 3gained 0sightings 6
open on reddit ↗
💬 2