← r/LocalLLaMA
▲
1102
+1
33👁
r/LocalLLaMA · u/Thin_Pollution8843 · 26d ago

3k$ 128GB VRAM + 256GB RAM DDR4 Server

post image

I finished my home inference server. First I tried Lenovo p620 workstation and while it’s a good value overall it pissed me off with a ton of proprietary Lenovo shit to deal with and I return it in the end.

Components:

4xV620 - 1400$

256GB DDR4 RDIMM 2666 - 610$

Huanandzhi D12D - 410$

EPYC 7452 - 170$

PSU ASRock 1600 - 220$

SSD Samsung 970EVO 1tb - Already had

Case//Fans//Misc \~ 200$

Power consumption is no shit ofc on such machine:

700-900w prefill
500-600w decode on Qwen3.8-next-flash Autoround W4A16

What it can do -

EDIT: Qwen3.8-next-flash Autoround W4A16 1.3k prefill and 70tg code/60tg prose on 128k+ context with MTP-2 on vllm fork.

I was disappointed with this machine and qwen3.8-27b speeds at first. But since Qwen3.8 next running good on it - I’m satisfied. Hope in more optimizations in future.

1110 0 1102 10/2 17:26 10/8 18:40 UTC
scorecomments33 sightings
first seen 2026-10-02 17:26 UTClast seen 2026-10-08 18:40 UTCscore then 1101score now 1102gained +1sightings 33
open on reddit ↗ 💬 288