← r/LocalLLaMA
▲
4
 
1👁
r/LocalLLaMA · u/mrgreatheart · 2h ago

Epyc inference rigs

I posted on here a while back asking for advice on upgrading to an Epyc system. Thanks to all who commented.

I’ve taken the plunge and ordered:

\- Epyc 7443 CPU (24 cores, 48 threads, 128 lanes)
\- H12SSL-NT-B motherboard (the variant with two 10Gb Ethernet ports)
\- 256GB 2666MHz 2Rx4 RAM

To this I will be adding 72GB of VRAM across four GPUs:

\- 3090 (24GB)
\- 5070 Ti (16GB)
\- 2 X 5060 Ti (16GB)

Unfortunately I have to wait 2 weeks for the motherboard, so naturally I’m wondering what difference it will make.

The upgrade should:

\- double the PCIe bandwidth of all four GPUs (x4 to x8 for the 5060s and x8 to x16 for the other two). They will also all be on proper CPU lanes instead of chipset and dodgy m.2 adapters.
\- increase the RAM memory bandwidth by at least 50% (96GB to 170GB/s theoretical, perhaps 140 in practice)
\- increase total RAM by 192GB (64 to 256)

And it will open up the possibility for more GPUs later.

Obviously this won’t make much difference to dense models although I’m hoping for a nice bump in prompt processing due to the increased link bandwidth.

But I’m excited about bringing my total memory pool up to 328GB and opening up access to things like GLM5.3-flash and larger quants of Qwen3.8-flash-next.

Has anyone here built something similar?

How does your system handle large MoE models that fit fully in RAM & VRAM?

posted Fri, 09 Oct 2026 13:27:55 GMTseen 1 time
open on reddit ↗ 💬 0