← r/LocalLLaMA
▲
1
 
2👁
r/LocalLLaMA · u/mrgreatheart · 15d ago

Epyc for inference

Hi. My current system is an Intel Ultra 7 with 64Gb DDR5 at 6000. It has 4 GPUs totalling 72Gb: A 3090 and 5070 Ti on x8 CPU connected PCIe slots plus two 5060 Ti - one on an x4 CPU connected m.2 socket, and the other on an x4 chipset connected PCIe slot. I am considering upgrading to an Epyc 7443 system with 256Gb of DDR4 8 channel. Given prices I’ll probably end up with 2400 speed sticks giving a theoretical memory bandwidth of 153Gb/s. The obvious benefits are getting all the GPUs on proper CPU connected x16 and x8 PCIe with room for another at some point plus enough RAM to overflow bigger models than I can fit in VRAM. I’m struggling to find clear information on whether this would actually be worth it for the cost. I am currently able to run Qwen3.8-flash-next in two rather painful configurations (not using the chipset connected 5060 because it hurts too much): \- IQ4\_XS at 300 pp / 40 gen in llama.cpp \- EXL3.05bpw at 1,200 pp / 25 gen in exllamav3 The low tg in exllamav3 appears to be because I have to offload 10 layers to CPU. Obviously the extra PCIe slots would bring the other 5060 into play but I’d like to know what possibilities the extra RAM and bandwidth would open up. Is anyone here offloading larger quants (or other large models) to CPU on Epyc, and if so how usable is it?

1 0 1 10/3 06:32 10/3 13:39 UTC
scorecomments2 sightings
first seen 2026-10-03 06:32 UTClast seen 2026-10-03 13:39 UTCscore then 1score now 1gained 0sightings 2
open on reddit ↗ 💬 50