infill and output tok/s speeds after "new" build compared to old questions
Let me start off by saying I am new to AI and have a lot to learn, basically I don't know shit. I started off a few weeks ago by throwing my 2 5090s I had from gaming pcs into a 9950x cpu with 64gb ddr5 6000 ram on a motherboard that was able to do gen 5 8x per card system. Running ubuntu 24.04 server and running unsloth studio, swift 1.5 qwen 3.8 27B Q8 with kv cache dtype at q8\_0 and 262k context I was getting 2500-3000 infill and 100-150 toks/s output.
I wanted the 5090s in my rack in the basement in my proxmox server, it has a 7402p cpu (24 core 48 thread (rome)). 256 gb ddr4 3200 ECC ram (8 channel) on a supermicro H12SSLNTO motherboard. I have the 5090s passed through (gen 4 16x each card) to a vm (using 128gb of ram, direct access, no ballooning etc) and a dedicated 1.8 tb nvme drive passed through dedicated for the ai server vm (actually the vm itself is using a pool on the proxmox server but all the ai stuff is sitting on and running from the 1.8tb nvme).
Everything is working okay. It is running ubuntu 26.04 server. I have unsloth studio running, running the same model and settings, infill is more, up to 3800 but output is like half or less around 60 tok/s. Ideas what may be causing the drop in tok/s output and what to look into if a system issue?
I have done a lot of memory bandwidth tests (theoretical is \~204 GB/s, double that of ddr5 dual channel) but from what I can find and because I only have a 4 ccd cpu I am only getting 90-120 GB/s memory bandwidth. I guess I can get that to the 160-180 range if I go with a 64 core 8 ccd cpu, which I am considering doing..but I don't even know if that has anything to do with anything, just a rabbit hole I went down.
Ideas what to look into for the drop in output tok/s?