← r/LocalLLaMA
▲
3
+2
9👁
r/LocalLLaMA · u/ziyaulhuk12 · 2d ago

What is everyone using to serve + monitor models across multiple GPUs/nodes? Trying to cut down on duct tape

Running inference across three machines - an AMD box (Ryzen 9 9950X + RX 7900 XTX), a smaller NVIDIA box (i5-10400 + RTX 3050), and a MacBook Pro M3 - all running LM Studio/Ollama. The serving/monitoring side is where I lose the most time: no single place to see what model/version is loaded where, token throughput, VRAM vs unified-memory pressure, etc., without checking each machine by hand.

What are you all actually using for:

  • Multi-node / multi-GPU serving + routing?
  • Observability that is not "wire up Prometheus on every box"?
  • Keeping track of model versions across nodes?

Happy to share my current setup if it is useful.

5 0 3 10/7 01:49 10/8 07:28 UTC
scorecomments9 sightings
first seen 2026-10-07 01:49 UTClast seen 2026-10-08 07:28 UTCscore then 1score now 3gained +2sightings 9
open on reddit ↗ 💬 13 (+13)