What is everyone using to serve + monitor models across multiple GPUs/nodes? Trying to cut down on duct tape
Running inference across three machines - an AMD box (Ryzen 9 9950X + RX 7900 XTX), a smaller NVIDIA box (i5-10400 + RTX 3050), and a MacBook Pro M3 - all running LM Studio/Ollama. The serving/monitoring side is where I lose the most time: no single place to see what model/version is loaded where, token throughput, VRAM vs unified-memory pressure, etc., without checking each machine by hand.
What are you all actually using for:
- Multi-node / multi-GPU serving + routing?
- Observability that is not "wire up Prometheus on every box"?
- Keeping track of model versions across nodes?
Happy to share my current setup if it is useful.
scorecomments9 sightings
first seen 2026-10-07 01:49 UTClast seen 2026-10-08 07:28 UTCscore then 1score now 3gained +2sightings 9
open on reddit ↗
💬 13 (+13)