SAGG — turning unreliable Gonka brokers into a reliable inference API (cascading failover, real data)
If you've used Gonka inference directly, you've probably noticed individual brokers aren't always consistent — one might be fast and reliable for a while, then slow down or drop requests, then recover. That's just how a decentralized network of independent nodes behaves.
SAGG takes a different approach: instead of relying on one broker and hoping it stays healthy, it holds several at once and automatically routes around whichever one is struggling at that moment. From the outside, you just get a normal, reliable API — the instability gets absorbed before it ever reaches you.
We didn't just build this and claim it works — we measured it properly, on real, sustained production traffic, two separate campaigns:
September 18 (1000 requests/line, standard prompt mix):
\- Standard line: 100% success (1000/1000 requests)
\- Super Deal line: 98.9% success (989/1000 requests)
September 30 recheck (200 requests/line, heavier prompt mix - longer context, code generation):
\- Standard line: 99.5% success (199/200 requests)
\- Super Deal line: 99.0% success (198/200 requests)
TTFT p50: \~190-490ms depending on line and load, p95 under 15s on heavier workloads.
Full methodology, raw data, and a reproduction script: github.com/privatedeskai/sagg-benchmark-data
For the technically curious: the hard part wasn't picking a backup broker — it was streaming responses specifically. Once a provider starts sending content to the client, you can't silently switch mid-stream without