SentryGate: An open-source AI Gateway with sub-10ms semantic vector caching and dynamic LLM routing (Ollama & OpenAI compatible)
Here is a common problem with building apps on LLMs:
Users ask the same question over and over.
Your app calls the model every single time.
You pay the API bill every time. Users wait 3–5 seconds every time.
**\*\*SentryGate\*\* is an open-source AI traffic controller that fixes this in literally one line of code.**
\### What it actually does:
\* ⚡ \*\*Lightning Fast (4ms)\*\*: If someone asks a question that was already answered, SentryGate serves the saved answer in 4 milliseconds instead of 4 seconds.
\* 💰 \*\*$0 on Repeat Questions\*\*: Bypasses the model completely on repeat or similarly phrased prompts.
\* 🧠 \*\*Doesn't Get Tricked\*\*: Basic caches get confused between "How to bake a cake with eggs" and "How to bake a cake WITHOUT eggs". SentryGate catches tricky negative words so it never serves the wrong answer.
\* 🕒 \*\*Knows What's Fresh\*\*: Real-time questions ("today's weather", "current stock price") automatically skip the cache.
\* 🔌 \*\*Zero Downloads / 1-Line Setup\*\*: No new libraries. Just point your existing OpenAI / LangChain \base\_url\ to SentryGate and keep your code 100% untouched.
Works locally on your machine with Ollama, or in the cloud. Completely open-source under the MIT license.
\* 🌐 \*\*Test the Live Playground (No login needed)\*\*: https://sentrygate-9ght.onrender.com
\* 💻 \*\*GitHub Repo\*\*: https://github.com/Prisha2004/Sentrygate
Note: Open-source project maintainer (MIT License, 100% free).
Feedback and stars are welcome!