on prem LLM stack for data that cant leave the building
Running the model locally is not a problem thats easy part but the leaks are the third party integrations along with it, like you designed everything perfect and then just added a cloud api along with it maybe a hosted judge for evals or a tracing saas or embedding point. one http call and the on prem things over
parts we already keep local are
- runtime: llama.cpp/ vllm /ollama
- models: qwen or llama family depending on rig
- vector db: pgvector or qdrant
some that leak but remain unnoticed:
- ingestion: pdfs and scans for some projects need a parse and the ocr step before chunking them and its where we often reach for a cloud parser and break the rule, although we can keep it local with liteparse sort of inbound parsers or other open source options on huggingface
- Eval: plenty of local setups still need prompts and outputs to a hosted judge or a tracing dashboard to see quality which is the same leak but seems different. instead a local score set or a local judge model and keeping it self hosted if possible handles the tracing part
I am curious to know about others end to end stack who keep it 100% local, eager to learn more
scorecomments16 sightings
first seen 2026-10-03 06:29 UTClast seen 2026-10-07 06:39 UTCscore then 7score now 6gained -1sightings 16
open on reddit ↗
💬 27 (+3)