← r/LocalLLaMA
▲
6
-1
16👁
r/LocalLLaMA · u/lucasbennett_1 · 9d ago

on prem LLM stack for data that cant leave the building

Running the model locally is not a problem thats easy part but the leaks are the third party integrations along with it, like you designed everything perfect and then just added a cloud api along with it maybe a hosted judge for evals or a tracing saas or embedding point. one http call and the on prem things over

parts we already keep local are

  1. runtime: llama.cpp/ vllm /ollama
  1. models: qwen or llama family depending on rig
  1. vector db: pgvector or qdrant

some that leak but remain unnoticed:

  1. ingestion: pdfs and scans for some projects need a parse and the ocr step before chunking them and its where we often reach for a cloud parser and break the rule, although we can keep it local with liteparse sort of inbound parsers or other open source options on huggingface
  1. Eval: plenty of local setups still need prompts and outputs to a hosted judge or a tracing dashboard to see quality which is the same leak but seems different. instead a  local score set or a local judge model and keeping it self hosted if possible handles the tracing part

I am curious to know about others end to end stack who keep it 100% local, eager to learn more

9 0 6 10/3 06:29 10/7 06:39 UTC
scorecomments16 sightings
first seen 2026-10-03 06:29 UTClast seen 2026-10-07 06:39 UTCscore then 7score now 6gained -1sightings 16
open on reddit ↗ 💬 27 (+3)