← r/LocalLLaMA
▲
3
+2
7👁
r/LocalLLaMA · u/kshitizsriv · 2d ago

Local embeddings and rerankers vs a hosted LLM for catalog matching?

I’m building a feature that matches free-form requests to a catalog of structured listings. Requests can contain several constraints and follow-up refinements. The results also need a short explanation of why each match was selected.

Our prototype uses a hosted LLM to rank a shortlist. I’m exploring whether a small locally hosted embedding model and reranker could deliver comparable quality at lower cost.

For anyone who has deployed a similar system: where did local retrieval start to fall short of an LLM? Did a hybrid approach work better? I’d especially appreciate real-world latency and cost figures around 10,000–100,000 requests per month.

3 0 3 10/7 05:16 10/7 19:23 UTC
scorecomments7 sightings
first seen 2026-10-07 05:16 UTClast seen 2026-10-07 19:23 UTCscore then 1score now 3gained +2sightings 7
open on reddit ↗ 💬 9 (+8)