Local embeddings and rerankers vs a hosted LLM for catalog matching?
I’m building a feature that matches free-form requests to a catalog of structured listings. Requests can contain several constraints and follow-up refinements. The results also need a short explanation of why each match was selected.
Our prototype uses a hosted LLM to rank a shortlist. I’m exploring whether a small locally hosted embedding model and reranker could deliver comparable quality at lower cost.
For anyone who has deployed a similar system: where did local retrieval start to fall short of an LLM? Did a hybrid approach work better? I’d especially appreciate real-world latency and cost figures around 10,000–100,000 requests per month.
scorecomments7 sightings
first seen 2026-10-07 05:16 UTClast seen 2026-10-07 19:23 UTCscore then 1score now 3gained +2sightings 7
open on reddit ↗
💬 9 (+8)