← r/LocalLLaMA
▲
11
+1
10👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 10d ago

What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects?

For example, directly using claude code or code, which is then hooked up to automatically delegate the actual code writing tasks to a local model like qwen 3.8 flash next, to save on cloud usage limits.

I’m imagining the loop would be:
User writes prompt
Claude/codex thinks about it and the plan
Claude/codex sends the specific and bounded coding instructions to the local model+harness (opencode, pi, etc) via api endpoint or MCP, with clear instructions on a defined endpoint
One the local model+harness hits the clear endpoint/“done” step, it sends a ping back to claude/codex
Claude/codex then verifies the output and then thinks about next steps to instruct the local model+harness on

Does this actually lead to improved savings on the cloud model usage while preserving code quality? Or does this end up being unnecessarily complex and not saving on any cloud usage

11 0 11 10/3 06:30 10/7 07:58 UTC
scorecomments10 sightings
first seen 2026-10-03 06:30 UTClast seen 2026-10-07 07:58 UTCscore then 10score now 11gained +1sightings 10
open on reddit ↗ 💬 21 (-1)