← r/LocalLLaMA
▲
2
 
18👁
r/LocalLLaMA · u/Adorable-Cost-3249 · 8d ago

Qwen3.8-27B Q4_K_M on one RTX 3090 + OpenCode: throughput, four coding tasks, and a reasoning-budget failure

I put an old RTX 3090 to work as a local coding agent with Qwen3.8-27B, llama.cpp, and OpenCode. Here are the setup and results, including what failed. This is a summary of my own blog post, linked below.

Setup

  • RTX 3090 24GB, Ryzen 7 5800X, 64GB RAM, Ubuntu.
  • Qwen3.8-27B Q4\_K\_M weights (\~16.8GB), all layers on GPU.
  • llama.cpp b11146, CUDA 12.8, flash attention, q8\_0 K/V cache, one generation slot.
  • 131,072-token context capacity; 8,192-token output allowance per response. Input and output share context, and reasoning uses the output allowance.
  • OpenCode 2.0.20 connected to llama-server's OpenAI-compatible API at http://127.0.0.1:8080/v1. OpenCode reads/edits files and runs tests; llama-server handles inference. Chat, tool-call round trips, and streamed tool calls worked in our checks.

Speed: fresh input versus a cached continuation

|Actual input|Generation|First token, fresh|First token, cached|
|:-|:-|:-|:-|
|2,073 tokens|36.4 tok/s|2.75 s|0.46 s|
|16,378 tokens|33.5 tok/s|17.00 s|0.47 s|
|65,537 tokens|25.9 tok/s|83.41 s|0.51 s|
|120,011 tokens|20.9 tok/s|183.01 s|0.63 s|

These throughput runs disabled thinking. The 2K row is the median of three fresh requests; larger rows have one fresh request and one continuation each. Cached continuations processed only 27–28 new input tokens, reusing almost the entire prefix. The subsecond figures depend on that reuse; they don't describe a new 120K prompt.

Peak sampled total GPU memory use was 22,162 MiB, including desktop use. It fit, with limited headroom. A separate \~120K synthetic retrieval check passed, but we did not evaluate coding quality at that length.

Four bounded Python coding tasks

Each task had a fresh session, medium thinking, an eight-minute deadline, and ten independent test methods kept outside the agent's workspace. First attempts ran serially without cloud fallback or network tools.

|Task|Independent checks, before → after|Outcome|
|:-|:-|:-|
|Expiring LRU cache|0/10 → 10/10|Completed in \~3m07s; strongest result|
|CSV ledger/refunds|1/10 → 10/10|Completed in \~5m44s; later review found gaps|
|Incremental build planner|1/10 → 1/10|No edits; exhausted its response allowance|
|Atomic SQLite transfers|1/10 → 10/10|Candidate passed, but timed out before final test rerun and handoff|

Three candidates passed the predefined checks; two completed the whole workflow within the deadline. The aggregate 31/40 includes one baseline pass from the unchanged build planner and is not a general coding success rate.

The build planner was the interesting failure: about 4,985 input tokens, then 8,192 output tokens entirely spent on reasoning, ending with length and no patch. This was an output-budget failure far below the context limit. A separate diagnostic with thinking disabled completed in 5m40s and passed 9/10 independent checks. That was one additional run at temperature 1, not evidence that disabling thinking is universally better.

Passing tests also missed defects. Further ledger review found Decimal rounding at a large numerical boundary and an unhandled I/O error. The wallet's own concurrency tests actually ran sequentially, and a separate boundary probe found SQLite converting an overflowing balance to REAL while recording success. Those later probes were not retroactively added to the forty checks.

For me, the useful workflow is a bounded task with clear acceptance criteria, followed by diff review and independent checks. I would repeat these tasks across thinking settings and response budgets before drawing stronger conclusions.

My full post, configuration, and measurement links. The downloadable kit contains the launcher, OpenCode configuration, throughput script, and records; it does not include model weights or the complete coding-task fixtures.

For others using a 24GB card with OpenCode: what reasoning setting and per-response output budget have worked best for bounded coding tasks?

The numbers and failure cases come from the linked experiment records.

3 0 2 10/3 06:28 10/8 02:03 UTC
scorecomments18 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-08 02:03 UTCscore then 2score now 2gained 0sightings 18
open on reddit ↗ 💬 14 (+3)