Qwen 3.8 with Pi harness constantly hallucinates that it is out of context?
With Qwen 3.8 Flash Next (FP8 on VLLM) on a fairly stock Pi harness, it constantly hallucinates some measure of available context that says it is almost out. It's to the point where it frequently refuses work or stops in the middle of something, claiming it shouldn't go any further because it's almost out of context, when I can see in the harness status bar that (256K) context is ~25% used.
When I ask how it determined that, it always says it "invented the number and the treated it as real data" or guessed, and that it'll stop doing that, but it keeps happening.
Is there anything in particular that would cause this?
Thanks!
scorecomments18 sightings
first seen 2026-10-05 13:40 UTClast seen 2026-10-08 22:03 UTCscore then 1score now 0gained -1sightings 18
open on reddit ↗
💬 36 (+31)