poorman inference engine for 16GB GPU and 35B moe Qwen 3.6for coding
I forked llama.cpp's server into AgrillaMoE, a dedicated build for Qwen3.6-35B-A3B (\~A4B) with Unsloth quants. On a (vant.ai) rented V100 16GB with the 2-bit UD-Q2\_K\_XL quant it generates at \~57-60 tok/s while running the full MoE-expansion profile — and it speaks both the OpenAI and Anthropic APIs, so Claude Code just works against it.
What is MoE expansion? Qwen3.6-35B-A3B has 8 routed experts active per token. The expansion patch raises that budget at runtime — no retraining, no file changes: --moe-experts 20 with an adaptive threshold keeps experts while p >= 0.8 × p(rank 8), applied to layers 25-39. You're literally consulting more of the 35B parameters per token — that's where the "retrieved intelligence" comes from, on GPQA-Diamond with Q8\_0 it scored 84.34% vs 81.82% stock top-8 (+2.5 pts) (miticooo!).
Same weights, better routing.
https://github.com/vagrillo/AgrillaMoE/blob/main/gpu16gbguide.md