Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM
Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4\_K\_M with llama.cpp at \~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM.
Speeds start at \~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they settle down - around 90k context it stabilizes at roughly 485 tok/s prefill and 15 tok/s decode, and holds there.
The vision projector runs on CPU (--no-mmproj-offload), which keeps VRAM usage under \~5.2 GB and avoids OOM / GPU crashes. Image encoding is slower on CPU, but it buys \~1GB of VRAM.
Most MoE expert layers also run on CPU (--n-cpu-moe 39), which is how a 35B model fits in 6GB VRAM in the first place.
Launch command:
bat
@echo off
cd /d "%\~dp0"
"%\~dp0llama-server.exe" \^
\-m "C:\\Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M\\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4\_K\_M.gguf" \^
\--mmproj "C:\\Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M\\mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" \^
\--no-mmproj-offload \^
\-ngl 99 \^
\--n-cpu-moe 39 \^
\-c 131072 \^
\-np 1 \^
\-t 6 \^
\-tb 10 \^
\-b 2048 \^
\-ub 2048 \^
\-fa on \^
\-ctk q8\_0 \^
\-ctv q8\_0 \^
\--load-mode mmap+mlock \^
\--jinja \^
\--reasoning-format deepseek \^
\--reasoning-preserve \^
\--spec-type none \^
\--image-min-tokens 1024 \^
\--temp 0.6 \^
\--top-p 0.95 \^
\--top-k 20 \^
\--min-p 0 \^
\--alias Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M \^
\--host 127.0.0.1 \^
\--port 8081
pause
Hardware: RTX 2060 6GB + 32GB DDR4 RAM + i5-10400F CPU
Context: 131072 tokens, Q8\_0 KV cache
Hope this helps someone. If anyone has tips to make the launch command even better, drop them in the comments - I'm out of ideas :D