← r/LocalLLaMA
▲
19
 
1👁
r/LocalLLaMA · u/Szadbaverem69 · 6h ago

Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM

Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4\_K\_M with llama.cpp at \~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM.

Speeds start at \~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they settle down - around 90k context it stabilizes at roughly 485 tok/s prefill and 15 tok/s decode, and holds there.

The vision projector runs on CPU (--no-mmproj-offload), which keeps VRAM usage under \~5.2 GB and avoids OOM / GPU crashes. Image encoding is slower on CPU, but it buys \~1GB of VRAM.

Most MoE expert layers also run on CPU (--n-cpu-moe 39), which is how a 35B model fits in 6GB VRAM in the first place.

Launch command:

bat

@echo off

cd /d "%\~dp0"

"%\~dp0llama-server.exe" \^

\-m "C:\\Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M\\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4\_K\_M.gguf" \^

\--mmproj "C:\\Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M\\mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" \^

\--no-mmproj-offload \^

\-ngl 99 \^

\--n-cpu-moe 39 \^

\-c 131072 \^

\-np 1 \^

\-t 6 \^

\-tb 10 \^

\-b 2048 \^

\-ub 2048 \^

\-fa on \^

\-ctk q8\_0 \^

\-ctv q8\_0 \^

\--load-mode mmap+mlock \^

\--jinja \^

\--reasoning-format deepseek \^

\--reasoning-preserve \^

\--spec-type none \^

\--image-min-tokens 1024 \^

\--temp 0.6 \^

\--top-p 0.95 \^

\--top-k 20 \^

\--min-p 0 \^

\--alias Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M \^

\--host 127.0.0.1 \^

\--port 8081

pause

Hardware: RTX 2060 6GB + 32GB DDR4 RAM + i5-10400F CPU

Context: 131072 tokens, Q8\_0 KV cache

Hope this helps someone. If anyone has tips to make the launch command even better, drop them in the comments - I'm out of ideas :D

posted Sat, 10 Oct 2026 05:41:18 GMTseen 1 time
open on reddit ↗ 💬 20