Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide
I'm running Qwen 3.8 27B Q4 XS with \~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.
Build llama.cpp with Vulkan:
cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j
Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then:
llama-server \
--model Qwen3.8-27B-UD-IQ4_XS.gguf \
--mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \
--n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \
--batch-size 2048 --ubatch-size 512 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \
--load-mode none --fit off \
--cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \
--no-context-shift --jinja --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080
---
Edit: The process is documented here: https://zenodo.org/records/23088880. Feedback will be integrated into the upcoming Qwen 4 setup.