← r/LocalLLaMA
▲
38
+7
32👁
r/LocalLLaMA · u/SultanGreat · 8d ago

What's the best setup for Qwen3.8 27b for a 16 gig VRAM?

Hello guys!

I have been experimenting with qwen 3.8 for a long time and I hadn't been able to get reasonable speed. I am on a 5060Ti 16 GB, and although this gpu can game, I am aware that AI demands more than 16 GB.

I am on a Fedora 44, AMD Ryzen 9600x and 16 GB system ram (16 GB system ram and 16 GB vram, totaling to 32 GB) and I would like to use llamacpp, although I would use any other tool if I could if it meant faster speed.

I am looking for a large context. Atleast 128k context. The first question is, what quantization to pick? In my experience Q3 UD was satisfying, but I am looking for uncensored model. In my experience, MTP has never lived up to its hype for me (and I don't know why!?), which is why I am thoroughly lost on making a good setup after an honest week of experimentation, which is why I have resorted to ask here as a last resort.

Update : Found a model, thanks to u/_wortkarg_

link : https://huggingface.co/RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3\_XXS-Uncensored

command (A better command would be appreciated and updated accordingly):

~/llama.cpp/build/bin/llama-server \
--model ~/Documents/Models/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf \
--alias "llamacpp" --host 0.0.0.0 --port 8001 \
-ngl 99 --flash-attn on --ctx-size 131072 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--parallel 1 --batch-size 512 --ubatch-size 256 \
--no-warmup --jinja \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

I am hitting at about 35 t/s+ speed with this one.

39 0 38 10/3 06:28 10/8 12:10 UTC
scorecomments32 sightings
first seen 2026-10-03 06:28 UTClast seen 2026-10-08 12:10 UTCscore then 31score now 38gained +7sightings 32
open on reddit ↗ 💬 84 (+3)