← r/LocalLLaMA
▲
0
-1
7👁
r/LocalLLaMA · u/Bulky-Priority6824 · 5d ago

QFN llama.cpp Any juice left to squeeze?

https://imgur.com/a/Ef2xyNu

Using this squished down ISTA model on 2x 5060ti 16gb and 32gb ddr4 ram I'm wondering if my settings are correct as I cant really find much consistent feedback for this model on this particular hardware.

What are people running in their config?

Qwen 3.8 FN GSQ RCO IQ1

|Field|Value|
|:-|:-|
|Name|Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002|
|Display|Qwen 3.8 FN GSQ RCO IQ1|
|Path|/opt/models/Qwen3.8-Flash-Next-GSQ-RCO-IQ1_M-00001-of-00002.gguf|
|Size|27.58 GB|
|llama backend|default|

Launch args

|Flag|Value|
|:-|:-|
|--host|10.210.44.126|
|--port|11434|
|--ctx-size|98304|
|--cache-type-k|q8_0|
|--cache-type-v|q8_0|
|--override-tensor|per_layer_token_embd=CPU|
|--gpu-layers|999|
|--load-mode|mmap+mlock|
|-fa|on|
|-b|2048|
|-ub|256|
|--temp|0.7|
|--min-p|0.05|
|--top-p|0.95|
|--top-k|20|
|--main-gpu|0|
|--parallel|1|
|--threads|8|
|--reasoning-format|deepseek|
|--reasoning-effort|medium|
|--reasoning|on|
|-sm|tensor|
|--tensor-split|1,1|
|--repeat-penalty|1.05|
|--presence-penalty|0|
|--fit|off|
|--alias|QFN|
|--n-cpu-moe|8|

Bench

|Metric|Value|
|:-|:-|
|Prompt|250.4 tok/s|
|Generation|30.5 tok/s|
|Config|tensor 1,1|
|Date|2026-10-04 16:36 UTC|

#

1 0 0 10/4 23:34 10/8 22:03 UTC
scorecomments7 sightings
first seen 2026-10-04 23:34 UTClast seen 2026-10-08 22:03 UTCscore then 1score now 0gained -1sightings 7
open on reddit ↗ 💬 11 (-3)