← r/LocalLLaMA
▲
7
+2
6👁
r/LocalLLaMA · u/Then_Blueberry7290 · 12d ago

LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF

Just recently stubled upon with this modell:LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF I'm just stay away from "magic" models, but this model size got my eyes on: With vision capabilities this is under 17GB, which means i can use it 32GB vram with full Context size (262k), bigger ubatch, and mtp4. Of course vision goes to ram, not gpu. Other similar model with nvfp4 line, usually 19-20GB in size or more. I tried in with llama.cpp, speed is 40-113 t/s (76 in my benchmark) with 262k context. Under normal agentic workin it is 45-65 t/s. (2x5060ti16GB OC) For example thinkingcap nvfp with vllm i can only have 160k context (cannot offload mmproj to ram) First glance it is the same as the other swift models (nvfp4) in quality. So My question is what is the tradeoff of this modell?

7 0 7 10/3 06:31 10/4 07:38 UTC
scorecomments6 sightings
first seen 2026-10-03 06:31 UTClast seen 2026-10-04 07:38 UTCscore then 5score now 7gained +2sightings 6
open on reddit ↗ 💬 11