← r/LocalLLaMA
▲
171
-3
23👁
r/LocalLLaMA · u/politefella0 · 13d ago

Qwen 3.8 flash next is based on Qwen 4 architecture, if the announced Qwen 4 27b is also the same architecture with n-grams does it mean I can actually have faster inference on a single 3090 without tweaking much?

I wish Qwen also released dataset and method to fully train a model ourselves but it is what it is. However, I come here with my stupid question because someone can answer it better.

And will the model still be an over thinker of faster inference will make up for that.

178 0 171 10/3 04:42 10/8 09:01 UTC
scorecomments23 sightings
first seen 2026-10-03 04:42 UTClast seen 2026-10-08 09:01 UTCscore then 174score now 171gained -3sightings 23
open on reddit ↗ 💬 55