Qwen3.8 Flash Next on 5060 Ti 16GB - 55 tok/s average, and a few demos
Hello guys! I've been out of the loop for a while. Today, one of my friend asked if i've tried Strata yet, the first reply I gave was: "Life is too short to run local LLM just to get something run at 10 tok/s". Hehe, I was an idiot.
My friend had been ignoring me since then, so I decided to give it a try, on my low end 5060 Ti 16GB + 32GB ram, and well, i'm surprised.
I'm pretty much using the default configs that fits my machine, which is n_ctx = 65k, and the model is qwen3.8-flash-next-coder-iq1_m. This is how the speed looks like:
https://preview.redd.it/d6wbv5swqvth1.png?width=1172&format=png&auto=…
On average, prompt processing is at 1k5 tok/s, and gen speed is at 55 tok/s.
Now, before you laugh at IQ1\_M, I decided to see how bad is the generation result, so I tried with a one shot prompt to create a simple landing page:
https://preview.redd.it/yo6kj7servth1.png?width=1834&format=png&auto=…
The total run time was about 2 minutes, at 46 tok/s. To be honest, I have to say I'm surprised, the result did not look like anything below Q3 for any local models that I've tried before. Here's a closer look at it:
https://preview.redd.it/z67sg0hlrvth1.png?width=1905&format=png&auto=…
There are some minor issues, but I have to say it's even better than the claudish style that I usually get with other frontier models. Maybe that kind of problem was well trained, so I decided to try another prompt, make an interactive 3d globe:
https://preview.redd.it/sccdghnzuvth1.png?width=1870&format=png&auto=…
This time, it ran for 8 minutes for the first version, and took about another minute to fix the JS errors. The result came out still impressive.
https://preview.redd.it/lzik4fezvvth1.png?width=3436&format=png&auto=…
You can see the two demos yourself here: