New website to download models in case HF starts censoring or limiting access.
New website to download models in case HF starts censoring or limiting access.
Lately because of the current hardware shortage, unfortunately or fortunately, we can’t just throw infinite cloud compute at our problems, but we’re forced to actually care about what’s happening under the hood. We’re tweaking inference engines, learning quantization math, and optimizing architecture just to squeeze as much performance as possible for the lowest possible setups.
Fact: Just recently, the forked llama.cpp(s) and halogen-flash-server of Strix Halo pushed the performance through the roof, achieving double performance in decode (52tok/s), 5-6x performance in prefill (1300tok/s) for Qwen 3.8 Flash Next (Q38FN), and Q38FN itself is another massive architecture improvement with Engram, making it not only small but also smart.
I still remember before the hardware shortage, as someone who loves tweaking and optimizing, people just told me to stop, tweaking is stupid, just buy more RAM, buy more GPU..
It reminds me of the early web.. Back when setting up a box or hosting a server meant digging through forum threads, troubleshooting on IRC, and freely sharing custom scripts just to make things work. That era didn’t just produce programmers; it built hyper-versatile, end-to-end thinkers who understood the stack from bare metal up.
Contrast that with where mainstream web culture ended up. Most platforms today like Tiktok, Facebook, Youtube... are engineered for zero-friction doomscrolling.. Endless feeds of short-form videos designed to keep us distracted and waste our time. We’ve been overpampered by convenience.
My point: When we have too little, we try to learn more. When we have too much, we get distracted and learn too little. This is the golden time of our Local LLM community, let's learn and improve!
I finished my home inference server. First I tried Lenovo p620 workstation and while it’s a good value overall it pissed me off with a ton of proprietary Lenovo shit to deal with and I return it in the end.
Components:
4xV620 - 1400$
256GB DDR4 RDIMM 2666 - 610$
Huanandzhi D12D - 410$
EPYC 7452 - 170$
PSU ASRock 1600 - 220$
SSD Samsung 970EVO 1tb - Already had
Case//Fans//Misc \~ 200$
Power consumption is no shit ofc on such machine:
700-900w prefill
500-600w decode on Qwen3.8-next-flash Autoround W4A16
What it can do -
EDIT: Qwen3.8-next-flash Autoround W4A16 1.3k prefill and 70tg code/60tg prose on 128k+ context with MTP-2 on vllm fork.
I was disappointed with this machine and qwen3.8-27b speeds at first. But since Qwen3.8 next running good on it - I’m satisfied. Hope in more optimizations in future.
https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
AA shipped a new benchmark last week as part of the Intelligence Index v4.3 update — a brand-new private eval that replaces τ³. Astra was farming a ton of points on it and used those to get even with Fable, but… looks like we have a new king.
So they changed the index twice in three days to make Astra look not-quite-worse than Fable, and then a random guy quietly took first place on it.
Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with a bit of secret sauce (hint On-Policy Distillation) and achieved great results (-58% thinking tokens, 1.95x speed up, <1% accuracy loss) so we wanted to open-source it and hear the feedback of the community.
This is the link to the model: https://huggingface.co/ukisai/Swift-Qwen3.8-27b
We also also providing a Free Research Purpose API (OpenAI compatible), courtesy of Nvidia who were kind enough to provide us with the GPUs. You can use it to try out the model if you do not have enough compute to run it, it's limited at 5RPM. https://ukisai.com/api/swift/v1/models
We also made a GGUF (Q1-Q8) and there's also a few nice community (Bartowski) quants with even lower/higher precision. The community also created amazing NVFP4, W4A16 and Uncensored versions of the model you can find on Huggingface.
IMPORTANT: Our training approach is not a replacement for the reasoning effort settings, chat templates or token caps but is complementary and targets a completely separate issue (overthinking and "anxiety-like" reasoning loops prior seen in PTQ, but as far as we identified also prominent in BF16 of this size class LLMs as well). Contrary to popular belief, these specific patterns do not contribute to answer quality when properly targeted. (our thesis being: reasoning length IS extremely important and should NOT be shortened by force, but rather optimized). This is also demonstrated bellow in our xhigh vs medium effort benchmark table. The goal is to keep xhigh accuracy while reducing only the unnecessary part of thinking.
I will TLDR you on our thought process, research, training and benchmarks.
The benchmarks: (raw benchmark files here - https://github.com/UkisAI/Swift-Qwen3.8-27B-evals/ )**
Swift-27B vs Qwen3.8-27B (BF16, all benchmarks ran x5, thinking effort xhigh)
|Benchmark|Qwen3.8-27B|Swift-27B|Median tokens|
|:-|:-|:-|:-|
|GPQA-Diamond|88.4%|88.3%|58% fewer|
|LiveCodeBench v6|76.8%|81.6% (+4.8pp, due to default truncation in LCB it is not performance gain)|46% fewer thinking tokens|
|Terminal-Bench 2.1|66.7%|65.8%|39% fewer|
|MMLU-Pro|85.5%|85.0%|28% fewer|
|C-Eval|90.0%|90.6%|19% fewer|
|IFBench|73.5%|71.8%|51% fewer|
|AIME 2026|98.7%|94.0%|50% fewer|
|HMMT (Nov 2025)|99.3%|96.0%|46% fewer|
|ERQA (vision)|67.5%|66.3%|55% fewer|
Token savings hold at every reasoning effort (mean thinking reduction): xhigh 41%, medium 23%, low 26% (albeit with accuracy loses of 1-4% on medium and 1-2% on low which we need further testing for)
Swift at xhigh vs the base's own effort settings on GPQA-Diamond (198 questions x 5 seeds):
|Model / effort|Accuracy|Median tokens|
|:-|:-|:-|
|Base xhigh|88.4%|6,642|
|Swift xhigh|88.3%|2,771|
|Base medium|84.1%|1,753|
So Swift keeps xhigh accuracy at under half the tokens, and beats base-medium by 4pp at roughly 1.6x its tokens.
End note:
While we are keen on complete open-source, we still need to keep a part of our training and data private, being a new lab. The license is not Apache 2.0, but it only affects companies >$1M. We hope this does not pose a problem for the community, but we are open to feedback on it.
We want to contribute as much as possible to the community and would really appreciate feedback on our work, quantization or Swift model requests. For context, we are working on Swift 3.8 Flash Next right now and have so far gotten up to -30% thinking token usage while maintaining xhigh accuracy, which we take as a strong indicator our methodology is reproducible across the Qwen model family. Will explore other families as soon as we have the capacity and would love to see which ones the community would love for us to optimize first.
With all the recent drama surrounding AI safety. It’s obvious that open source could be caught in the crossfire.
From initial testing it seems pretty solid so far. Asked it to compile the latest llama.cpp for CUDA and its doing well so far. If this thing holds up to its score then its SHOCKINGLY good for its size.
Xi emphasised the need to step up cooperation.
https://www.cnbc.com/2026/09/13/china-xi-ai-tech-brics.html
https://www.yahoo.com/news/world/articles/xi-pushes-china-open-source-102711002.html
https://www.chosun.com/english/world-en/2026/09/14/C35V7T5MAVDN7LKEBESHBDBJ2Y/
https://it.euronews.com/next/2026/09/14/xi-jinping-propone-ai-brics-una-zona-di-ia-open-source
Note - This is translated from the actual blog link right at the bottom.
A few days ago, DeepSeek v4.1 was released. It raised the ability of small models to a new level.
AI is improving much faster than anyone expected. From the first ChatGPT that could only chat simply with a few thousand tokens of context, to models with real reasoning like OpenAI o1, DeepSeek R1, and Kimi K1.5 Thinking — that only took about two years. From reasoning models to agents that can smoothly use tools, run commands, and finish complex tasks — that took only about a year and a half. It’s hard to imagine what AI will be like in one, two, or three more years. How powerful will it be? Will it already be able to improve itself and deeply enter areas like embodied intelligence?
AI is getting better and better at writing operators
In the field I work in — designing and writing operators — AI has also improved very quickly. In just one year, it went from a small helper that could look up documents, read code, and find bugs, to an expert that can independently read CUDA, PTX, and SASS code, use professional tools to analyze the stall time of every instruction, and then optimize operators by itself. I believe that soon it will also be able to design operator schedules on its own, evaluate different schedules, implement them, and optimize them.
Of course I am proud of DeepSeek v4.1’s success — after all, its main Attention operator was written by me \[1\]. Its good performance is partly a recognition of my work. But the times keep moving forward, and technology cannot be stopped. I know clearly that in half a year or one year, the operators written by AI will most likely be as good as mine, or even better. AI can think 300 tokens in one second, type a command in half a second, and finish a piece of code in twenty seconds. I cannot. AI can keep improving in model depth, thinking strength, tool use (how often it interacts with the environment), and even parallelism. I cannot.
Humans have never hesitated when it comes to destroying themselves. Why do I still work hard to optimize operators, even though I know that the better my operators are, the faster our new models will train and run, the faster model ability will improve, and the sooner I will be replaced? One reason is that writing operators feels like playing a game to me. It gives me a lot of joy. When I invent a new technique or see the performance of my operator go up, I feel as excited as a speedrunner who breaks their own record. And when I see that my operator is much better than the official ones from the vendors, I feel very proud. But a more important reason is this: even if I give up or deliberately slow things down, other companies’ models will still keep improving and will replace me anyway. “Of course I hope I won’t be revolutionized. But if it has to happen, I hope the person who revolutionizes me is myself.” When everyone is so determined to destroy themselves, I have no choice but to join this cruel arms race.
What about me?
When the day comes that AI writes operators better than I do, what will happen to me?
My judgment is: I probably won’t lose my job completely, but I will have to change careers. I can still keep a job, but I may never again be able to do the work I once loved.
I once made a judgment about the changing times and my own future: because things are changing so fast (the AI progress above is a good example), I cannot predict what will happen in five or ten years. But no matter what, I believe that with my vision, judgment, initiative, and intelligence, I can stay in the game and stand at the front of the times again. However, this judgment only guarantees that I won’t become unemployed. It does not guarantee that I won’t need to change careers. In fact, it encourages me to change careers in order to avoid unemployment.
What does changing careers mean? It means I have to give up the field of operator design, writing, and optimization that I have worked in for a long time and loved deeply, and instead become a “mecha pilot” for Agents. Before, my interests, what I was good at, and what industry needed were basically aligned. Now, AI has made what I am good at into something it is even better at, and industry demand has shifted from “people who can write high-performance operators” to “people who can use AI to produce high-performance operators faster.” To meet industry needs, I will have to leave the direction I loved and move to an unknown new direction. I believe that with my understanding of engineering, upper-level model needs, and lower-level hardware, I can still produce operators with high quality and high efficiency. I also know I might come to love this new direction (or I might not). But the feeling of having my passion taken away is really not nice. That quiet joy of sitting at my desk and calmly writing operators for a whole afternoon may become a final song this summer. I have to bury my talent in yesterday and become a mecha pilot. My hands hold more gears, but my heart has fewer rhythms.
Here is a simple comparison: You are an expert at knitting sweaters. You are especially good at creating patterns and matching colors. The sweaters you make are high quality and beautiful, so rich people from near and far ask you to knit for them, and you make good money. At the same time, you really enjoy sitting by the window with a cup of tea, looking at the green mountains, water, cows, sheep, and cooking smoke, and quietly knitting for a whole afternoon. But one day someone invents a magical machine. You only need to give it yarn and a pattern, and it automatically knits a sweater. The quality and texture are as good as yours, and it is much faster. You know that your colleagues can easily reach your old level with this machine, so you have to use it too. You also know that with the knitting skills you built over twenty years, even when everyone has the machine, your speed and quality can still be better than others. But that feeling of listening to the rain by the window, slowly pulling the needle and thread, and enjoying the quiet time is crushed by the noise of the machine.
I know this is helpless, but there is no other way. I can keep my job, but my old passion will most likely have to be given up. I am a person whose rational side and emotional side are quite separate. When I need to be rational, I can be very rational, but sometimes I also show my emotional side. I remember when I moved out of the rental apartment I had lived in for a year, I cried a lot because I didn’t want to say goodbye to the memories. Saying goodbye today to the era of hand-writing operators and optimizing them with the human brain is even more cruel.
I don’t know if any readers feel the same way, but I think this is just how things are.
What about people?
While AI keeps improving, I also worry about some questions:
Will students now be much more likely to use AI to finish homework, especially practical labs? Imagine there are two choices: one is to spend eight hard hours finishing a lab and maybe not even get full marks; the other is to start an AI model, spend a few cents and a few minutes, and let AI write full-mark code. Which one will most students choose?
The point above will cause many students to have seriously weak engineering skills — things like organizing code, building systems, thinking about future needs and designing for them in advance, and abstraction ability. As AI keeps getting stronger, are these engineering skills still necessary? Will they be abandoned by the times like the old skill of “writing x86 assembly fluently,” or will they always be valuable like the ability to “understand the whole computer system from software to system to hardware”? If it is the latter, then it is dangerous — a person with poor engineering skills, when paired with AI, can produce messy code several times faster than before, planting all kinds of problems in systems and making the world more of a “clown stage.”
In future society, will power become more important than technology or intelligence?
These questions may need to be answered by the times themselves.
Conclusion
With the development of AI, future society may move toward two extremes: communism or Cyberpunk 2077. In the first, productivity is greatly liberated and people’s living standards improve a lot (I’ll stop here so I can pass review). In the second, a few tech companies control most resources. Only a very small number of people can use the most advanced AI and technologies and get close to “mechanical ascension.” Most people can only use very weak AI. Crossing social classes will become harder and harder: you need the strongest AI first in order to cross classes, which creates a dead loop.
Guess what: if Anthropic forever holds the most advanced AI in the world, will future society become communism or 2077? You guess?
So I still believe that the most advanced intelligence should be provided to everyone in an open and cheap way. I do not trust that Anthropic or OpenAI will do this. Especially, I do not want Anthropic to hold the most advanced artificial intelligence or AGI. To put it strongly, that would be as serious as letting Hitler get atomic bomb technology before the Allies. That is why I chose and continue to stay at DeepSeek: we research powerful, fast, and widely beneficial artificial intelligence and open-source it. Maybe this can pull the world a little bit back from the 2077 side.
May the future world be well. May all the beauty be blessed.
\[1\] “Main Attention” only includes the MQA attention with head dim = 512. It does not include the indexer used to select the top-k important tokens. That part was written by other (also very strong) colleagues (and their AI Agents).
Prices have been going crazy, but it's about to get worse me thinks.
Especially the 7B one seems very interesting, it casually destroys muse glimmer with a way smaller size. And they open source literally everything, every step of the way. Anyone tried that model? It can be a new milestone if 7b and 3.7b ones are actually good, and not just benchmaxed.
Now that the dust has settled a bit - here's a write-up on running Qwen3.8-Flash-Next (125B-A6B MoE + 51B n-gram table) on relatively middle-tier hardware (RTX 4070 12GB + 64GB DDR5-5600 + Gen4 NVMe on Linux).
I started out with bare 6 tok/s and through latest patches and optimizations getting close to 20 tok/s generation. You just need enough RAM.
For me this is the most intelligence possible on this machine right now. The 27B dense is not a choice because of low VRAM but may make more sense for other configs like 24GB VRAM owners. It actually surpasses the 27B model on most tasks as well so its great for Low VRAM, High/fast RAM configs.
PP is still a bit low at 300-350 tok/s.
What helped
\- Using AtomicChat's 4.27 bpw quant https://huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF
\- Ngram SSD offloading (lazy-mode)
\- --fit on --fit-target 512 helps automatically select the right params.
\- Master branch (19.35 t/s): Latest commit with MoE improvements.
- MTP Variant - PR #28243 + Compact MTP (20.65 t/s): MTP support is not yet merged so need to apply this PR enables Daniel Han's 1.78 GB \shared-Q4\_K\_M\ compact head. Combined with \-ncmoe 45\, it yields 77–96% acceptance and breaks through the 20 t/s barrier on every tested task (coding, summarization, creative).
With such low VRAM, MTP is not a huge jump because you have to give up a few layers to store the MTP head in VRAM. Only the shared + Q4\_K\_M in MTP gets a beneficial uptick.
Using commercial models to research, optimize and benchmark inference for local models helps a ton (GLM 5.3 flash with opencode go, so did Astra, Gemini 3.8 etc.)
Lot more details in the post (AI-assisted).
The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They forgot to include "Sponsored by OpenAI" at the end of the articles, sure that was just an oversight?
This is what the end of a bubble looks like, desperate attempts to get some sort of regulatory capture in place to keep the business model from collapsing in upon itself. My days next week are 100% booked talking to companies about how to get off frontier models, one large, and a bunch of smaller customers, including one who's flying me out to them to sit down and get a plan in place immediately (the controversy around that math problem really spooked some CEO/CIO's about data privacy using cloud models).
Gonna be an interesting few weeks. Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :)
20B-A1B model is coming, good for low VRAM people?
The first generation of our 150M model has just been released
Its performance is similar to that of GPT2-Small
The benchmarks:
PIQA: 62.24%
Hellaswag: 32.20%
Arc-Easy: 44.91%
Arc-Challenge: 25.00%
Arithmark 3.0: 33.90%
CapitalBench: 36.55%
It was trained on 7B tokens, using an RTX Pro 6000
an example inference script to try it out yourself is available in the Huggingface repo
If there's any question, I'll gladly answer them!
I'm like you guys and am constantly experimenting with new models, seeing what they're all good at, how I can make use of them for certain projects and goals. I've been using Qwen 3.8 27b for minor coding work and it has been impressive.
But with just regular chatting I have been impressed with Muse Glimmer.
It seems to be able to have the ability to follow and hold good, deep and meaningful conversations without coming off as a typical chatbot.
No repeated statements like "I hear what you're saying", "that sounds really deep..." none of what sounds generic or like it's blowing smoke up your ass. I was impressed with how natural it comes across just in natural conversation. I think it's one of the best "chat" models you could get right now as it's one of the only local models that doesn't feel like you're chatting with an AI when having a conversation.
I'm thinking of finding a way to run both Qwen3.8 and Muse at the same time. It's fun to play with these things.
I have a full model, it's ready to train. It's \~9b parameters.
9.4b to be more exact. That includes a 1/2/3 Engram table, Moonshot's AttnRes modeling, and RoPE / NoPE layering at 3:1 as more or less validated by most major labs. It uses the Llama 3 series tokenizer and LM Head as an initial start. The data fed in is logit level extraction from a Llama 3 teaching model.
I've already run the first training steps to test that the model is stable, etc.
I'm willing to sit and do the pre-IT training on the model. I don't know what task people really wanna do with this thing to be exact. So the focus of the IT training is a bit more vague other than giving it "Thinking" as per one of the open standards.
Honestly? I don't care what people want to use it for, just that they want to use it. I figured I'd just give it into the ether and LocalLlama was a place I figure I could easily give it to.
In theory the model should be more capable than any of the \~9b's we have running around with enough training. I'm also NOT a lab, so I don't have their training budgets. I basically ran all of the data production, etc on a 4090 + rented hardware.
The training code is deeply optimized to run on an RTX 6000 Pro series card. A single card.
All my engram research was being done before the Qwen model dropped. Qwen showed I only needed a single table injected at a layer, versus the 2 I used. 1 gave the majority of the benefit over 2.
The code would be open source, the data, all of it. IDGAF. It's technically already open source as it's all in public repos at the moment. I just stare at it and question if it's worth the time. At the least I'll throw a couple hundred at it to build a "functional" pre-IT model and build the IT dataset. It will need more pretraining before the IT set probably. It's kind of unknown because no one publishes the exact figures on this type of training.
If someone has a datacenter contact with a system they want to allow it to run on, we can all have it as a public model we watch. I tentatively named it "Budget" but... Localllama can name it whatever they want.
It's been fun to write the full end to end, generate data, etc. Even found issues in vLLM and reported to their repo that might end up helping you guys anyways. There was some prompt loading code that could be \~10x to \~100x faster I gave examples of to them.
If you read this far, thanks,
Signed some ML dude who reads too many research papers and has too much spare time.
edit; I went to review some Apache 2.0 licensing issues and noted a glaring hole in using Llama's tokenizer OR data. You can't even use synthetic data from their models without some licensing. I'll just have to rewrite the target to use OLMo 3 series I think. I guess I'll be back in a week after data generation and code fixes.
Open source uncensored Goon model or what? I'm trying to find a useful niche to develop the model so it gets used and ends up more than a research artifact. I'm planning on building the research artifact, it's a matter of whether I can find a group of users that will actually want to use it. It's all free, so I'm not trying to monetize you. The code and data lives on GH/HF.
It's still a dream for many folks to run medium size(30B range) models @ Q8 with Unquantized KVCache (256K Context) on their GPUs.
It would be awesome to have DeepSeek-V4.1-Flash's KVCache + Engram for all Upcoming models. Even for big models.
Engram - Heard that approximately 1/3-1/2 of Model size. Might come in different size range too. So 10-15 GB for 30B models.
Here few models with approximate numbers. Current models in Odd rows & Future/Fictional models in even rows(Bold). I just put 1GB for 256K context below though DeepSeek-V4.1-Flash takes only same 1GB for 1 million context.
|Model|Model Size|256K KVCache F16|MTP|Vision|Total GB|
|:-|:-|:-|:-|:-|:-|
|Qwen3.8-27B-Q8|29|16|1|1|47|
|Qwen4.0-27B-Q8|29|1|1|1|32|
|Qwen3.8-27B-Q4\_K\_M|17|16|1|1|35|
|Qwen4.0-27B-Q4\_K\_M|17|1|1|1|20|
|Muse-Glimmer-30B-Q8|30|16|1|1|48|
|Muse-Glimmer-2-30B-Q8|30|1|1|1|33|
|Gemma-4-31B|33|16|1|1|51|
|Gemma-5-31B|33|1|1|1|36|
|Qwen3.6-35B-A3B-Q4\_K\_M|23|6|1|1|31|
|Qwen4.0-35B-A3B-Q4\_K\_M|23|1|1|1|26|
|Gemma-4-26B-A4B-Q8|27|6|1|1|35|
|Gemma-5-26B-A4B-Q8|27|1|1|1|30|
Possibly there might be few more things(Please share those) to keep these number down. So I think 32GB VRAM is more than good enough for Upcoming (Optimized Smarter) Models. RAM is enough for Engram.
By above logic(based on Qwen4.0-27B), people could run Q4 of 54B models with same 32GB VRAM.
Maybe next year onwards, inventions could make 24GB enough for similar size models.
from internlm:
We introduce Intern-S2-397B, our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-S2-397B scales along three critical dimensions: pre-training, reinforcement-learning task coverage, and interactive agent environments. By combining a new vision-language pre-training paradigm with large-scale multi-task reinforcement learning and long-horizon agent reinforcement learning, Intern-S2-397B delivers a step change in general reasoning, scientific problem solving, and agentic capabilities.
This took roughly 5 hours to create, using 2 different configured harnesses, same model. RTX 3090, overclocked +12% gain (MSI Afterburner), Q4KM - built this for fun, will be throwing it on GitHub, opensource for people to get an idea of a project created to the near ceiling of performance & capability for q3.8 27b. & also maybe ya’ll can contribute to the game only iterating locally. It would be a fun little experiment.
\> "make an svg of a frog playing on a chello on the back of a whale with carribean island in the back."
interestingly the svg looks different in the OpenWebUi preview then when looked at in preview (osx). The palms and music notes are missing in the browser. I am pretty impressed by the result, is suggested to add some parameters to animate the whale and the water.
Qwen3.8-Flash-Next-IQ4\_XS on llama.cpp with 256K q8 context
openwebui reports:
input\_tokens: 27711
output\_tokens: 41562
total\_tokens: 69273
I know I know, it's great, we know. I've been working on tweaking inference engines for a week now and it's been one shotting most of my vague prompts without any issues. It will even write tests and validate the changes without me asking. It's actually nuts.
Last time I did something with advanced math I was making a game using Sonnet. It took many iterations to get physics to work correctly.
Such a good model. I'm so glad I went all in on local months ago. I was so tired of Claude making every excuse it could to try to force a new turn.
I had all the data saved from AA's v4.1 index, so when they upgraded it in the wake of Astra's release, I could actually generate a before/after comparison.
All models are the same. The only thing that changes is the weighted sum of the benchmarks that compose the Intelligence Index.
Highlights
Does anyone think the gurus on the DGX Spark forum are going to figure out how to magically fit DeepSeek 4.1 Flash on a 2x cluster, or is it only possible on 3 or 4 Sparks?
VLLM Benchmark:
Prefill, Prompt processing
\- avg, 871.93 tok/s (3 hours constant running xhigh)
\- 10K prompt, 1000.26 tok/s (16 runs)
\- 90K prompt, 743,59 tok/s (16 runs)
Decode, tok gen
\- avg, 38.39 tok/s (3 hours constant running xhigh)
\- 10K, 42.3 tok/s (16 runs)
\- 90K, 34 tok/s (16 runs)
Preamble: I am on WSL2. Running the 27B Q5 UD GGUF through llama.cpp with 81,920 context plus MTP gives me around 25-30 tok/s. Then I found this GitHub repo: https://github.com/noonghunna/club-3090
It is basically a recipe and Docker configuration for running the model.
So, 30 tok/s itself is fine, but I just got bored waiting for RunPod to open its GPUs. I finally brought my vLLM tuning back from the back burner, and here I am.
I often forget that Inductor/Triton compilation and CUDA Graph capture require additional VRAM while testing the configuration and kernel calls. When JIT compilation failed because of an OOM, I never bothered trying AOT.
FYI, AOT and JIT are compilation strategies. AOT means Ahead of Time, while JIT means Just in Time.
If you OOM on the first startup, try it one more time. Inductor might have already compiled and cached part of the configuration before the OOM, allowing the next run to reuse it if the configuration has not changed. This is not guaranteed, but it worked for me.
And yes, it was trial and error. It was kinda tedious and pain in the ass, starting from 32K, then 64K, 80K, 128K, and finally 144K. The practical ceiling for my conf at 154K, but I chose 144K. I also started the batch size at 256 and climbed to 1024, although I might be able to squeeze in 1280-1536.
Also, beware of your vLLM compilation cache. It might grow to 5-6GB after testing many configs. Personally, I delete the old cache and run the final configuration again twice so it rebuilds only what I currently use.
My current setup runs Qwen3.8-27B with INT4 AutoRound weights through vLLM while using an FP8 E4M3 KV cache. It fits on one GPU with a configured context window of 147,456 tokens.
Although I should say, with GDN, or really any linear-attn, vLLM can be kinda bad at predicting how much VRAM the KV and state-cache pools will require.
Benchmark:
https://github.com/noonghunna/benchlocal-cli
This is the deterministically scored, no-Docker portion of BenchLocal: 75 scenarios covering tool calling, instruction following, structured output, data extraction, and reasoning/math.
|Pack|Score|p50|
|:-|:-|:-|
|ToolCall|14/15 (93%)|3.21s|
|InstructFollow|15/15 (100%)|6.97s|
|StructOutput|14/15 (93%)|7.74s|
|DataExtract|14/15 (93%)|11.63s|
|ReasonMath|14/15 (93%)|9.50s|
|Total|71/75 (94.7%)|—|
Thinking was forced on with reasoning\_effort=low. The run took about 15 minutes. Yep, even with low reasoning effort and INT4 weights, it passed 71/75.
Setup
The important vLLM settings were:
Full command : https://gist.github.com/komikndr/b17955e1a80ce6ede9a3115f16216bc5#vllm-qwen-27b-38-just-remove-or-add-flag-as-you-like
Forgot to mention, no MTP and no MultiModal, i max the CTX, multimodal is at 64-65K ish but at that point i'll just use Llamacpp. And also again this is WSL2, if you are on baremetal, you could improve more speed
The full K2 Horizon lineup is out on Artificial Analysis.
The AA intelligence vs. parameters plots show that
\- 0.9B and 375B are bad
\- 3.7B and 7B are SOTA
\- 36B A4B is SOTA for hardware with poor memory bandwidth (spilled experts, Strix Halo, DGX Spark).
I'm going to take the AA Intelligence Index at face value here. This post is not about it.
The problem is that these models have a god-awful KV cache design. This means that you really can't use the number of parameters for "best in class" considerations, because these models heavily shift to the right on the plot if you replace parameter count on the X axis with RAM requirements.
For Q4\_K\_M weights, no drafter, no vision, 128k kvarn4 KV cache:
Compare them to
Notes: I don't advise compressing 2\~4B models to Q4 and I haven't tested these models' tolerance to weights and kv cache quantization yet. The above choices are just to keep the comparison fair.
This awful context design means that
Hello Reddit. Posting this for fun. I thought it was a lonely and silly journey to set up Qwen 3.8 Next on a system that doesn't really run it properly—it was a challenge that might help the community. I have yet to benchmark this specific REAP version versus Qwen 3.8 27B QK4, but my assumption is that it will do much better, despite the hemorrhaged world knowledge.
As you can see, we have about 44.5 GB of actually addressable system and VRAM available. The iGPU is taking care of the OS to make sure the GPU is totally free—but still, this is barely enough to hold everything together. This config actually totally fails with any of the Unsloth quants—no, I needed something more aggressive. I found the perfect thing—this REAP:
https://huggingface.co/AnonimousA/Qwen3.8-Flash-Next-REAP-320-GGUF
What's so great is the total size—a cool \~68.95 GB. The couple of Gigs we have shaved are absolutely key for making this all work.
Here is the breakdown of the model weights. We have the famous new n-gram portion, the experts, the active layers, the attention/SSM layers, and the extra space needed for the KV cache:
|Component|Weight (Approx)|Notes|
|:-|:-|:-|
|N-gram / PLE Embedding|\~29.48 GB|The massive lookup table|
|MoE Routed Experts (320)|\~34.89 GB|The main expert slab (pruned from 512)|
|Attention / SSM / Router|\~4.33 GB|Core architecture weights|
|KV Cache|\[TBD\]|Context memory overhead|
Obviously, running this model over SSD would make the speeds notoriously bad. Turning on mmap means that llama.cpp won't actually try to keep the model in RAM at all (it relies on the OS page cache instead), which results in \~2 tok/sec speeds—effectively useless.
The answer is to stick everything in RAM (using --load-mode none). The great thing is that the N-gram section of the model can be streamed over SSD via lazy mmap without this causing much issue—it's a massive lookup table that doesn't require heavy computation.
That's the huge win that allows an MoE model of this size to actually run well.
68.9 GB - 29.48 GB = 39.42 GB.
We just need to cram that 39.42 GB, along with the compute buffers and KV cache, into GPU and system memory, and we are golden—just barely. To do this, we need --lazy-mode on—that's what keeps the N-gram portion in RAM.
After that, it's a matter of fitting as many layers as possible onto the GPU. It's essential to completely fill the GPU as much as can be filled, so that we keep a precious few GBs in system RAM to run the OS. I found that having less than 2 GB left really started to destroy Fedora, but I think you could do better if you dropped the GUI—I just didn't want to in my case.
This leads me to --n-cpu-moe 34. This controls how many layers go to CPU. In my case, this was the exact limit needed to run this with 64k context on the GPU, quantized to Q4. Any more—GPU out of memory. Any less—total system meltdown, as the OS panicked and tried to put everything on the swap. You'll need to play around with this, but that was my exact number.
Settings used:
CUDA0 + --load-mode none --lazy-mode on
--n-cpu-moe 34
-c 65536 -b 512 -ub 128 -t 7 -ngl 48 -fit off -fa on
-ctk q4_0 -ctv q4_0 -kvo --cache-ram 0 --jinja --no-warmup
Results (64k Context, Q4):
I think this is in a somewhat usable state—but Qwen 3.8 27B GSQ IQ3S remains my daily driver; it's able to prompt process 5 times faster, I can fit in the mmproj and MTP layers, and it doesn't seem likely to set my desk on fire. But maybe for really hard tasks, I'll use the next model. It is smarter, it runs at a reasonable speed, and it was a good learning experience.
I'm curious if anyone else is able to get this model or just large MoEs working on a GPU and RAM config similar to mine. LMK. Also, I'm a total noob to this stuff, any advice is appreciated.
(Also, heading off all the obnoxious "why did you quantize the cache - unusable - just get a better computer" ragebait posts. This is a human being writing this post, to help others and just enjoy pushing something to its limits. And in my limited testing, the next model seems much better at pixel art than the 27B version.)
Final Note: If you have a larger pool of system memory, like 64 GB (because you can spend $899 on Amazon on a kit of DDR5 somehow), you would be better served by using this fork of llama.cpp, which has optimized flags for this exact setup and wonderful guides. For me in particular, with my limited hardware, this seemed to work better—their cache kept OOMing unless I turned on mmap—but I think with more system RAM, their setup and guides are optimal.
I am considering dumping my ChatGPT Plus subscription and go full local, but to do so I would first need to reach a decent result for quality (and reasonable speed).
My 4090 fried itself out of nowhere, so I am not stuck with a 3070 until I get something better.
I am kind of suspicious about the claims the companies are doing lately about the dangers of AI and how they are pumping the prices intentionally to either law out the open source or price out the open source, so I want to just go full local even more now.
I have a MSI MAG X670E Tomahawk WiFi which theoretically supports 3 GPUs?
What would you end up with?
p.s. I am ruling out Macs to be open and easier to setup in case I will want to use them for my homelab
edit: typo
Pushed out this new update, hopefully decreases the instances of crashes. I torture tested this one for \~12 hours after my fixes and found no instability. Q4 K XL needs more fine tuning, which I will work on in the future. I am simultaneously juggling this + a legitimate inference engine + finalizing work on a deep research/site builder application I've been working on for about 6 months. After those get pushed to prod I will refocus here. Thanks!
Plug: join the Launch80 discord if you are in to the cutting edge of RDNA4 optimization! There are guys pushing out even better numbers and configurations than mine here on other quants. I think we are starting to get closer to the ceiling on these configurations where the model isnt fully VRAM resident.
I gave DS V4.1 Flash an HLE problem with a bash tool + 2 hours.
Hour 1: it wrote three MILP solvers. (225,200)
Hour 2: it downloaded the HLE dataset from Hugging Face, found the question, read the answer key (225,600), and concluded its own answer (225,200) was better.
I'm equal parts impressed & terrified.
I'll start by saying I'm not talking about the models themselves, I'm aware that I can't come close to something like Fable's intelligence locally. Just wanted to get that out of the way. Basically. Over the last year I've gotten quite comfortable with claude code, and it seems likely there were be a gradual cost rug pull, and I'd like to put myself in a better position when that happens for local use. I am already used to running local models (such as Qwen3.8_Q5) in things like lmstudio, but I have no experience with other harnesses. I'd like to know, what harness, right out of the box would feel most at home for current Claude Code users. I say this as someone who was not coding prior to "vibe coding". I'm looking for the path of least resistance, though I will no doubt eventually spread out into tools that give me more control. But for now, I'm just looking for a life raft. Just needs to be local, opensource, and free of spyware. In case someone wants to know 24GB VRAM (rtx 3090) and 64GB DDR4.
Edit: Spelled Baseten not Base-10
Full episode of this available at https://www.youtube.com/watch?v=PrSf7IOYu-I
It's interesting to see how Dwarkesh has had to come around to the evidence that we are well on our way to creating AGI and even RSI in the last few months, despite historically being very skeptical.
I highly recommend people interested in large language models check out this particular episode, because it dispels a lot of mythology about stuff like plateaus from lack of data etc. For those who thought we were hitting a wall a year ago, it turns out there was a ton of low hanging fruit and the researchers in this episode discuss what that fruit was. They also extrapolate these trends into the future.
It's funny this subreddit is becoming rather skeptical of AI progress, which to put diplomatically, I think is based on a lack of information and too much time on Reddit.