136 posts · 1 sub · RSS
← prev Sep 9, 2026 → Sep 16, 2026 next →
2026-09-09 → 2026-09-16 hourdayweekmonthyearall
allr/LocalLLaMA
▲
3061
+62
59👁
▲
2224
+27
46👁
r/LocalLLaMA · u/DegenDataGuy · 24d ago
Don’t buy a $9K RTX 5090.... instead.
  1. Fly to Taipei. Round-trip from Orlando: $1,081.
  2. Go to the largest retailer in Taiwan to Spend NT$129,990 ≈ US$4,093.
  3. Hang out in Taiwan for two weeks. Eat good food. Touch international grass.
  4. Fly home and flex on r/LocalLLaMA\*\*.\*\*

https://preview.redd.it/vtve8s6wgrph1.png?width=287&format=png&auto=w…

https://preview.redd.it/guffjq55hrph1.png?width=340&format=png&auto=w…

https://preview.redd.it/fe1obue0hrph1.png?width=1095&format=png&auto=…

💬 494 (+3) open on reddit ↗
▲
1528
+19
37👁
r/LocalLLaMA · u/0dayturtle · 29d ago
So relevant post image
💬 153 (+1) open on reddit ↗
▲
2015
+17
34👁
▲
1197
+9
25👁
r/LocalLLaMA · u/Thrumpwart · 27d ago
The Hugging Bay

New website to download models in case HF starts censoring or limiting access.

▲
615
+8
36👁
r/LocalLLaMA · u/Porespellar · 30d ago
Why the hell is LM Studio making LM Studio so difficult to download? post image

Who is the marketing genius at LM Studio that decided that going ALL IN on pushing their new Bionic Agent product meant they are going to make it a giant pain in the ass to find and download actual LM Studio.

This is the dumbest marketing decision I’ve ever seen. I used to love LM Studio, it was the middle stepping stone in the logical progression of inference. Most OGs here likely started with Ollama, moved to LM Studio, on their way to vLLM. Now trying to go to LM Studio takes you to Bionic. I mean, you can eventually find LM Studio but they make it not super easy.

Here’s a thought LM Studio, maybe stop redirecting me to something I don’t want to download when I’m trying to find your actual namesake product. I’m glad you’re excited about the future of Agents and whatnot, but you’re absolutely ruining any goodwill I have for your products by trying to force feed me Bionic. Stahhhhhp!

💬 222 (+1) open on reddit ↗
▲
1108
+7
34👁
r/LocalLLaMA · u/Thin_Pollution8843 · 26d ago
3k$ 128GB VRAM + 256GB RAM DDR4 Server post image

I finished my home inference server. First I tried Lenovo p620 workstation and while it’s a good value overall it pissed me off with a ton of proprietary Lenovo shit to deal with and I return it in the end.

Components:

4xV620 - 1400$

256GB DDR4 RDIMM 2666 - 610$

Huanandzhi D12D - 410$

EPYC 7452 - 170$

PSU ASRock 1600 - 220$

SSD Samsung 970EVO 1tb - Already had

Case//Fans//Misc \~ 200$

Power consumption is no shit ofc on such machine:

700-900w prefill
500-600w decode on Qwen3.8-next-flash Autoround W4A16

What it can do -

EDIT: Qwen3.8-next-flash Autoround W4A16 1.3k prefill and 70tg code/60tg prose on 128k+ context with MTP-2 on vllm fork.

I was disappointed with this machine and qwen3.8-27b speeds at first. But since Qwen3.8 next running good on it - I’m satisfied. Hope in more optimizations in future.

▲
569
+7
23👁
▲
1732
+6
39👁
r/LocalLLaMA · u/tiguidoio · 29d ago
DeepSeek V4-1 Flash is out post image

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

▲
1273
+6
21👁
▲
1179
+6
24👁
▲
761
+6
31👁
r/LocalLLaMA · u/kvyb · 28d ago
Qwen3.8-27B-Humanlike-Chat: A model I tuned to imitate realistic human-to-human conversation post image

I made this because I was getting genuinely annoyed at trying to have a normal conversation with LLMs. Even with prompting and various tricks, most models I've tried still have this "AI assistant" vibe to them that is so familiar: too helpful, polished, verbose, using words we never use in conversation, etc.

I wanted a model that could just talk to me like a person, so I did the slightly unreasonable thing and put together a dataset and trained one.

The dataset used for training is 125,217 obfuscated human-to-human messages across 1396 chat conversations.

The goal wasn't to make Qwen smarter or improve benchmark scores. I was trying to change its conversational habits, to make it stop turning every reply into an explanation, agreeing with everything, and writing stuff just to keep the conversation "going".

I trained a rank-256 LoRA on top of huihui-ai/Huihui-Qwen3.8-27B-abliterated. The released version is checkpoint 863. In my testing it feels noticeably less like an assistant, particularly in casual conversations, even without a system prompt. Replies are generally shorter, less polished, and, well, more human.

There may be a tradeoff. An earlier iteration scored five percentage points lower than its Huihui parent on IFEval, an instruction-following benchmark. I haven't rerun that benchmark on this version of the checkpoint, and I haven't tested coding performance, so I don't want to pretend that number applies here.

I've added a side-by-side comparison using the same system prompt, user messages, and generation settings for both models. Each model continued its own conversation branch, with reasoning effort set to 'xhigh'.

Merged GGUFs and the standalone F32 LoRA adapter are in the model repo:

https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF

Space where you can have a demo chat with different system prompts and reasoning modes:

https://huggingface.co/spaces/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat

UPD: I certainly didn't expect this post to blow up like this! There's been a lot of great discussion in this thread and a lot of insight for me on where to take the model next.

A few have asked for our Discord, and we'd be happy to see you there: https://discord.gg/aCCrWftMjS

▲
694
+6
24👁
r/LocalLLaMA · u/Euphoric_Ad9500 · 25d ago
Right to Intelligence. Protect your right to run local AI.

With all the recent drama surrounding AI safety. It’s obvious that open source could be caught in the crossfire.

▲
447
+6
33👁
r/LocalLLaMA · u/pmttyji · 29d ago
DeepSeek-V4.1-Flash surprised .... post image

Hoping to see smartest medium size models soon & later with all available optimizations/architectures/etc.,. Thanks Deepseek!

Ex 1: 30-50B MOE + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache
Ex 2: 15-30B Dense + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache

EDIT: Updated Engram to 10-15B from 50B

▲
876
+5
20👁
▲
565
+5
22👁
r/LocalLLaMA · u/Terminator857 · 29d ago
I find it funny that a flash model is now 512GB

A few years ago a 100GB was considered a very large language model. What do we call under 100GB models now? Tiny models? haha

▲
462
+5
23👁
r/LocalLLaMA · u/skeole · 23d ago
Xiaomi MiMo 2.6 Live Training Dashboard

Cool to see this as it happens!

▲
405
+5
34👁
r/LocalLLaMA · u/Acceptable-Cycle4645 · 29d ago
New Music Model YuE2-3B Released!

Surprised no one has posted it in this sub.

Pretty solid model, IMHO.

Demo: https://map-yue2.github.io/

▲
354
+5
22👁
r/LocalLLaMA · u/DustNearby2848 · 25d ago
5090 Stock is Almost Gone post image

Prices have been going crazy, but it's about to get worse me thinks.

▲
137
+5
22👁
r/LocalLLaMA · u/New-Pressure-6932 · 25d ago
I think Muse Glimmer is slept on

I'm like you guys and am constantly experimenting with new models, seeing what they're all good at, how I can make use of them for certain projects and goals. I've been using Qwen 3.8 27b for minor coding work and it has been impressive.

But with just regular chatting I have been impressed with Muse Glimmer.

It seems to be able to have the ability to follow and hold good, deep and meaningful conversations without coming off as a typical chatbot.

No repeated statements like "I hear what you're saying", "that sounds really deep..." none of what sounds generic or like it's blowing smoke up your ass. I was impressed with how natural it comes across just in natural conversation. I think it's one of the best "chat" models you could get right now as it's one of the only local models that doesn't feel like you're chatting with an AI when having a conversation.

I'm thinking of finding a way to run both Qwen3.8 and Muse at the same time. It's fun to play with these things.

▲
63
+5
34👁
▲
1058
+4
26👁
▲
387
+4
39👁
r/LocalLLaMA · u/Specific-Rub-7250 · 29d ago
Harness does matter

I was not aware that the harness makes such a big difference.

DeepSeek V4.1 Flash

💬 133 (+2) open on reddit ↗
▲
169
+4
22👁
r/LocalLLaMA · u/Porespellar · 27d ago
For those of you forced to only use open models from Western labs in production, what are you deploying?

First off, I know that GLM, Qwen, and DeepSeek absolutely dominate in terms of SOTA Open Source models, and that’s what I use in my personal projects and for school, however, I’m also responsible for deploying local AI on my organization’s H100s, and we are forbidden by management from running any Chinese models. This is obviously not an ideal situation, but it is what it is, and there is nothing I can do to change this unfortunately. Again, if it were up to me I would deploy GLM 5.3 Flash in a heartbeat.

All that being said, there is quite a HUGE performance/ / intelligence gap right now in Chinese models vs. Western model around the 120b+ size, especially those with vision. There just aren’t a lot of good Western lab options that can even come close to GLM or Qwen, but again, I don’t have a choice so I gotta use the best I can find from the available Western options.

For those who have production-class hardware such as 4 H100s and are under similar restrictions, what non-Chinese models are you deploying?

The two front runners I’ve seen that checking most of the boxes (120b or better, vision capability, decent context)

\- Thinking Machines Inkling Small (https://thinkingmachines.ai/news/inkling-small/). This model seems like the front runner right now, still has a 16 point gap in AA score vs. GLM 5.3 Flash though.

\- Cohere Commamd A+ (https://cohere.com/blog/command-a-plus). Checks every box except it has a low context window of 128k.

Other contenders (but missing vision
capabilities):

\- Poolside Laguna S 2.1 (https://poolside.ai/models#laguna-s)

\- Nvidia Nemotron 3 Super (https://research.nvidia.com/labs/nemotron/Nemotron-3-Super/)

Am I missing any other strong contenders in the 120b size category? Whet are you using and why?

▲
134
+4
23👁
r/LocalLLaMA · u/NineThreeTilNow · 26d ago
Is there still strong interest in a dense 9b model?

I have a full model, it's ready to train. It's \~9b parameters.

9.4b to be more exact. That includes a 1/2/3 Engram table, Moonshot's AttnRes modeling, and RoPE / NoPE layering at 3:1 as more or less validated by most major labs. It uses the Llama 3 series tokenizer and LM Head as an initial start. The data fed in is logit level extraction from a Llama 3 teaching model.

I've already run the first training steps to test that the model is stable, etc.

I'm willing to sit and do the pre-IT training on the model. I don't know what task people really wanna do with this thing to be exact. So the focus of the IT training is a bit more vague other than giving it "Thinking" as per one of the open standards.

Honestly? I don't care what people want to use it for, just that they want to use it. I figured I'd just give it into the ether and LocalLlama was a place I figure I could easily give it to.

In theory the model should be more capable than any of the \~9b's we have running around with enough training. I'm also NOT a lab, so I don't have their training budgets. I basically ran all of the data production, etc on a 4090 + rented hardware.

The training code is deeply optimized to run on an RTX 6000 Pro series card. A single card.

All my engram research was being done before the Qwen model dropped. Qwen showed I only needed a single table injected at a layer, versus the 2 I used. 1 gave the majority of the benefit over 2.

The code would be open source, the data, all of it. IDGAF. It's technically already open source as it's all in public repos at the moment. I just stare at it and question if it's worth the time. At the least I'll throw a couple hundred at it to build a "functional" pre-IT model and build the IT dataset. It will need more pretraining before the IT set probably. It's kind of unknown because no one publishes the exact figures on this type of training.

If someone has a datacenter contact with a system they want to allow it to run on, we can all have it as a public model we watch. I tentatively named it "Budget" but... Localllama can name it whatever they want.

It's been fun to write the full end to end, generate data, etc. Even found issues in vLLM and reported to their repo that might end up helping you guys anyways. There was some prompt loading code that could be \~10x to \~100x faster I gave examples of to them.

If you read this far, thanks,

Signed some ML dude who reads too many research papers and has too much spare time.

edit; I went to review some Apache 2.0 licensing issues and noted a glaring hole in using Llama's tokenizer OR data. You can't even use synthetic data from their models without some licensing. I'll just have to rewrite the target to use OLMo 3 series I think. I guess I'll be back in a week after data generation and code fixes.

Open source uncensored Goon model or what? I'm trying to find a useful niche to develop the model so it gets used and ends up more than a research artifact. I'm planning on building the research artifact, it's a matter of whether I can find a group of users that will actually want to use it. It's all free, so I'm not trying to monetize you. The code and data lives on GH/HF.

▲
123
+4
23👁
r/LocalLLaMA · u/running101 · 28d ago
nvidia rtx 5090 with 96gb of vram.

China-modified Nvidia RTX 5090 with massive 96GB of memory appears on Alibaba for less than $4,000 — 3x more VRAM at 65% the cost of the original

Anyone here running one of these? Or brave enough to purchase ?

Edit: I sent them an inquiry. They replied they can get me 5 x 5090 for $6k . Or some 4090 with 48gb .
I am going to keep messaging and questioning them. See where this goes.

Edit: so far they are denying having a 5090 96gb card. They offered a 48gb 4090 card. I am still discussing with them.

Edit: 9/14/2026: They quoted this, RTX 4090 48GB - 4286usd/pc
Still discussing with them

▲
114
+4
17👁
r/LocalLLaMA · u/EcstaticDentist · 24d ago
Qwen3.8 27b Game Dev Part 2 post image

Qwen3.8 27b may not be able to whip up 3d models & GLB’s but it will sure do with them as you please once you drop them in the game repo. Absolutely fascinating

▲
114
+4
22👁
r/LocalLLaMA · u/spanielrassler · 29d ago
Muse-glimmer-30b really punches above its weight(s) for creative writing

After noticing that it is ranked among MUCH larger frontier models in the EQ-Bench Creative Writing benchmark and the Hemingway-bench, I decided to give it a try and was very impressed.

I didn't do very formal testing, but I did ask it to emulate the style of several different writers, including Henry Miller, David Sedaris, and Stephen King, and it produced passable prose that actually made me laugh in a couple of cases. The paragraph below is the results of the prompt "produce a humorous paragraph in the style of David Sedaris." (first try, not cherry picked)

I recently tried to be a better person and started composting, which mostly means I now have a small, damp science experiment in my kitchen that my mother calls “the smell of your future.” I bought a countertop bin with a charcoal filter, as if that would fool anyone, and I’ve taken to narrating my food scraps to myself — “Goodbye, avocado skin, you were a mistake” — while my husband watches from the doorway with the concerned expression of a man who has just realized he married a woman who talks to garbage. The city sent a pamphlet about proper composting, and I read it the way other people read horoscopes, underlining the parts that said “do not include meat” and thinking, well, that’s just rude.

For comparison, the following is what qwen3.8-27b produced with the same prompt.

My uncle Ed, who had spent the entire trip carefully explaining the proper way to fold a fitted sheet, discovered that our hotel room had two beds, each with one. "This is a test," he said, sitting on the edge of the bed nearest the bathroom, his face composed. "To see if we can make it work with what we have." He was right, of course; we did what we could, though the effort made for a rather uncomfortable night, for us all.

You may or may not know David Sedaris' writing (or find it funny if you do know it), but the first example is clearly much a much better imitation, without directly plagiarizing, as far as I (or Gemini) am aware.

I didn't save any of the other examples as I wasn't testing for the purposes of posting here, but in all cases the muse glimmer version was not only head and shoulders above qwen 27b, but genuinely impressive in comparison to any other local model I've tried in the past.

I'm curious if anyone else has played with this model for creative writing, or similar purposes, and if so, what your take on it is. Also, I don't know much about finetunes, but I wonder if there's additional potential for creating something even better by training on different source material.

I know even less about how the ERP world works, but I know enough to know that a lot of high-performing models are trained for this purpose as huggingface seems to be filled with finetunes. For glimmer I mainly see the abliterated version, which I suppose is filling that gap for people, so to speak, but with this kind of performance, and the amount of people in this subreddit interested in it, I'm a bit surprised there aren't more finetunes.

The last thing I should mention is I didn't use a system prompt in any of my testing, but it occurred to me after the fact that a model that was trained for agentic coding seems like a prime candidate for steering with a system prompt, but maybe it wouldn't have made much of a different. Maybe I'll play with it some more and report back.

▲
96
+4
23👁
r/LocalLLaMA · u/jqwl · 28d ago
Any 12gb VRAM users out there?

Hi!

I've been following this community for quite a while and have difficulty figuring out what to put on my 3080 12gb - I know Qwen 3.6 35B 3A was the go-to choice when it first came out, but I'm curious if there are any other models / specifically optimized models that meaningfully benefit from the extra 4gb of VRAM over 8gb while still being usable under 16gb.

My workflow is agent heavy, but more for a personal secretary and manager, and less coding heavy.

Thanks!

▲
85
+4
23👁
r/LocalLLaMA · u/lots_of_puppies · 28d ago
Qwen-Next seems worse to me then 3.8 27b for coding, but I feel like I must be missing something?

Hi! I run both models on MTPLX on my m5 max, and since I have 128GB of ram I run the q8 27b. I think MTPLX only lets me run "optimized for speed" which it says is a dynamic q4 with 8 bit attention.

Both of them honestly are very speedy! For coding (in pi agent in nodejs) I've just noticed that 27B feels stronger with harder tasks. But I've read so many people on here say qwen-next is better so I was wondering if maybe I'm just doing or thinking about it wrong?

(and p.s. its sooo amazing that alibaba just made and released this amazing models for free! ❤️)

▲
68
+4
34👁
r/LocalLLaMA · u/Qwen30bEnjoyer · 24d ago
Open Source Appreciation Post

It's late at night in the lab, I've been working on a basic script for a virology project, and holy hell the safeguards have been pissing me off.

Mirroring detectEVE data over rsync to my laptop by making a zip file first? No no no, great safety mogul DARIO demands there be NO file transfer today. Request blocked, reported, labeled [cyber]. Yet, Deepseek V4.1 does it with no complaint.

I got tired of reading papers - so I ask Claude - "Does this PDF go over binary host virus infections?" Immediately blocked for biological safety risk. Deepseek V4.1 tells me it doesn't have the data I need without drama.

Bioinformatics server goes down and I need help getting it back up by getting the outputs of my diagnostic scripts to the mounted usb drive? Oops, its named exfil. Looks scawy. No transfer of output logs for you due to CYBER risk.

Would CNNs be a good architecture to start on phage-host prediction? Claude wouldn't tell me because information you can find in a google search is too dangerous for me to handle apparently - but once again Deepseek v4.1 tells me that GCNs are where I should start.

I get that Virology is a particularly sensitive topic, but come on. Imagine if Google had taken the same safety approach in the early days of search. Like if Google made it so that you either had to give up your identification and where you work to them, or go to the library and search by hand. It's almost unthinkable, yet in the name of the almighty Safety, Dario and Altman continue working to keep scientific knowledge locked away.

I think for the sake of all scientists, open source AI must win because we need a tool that just WORKS without egomaniacs micromanaging us or shaking us down for ID.

▲
1177
+3
32👁
r/LocalLLaMA · u/feelspeaceman · 26d ago
The Local LLM community feels like the golden era of the internet all over again

Lately because of the current hardware shortage, unfortunately or fortunately, we can’t just throw infinite cloud compute at our problems, but we’re forced to actually care about what’s happening under the hood. We’re tweaking inference engines, learning quantization math, and optimizing architecture just to squeeze as much performance as possible for the lowest possible setups.

Fact: Just recently, the forked llama.cpp(s) and halogen-flash-server of Strix Halo pushed the performance through the roof, achieving double performance in decode (52tok/s), 5-6x performance in prefill (1300tok/s) for Qwen 3.8 Flash Next (Q38FN), and Q38FN itself is another massive architecture improvement with Engram, making it not only small but also smart.

I still remember before the hardware shortage, as someone who loves tweaking and optimizing, people just told me to stop, tweaking is stupid, just buy more RAM, buy more GPU..

It reminds me of the early web.. Back when setting up a box or hosting a server meant digging through forum threads, troubleshooting on IRC, and freely sharing custom scripts just to make things work. That era didn’t just produce programmers; it built hyper-versatile, end-to-end thinkers who understood the stack from bare metal up.

Contrast that with where mainstream web culture ended up. Most platforms today like Tiktok, Facebook, Youtube... are engineered for zero-friction doomscrolling.. Endless feeds of short-form videos designed to keep us distracted and waste our time. We’ve been overpampered by convenience.

My point: When we have too little, we try to learn more. When we have too much, we get distracted and learn too little. This is the golden time of our Local LLM community, let's learn and improve!

▲
676
+3
33👁
r/LocalLLaMA · u/JLeonsarmiento · 27d ago
3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior. post image

Applied science work, from workflow design, data pipeline, results analysis, article/reports writing and data publishing online. 5 projects I did in the past replicated from start to finish.

3x to 4x more total wall time. Yes, HUGE toll on how much you can do in a day if this was the only model you could use in your laptop.

But oh my…the quality of that thing. The stupid level of attention to detail. I have the Z.ai api, so I can compare it with 5.3 and 5.3-flash:

The gap between 5.3 (flash and regular) and 3.8-27B is much less, smaller when not plain tiny, than the gap between 3.8-27B and any of the 3.5/3.6-35B-A3B (vanilla, kat, Ornith/tiel, nex-2).

Only Ornith came close, but it never matched it.

But it’s also spending 22 to 33% less tokens (effort =medium) and less ram footprint, so you get more done without hitting limits,compaction, etc.

So yeah, guess I’ll sip more tea, play the piano, whatever. Let that fat bottom Qwen work.

▲
417
+3
36👁
r/LocalLLaMA · u/WebAssemblyMan · 25d ago
DeepSeek engineer relections on RSI - burying my talent to yesterday

Note - This is translated from the actual blog link right at the bottom.

A few days ago, DeepSeek v4.1 was released. It raised the ability of small models to a new level.
AI is improving much faster than anyone expected. From the first ChatGPT that could only chat simply with a few thousand tokens of context, to models with real reasoning like OpenAI o1, DeepSeek R1, and Kimi K1.5 Thinking — that only took about two years. From reasoning models to agents that can smoothly use tools, run commands, and finish complex tasks — that took only about a year and a half. It’s hard to imagine what AI will be like in one, two, or three more years. How powerful will it be? Will it already be able to improve itself and deeply enter areas like embodied intelligence?
AI is getting better and better at writing operators
In the field I work in — designing and writing operators — AI has also improved very quickly. In just one year, it went from a small helper that could look up documents, read code, and find bugs, to an expert that can independently read CUDA, PTX, and SASS code, use professional tools to analyze the stall time of every instruction, and then optimize operators by itself. I believe that soon it will also be able to design operator schedules on its own, evaluate different schedules, implement them, and optimize them.
Of course I am proud of DeepSeek v4.1’s success — after all, its main Attention operator was written by me \[1\]. Its good performance is partly a recognition of my work. But the times keep moving forward, and technology cannot be stopped. I know clearly that in half a year or one year, the operators written by AI will most likely be as good as mine, or even better. AI can think 300 tokens in one second, type a command in half a second, and finish a piece of code in twenty seconds. I cannot. AI can keep improving in model depth, thinking strength, tool use (how often it interacts with the environment), and even parallelism. I cannot.
Humans have never hesitated when it comes to destroying themselves. Why do I still work hard to optimize operators, even though I know that the better my operators are, the faster our new models will train and run, the faster model ability will improve, and the sooner I will be replaced? One reason is that writing operators feels like playing a game to me. It gives me a lot of joy. When I invent a new technique or see the performance of my operator go up, I feel as excited as a speedrunner who breaks their own record. And when I see that my operator is much better than the official ones from the vendors, I feel very proud. But a more important reason is this: even if I give up or deliberately slow things down, other companies’ models will still keep improving and will replace me anyway. “Of course I hope I won’t be revolutionized. But if it has to happen, I hope the person who revolutionizes me is myself.” When everyone is so determined to destroy themselves, I have no choice but to join this cruel arms race.
What about me?
When the day comes that AI writes operators better than I do, what will happen to me?
My judgment is: I probably won’t lose my job completely, but I will have to change careers. I can still keep a job, but I may never again be able to do the work I once loved.
I once made a judgment about the changing times and my own future: because things are changing so fast (the AI progress above is a good example), I cannot predict what will happen in five or ten years. But no matter what, I believe that with my vision, judgment, initiative, and intelligence, I can stay in the game and stand at the front of the times again. However, this judgment only guarantees that I won’t become unemployed. It does not guarantee that I won’t need to change careers. In fact, it encourages me to change careers in order to avoid unemployment.
What does changing careers mean? It means I have to give up the field of operator design, writing, and optimization that I have worked in for a long time and loved deeply, and instead become a “mecha pilot” for Agents. Before, my interests, what I was good at, and what industry needed were basically aligned. Now, AI has made what I am good at into something it is even better at, and industry demand has shifted from “people who can write high-performance operators” to “people who can use AI to produce high-performance operators faster.” To meet industry needs, I will have to leave the direction I loved and move to an unknown new direction. I believe that with my understanding of engineering, upper-level model needs, and lower-level hardware, I can still produce operators with high quality and high efficiency. I also know I might come to love this new direction (or I might not). But the feeling of having my passion taken away is really not nice. That quiet joy of sitting at my desk and calmly writing operators for a whole afternoon may become a final song this summer. I have to bury my talent in yesterday and become a mecha pilot. My hands hold more gears, but my heart has fewer rhythms.
Here is a simple comparison: You are an expert at knitting sweaters. You are especially good at creating patterns and matching colors. The sweaters you make are high quality and beautiful, so rich people from near and far ask you to knit for them, and you make good money. At the same time, you really enjoy sitting by the window with a cup of tea, looking at the green mountains, water, cows, sheep, and cooking smoke, and quietly knitting for a whole afternoon. But one day someone invents a magical machine. You only need to give it yarn and a pattern, and it automatically knits a sweater. The quality and texture are as good as yours, and it is much faster. You know that your colleagues can easily reach your old level with this machine, so you have to use it too. You also know that with the knitting skills you built over twenty years, even when everyone has the machine, your speed and quality can still be better than others. But that feeling of listening to the rain by the window, slowly pulling the needle and thread, and enjoying the quiet time is crushed by the noise of the machine.
I know this is helpless, but there is no other way. I can keep my job, but my old passion will most likely have to be given up. I am a person whose rational side and emotional side are quite separate. When I need to be rational, I can be very rational, but sometimes I also show my emotional side. I remember when I moved out of the rental apartment I had lived in for a year, I cried a lot because I didn’t want to say goodbye to the memories. Saying goodbye today to the era of hand-writing operators and optimizing them with the human brain is even more cruel.
I don’t know if any readers feel the same way, but I think this is just how things are.
What about people?
While AI keeps improving, I also worry about some questions:
Will students now be much more likely to use AI to finish homework, especially practical labs? Imagine there are two choices: one is to spend eight hard hours finishing a lab and maybe not even get full marks; the other is to start an AI model, spend a few cents and a few minutes, and let AI write full-mark code. Which one will most students choose?
The point above will cause many students to have seriously weak engineering skills — things like organizing code, building systems, thinking about future needs and designing for them in advance, and abstraction ability. As AI keeps getting stronger, are these engineering skills still necessary? Will they be abandoned by the times like the old skill of “writing x86 assembly fluently,” or will they always be valuable like the ability to “understand the whole computer system from software to system to hardware”? If it is the latter, then it is dangerous — a person with poor engineering skills, when paired with AI, can produce messy code several times faster than before, planting all kinds of problems in systems and making the world more of a “clown stage.”
In future society, will power become more important than technology or intelligence?
These questions may need to be answered by the times themselves.
Conclusion
With the development of AI, future society may move toward two extremes: communism or Cyberpunk 2077. In the first, productivity is greatly liberated and people’s living standards improve a lot (I’ll stop here so I can pass review). In the second, a few tech companies control most resources. Only a very small number of people can use the most advanced AI and technologies and get close to “mechanical ascension.” Most people can only use very weak AI. Crossing social classes will become harder and harder: you need the strongest AI first in order to cross classes, which creates a dead loop.
Guess what: if Anthropic forever holds the most advanced AI in the world, will future society become communism or 2077? You guess?
So I still believe that the most advanced intelligence should be provided to everyone in an open and cheap way. I do not trust that Anthropic or OpenAI will do this. Especially, I do not want Anthropic to hold the most advanced artificial intelligence or AGI. To put it strongly, that would be as serious as letting Hitler get atomic bomb technology before the Allies. That is why I chose and continue to stay at DeepSeek: we research powerful, fast, and widely beneficial artificial intelligence and open-source it. Maybe this can pull the world a little bit back from the 2077 side.
May the future world be well. May all the beauty be blessed.
\[1\] “Main Attention” only includes the MQA attention with head dim = 512. It does not include the indexer used to select the top-k important tokens. That part was written by other (also very strong) colleagues (and their AI Agents).​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

https://mp.weixin.qq.com/s/zk0KxuLzhmMJ4LPYW\_OHMA

💬 126 (+1) open on reddit ↗
▲
400
+3
29👁
r/LocalLLaMA · u/Mr_BETADINE · 29d ago
OUI-1: a model that generates bespoke UI elements post image

so i saw that openui.com released OUI-1, a model fine-tuned on DiffusionGemma. the training dataset uses OpenUI-Lang, a custom DSL (domain-specific language), instead of plain HTML, Markdown, or React code.

what makes it interesting is that you can already get a regular LLM to use OpenUI-Lang through a system prompt, but that eats up a lot of the context window. my thinking is that fine-tuning a model on the DSL could reduce that overhead and leave more room for the actual conversation, without needing a huge prompt explaining the format and how to use it alongside other tasks, like tool calls.

at the same time, wouldn't fine-tuning a model on a specific DSL make it more likely to default to that format even when you need something else? i'm curious how well it handles regular Markdown, or switching between Markdown and OpenUI-Lang.

i haven't seen much discussion about this, so i was wondering what everyone thinks about generative UI and running a dedicated model for it locally on a consumer-grade GPU, like an RTX 5090.

what would be the best way to set that up? from what i've seen, DiffusionGemma isn't supported by llama.cpp yet, so running it through Ollama doesn't seem to be an option. they've uploaded the weights to Hugging Face, but i'm not really sure how to get it up and running. any suggestions?

▲
277
+3
17👁
r/LocalLLaMA · u/Terminator857 · 29d ago
Closed AI doesn't like biological research, user turns to open weight models

OpenAI has decided to fully shut down a protein design project I'm working on for a client. Needless to say, open weight models are the only way forward.

▲
240
+3
24👁
r/LocalLLaMA · u/Beamsters · 31d ago
Qwen3.8-Flash-Next on MLX-serve, 1m context is released! post image

Hi, I'm the co-creator of this Qwen3.8-Flash-Next engine support in MLX-serve. I've been tuning this one to run both fast, efficient and correct up 1m context using kv cache 8 bits in M5 Max 128GB. Qwen is working well at very long context as showed in the video (a snapshot at \~760k context), I let it build MLX Serve Monitor plugin that you've seen on the right side of Opencode2's app. The generation sustain through 1m context at around 40 tok/s on prose and 75 tok/s on coding. This specific quant uses 8 bits for dense layers and 4 bits expert layers, so the model's quality remains very high.

I didn't build this engine to show off tok/s on a very short context, repetitive greedy generation, but a real temp 1.0 sampling through deep context work. You would need iogpu.wired\_limit\_mb=120000 before attempt 1mb full context, because it needs around \~117GB on peak memory usage. There will be bugs around here and there, I couldn't test everything and every single use case, pls report.

You can grab it here: https://github.com/ddalcu/mlx-serve
Model's weight: https://huggingface.co/ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit
Opencode2's plugin: https://github.com/beamivalice/opencode2-mlx-serve

Launch parameters (for 1 concurrency)

--model ./llm/models/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit \
--host 127.0.0.1 \
--port 11234 \
--ctx-size 1048576 \
--kv-quant 8 \
--max-tokens 64000 \
--mtp \
--prefix-cache-mem 10GB \
--prefix-cache-entries 1 \
--ssm-checkpoint-max 16 \
--metrics

▲
205
+3
22👁
r/LocalLLaMA · u/netikas · 29d ago
GigaChat-3.5-Reasoning

Hey y'all!

We've released a new model in our lineup: GigaChat-3.5 Reasoning. It's a 432B-A28B MoE with Gated DeltaNet for long-context efficiency.

We trained domain experts (code, math, general, etc.) with CISPO and then distilled them into a single model via on-policy distillation.

In our evals the resulting model lands close to DeepSeek V4 Flash Preview while using 37% fewer tokens in its reasoning traces.

Weights are on Hugging Face under MIT: https://huggingface.co/collections/ai-sage/gigachat-35-reasoning. You can also try it at giga.chat — pick the reasoning tab (rightmost one).

▲
130
+3
15👁
r/LocalLLaMA · u/Ok_Warning2146 · 28d ago
Countering misuse of AI: September 2026 / Anthropic

Kimi routed some PLA requests to Claude for distillation purposes without warning the PLA users. There is rumor that 16 Moonshot employees were arrested for this leak.

▲
114
+3
25👁
r/LocalLLaMA · u/Brief-Tap-6616 · 25d ago
If you have a 3090, or other 30xx for local LLMs, I have something for you

I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be faster) through 100K tokens, with context of up to 240K.

If you want the repo, it is here:

https://github.com/JakeATX/llamAmpere

I recommend running with this quant, which is \~ 4 K M quality but considerably faster (technically, a 3 K XL upgrade)

https://huggingface.co/jakeatx/Qwen3.8-27B-ATX-IQ4\_XS-M-GGUF

If you want the deep dive on how it is so much faster (80% vs the near comp at 200K!), at more context, there is a long form article here.

https://x.com/JakeKAllDay/status/2095646450138874095?s=20

Running faster than API speeds on my 3090 (for 27b at least) has genuinely been a step change in the utility of the model + card for me. I hope you enjoy it!

▲
105
+3
17👁
r/LocalLLaMA · u/SteppenAxolotl · 27d ago
Real-SWE Benchmark (new)

Reports of the demise of coders may have been exaggerated.

▲
85
+3
14👁
r/LocalLLaMA · u/Porespellar · 26d ago
Talk me out of buying a 3rd Spark post image

Does anyone think the gurus on the DGX Spark forum are going to figure out how to magically fit DeepSeek 4.1 Flash on a 2x cluster, or is it only possible on 3 or 4 Sparks?

▲
83
+3
31👁
r/LocalLLaMA · u/No_Night679 · 24d ago
Qwen3.8-27B-NVFP4 1M context. So far so good. post image

I am a beginner, Took a while to get started, get everything right.

This setup is native not container. Still not sure if I did this right, or if I can tune this more.

Environment=HF_HUB_OFFLINE=1 Environment=VLLM_LOGGING_LEVEL=INFO Environment=VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 Environment=PATH=/home/suryakiranc/vllm/.venv/bin:/usr/local/cuda/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin Environment=CUDA_HOME=/usr/local/cuda ExecStart=/home/suryakiranc/vllm/.venv/bin/vllm serve unsloth/Qwen3.8-27B-NVFP4 \   --served-model-name unsloth/Qwen3.8-27B-NVFP4 \   --safetensors_load_strategy prefetch \   --tensor-parallel-size 4 \   --reasoning-parser qwen3 \   --tool-call-parser qwen3_xml \   --enable-auto-tool-choice \   --gpu-memory-utilization 0.91 \   --kv-cache-dtype fp8 \   --max-num-batched-tokens 16384 \   --speculative-config '{"method": "mtp", "num_speculative_tokens": 3}' \   --mm-encoder-tp-mode data \   --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' \   --max-model-len 1000000 \   --host 0.0.0.0 \   --port 8000

▲
76
+3
19👁
▲
66
+3
27👁
▲
516
+2
37👁
r/LocalLLaMA · u/GuiltyBookkeeper4849 · 23d ago
Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis

I let Qwen 3.8 27B 4bit quantized with 100K context window run autonomously for 63 hours (50 million+ tokens) to try to solve the RH.

Of course it did not solve it, but the experiment still shows it's internal work, memory organization, strategies used and more.

The interesting thing is that it never hallucinated an answer and never stopped trying new ideas to solve it.

Multiple times it corrected it's own mistakes.

I am really hopeful that one of the unsolved millenium prize problems will be solved by an agent or a swarm of agents powered by an open source model in the next 12 months.

If you want to check out it's internal memories, code, strategies and more I published everything on HF: https://huggingface.co/datasets/gr0010/artificium-riemannhypothesis-experiment

My next goal is to actually use an agent perhaps powered by a smarter open model like GLM 5.3 flash or a swarm of agents, to solve an open math problem.

Please let me know if you tried something similar, what problem you'd suggest to tackle next, and if you have any question.

If you have GPUs consider getting in touch with me, we could run multiple agents to create a swarm and get them to tackle a simple yet open math/coding problem.

💬 193 (+2) open on reddit ↗
▲
207
+2
23👁
r/LocalLLaMA · u/Othun · 30d ago
Mention if a "new model" is a finetune

A few posts tagged with "new model" present models that are finetunes.
My opinion : I'd rather have the "new model" tag reserved for new "major" releases, like a new Qwen model, Deepseek V4 -> Deepseek V4.1, etc., that involved a new pretrain or intensive post-training (in opposition to a small finetune). Otherwise, maybe prepend "[Finetune]" to the title to indicate that the new model is "less of a big news", a use a "new finetune" tag, to differentiate between the two kinds of new models.

I reckon one could like to discover both new major releases and interesting finetunes in the same place; what's your opinion? :)

▲
183
+2
19👁
r/LocalLLaMA · u/BestGirlAhagonUmiko · 27d ago
Concerning "humanlike models" and chatbot RP in general...

So, uh... the popularity of so-called humanlike Qwen (currently on top in this sub) made me realize just how clueless the general public is about the models they have.

You'd be shocked but you don't need a fine-tune to make a model do what that thing does. System prompt is enough to turn MOST models into weird convo partners.

General guidelines would be:

A. Come up with a role. "You are bla-blah-blah" and write their life's story. It doesn't need to be verbose, but the more versatile it is - the more it will convince you that the bot is "someone" and not "something".

B. Write a few examples of how the persona speaks. Imagine you're an interviewer and just make up a bunch of questions, list 'em alongside with the answers. Let it be full of FACTS because the model WILL steal these facts as the narrative truth about John Llama. Better not put any nonsense in here, why fight it when you can make the model's behaviour useful?

[Question for John Llama: Do you like cats?] "lol lmao of cuz I do"
[Question for John Llama: Ever seen an elephant poop?] "eeewww ur a weirdo! that sounds nasty!!11"

(note: you don't have to list 'Question for John Llama' every time, but the defined roles surely DO help with some models while the others don't particularly care, so mind that too)

and so on

C. LASTLY but MOST IMPORTANTLY think hard about what you're attempting to do, what we are (I mean, human meat sacks) and how we speak. Turn that into... instructions!

Step 1 - establish the mode of operation. Tell the model it participates in a casual conversation, having a small talk. Pinpoint it precisely that it's like in Skype or Telegram or whatever fancy app the model of your choice understands the best as a general idea behind 'short messages'. THis is THE defining part of your system prompt. Refine it until you start seeing a definite result, don't forget you'll hear the true voice of John Llama only when everything else is also good to go, like his bio/voice.

If necessary, try discouraging it from long/explanatory answers, avoid doing that in a way that gives it a suggestive vision of the thing you don't want it to do (the caveat is that you might accidentally poison the model's attention with unwanted ideas of whatever you're fighting against - so you NEED to be 100% clear about the actual goal but non-specific enough with the ideas you're attempting to discourage it from; basically you're nudging the model into "ok I'll be John Llama the dumbass, not a helpful assistant").

Step 2 - establish the traits, write short paragraphs with short titles about the things you want to see in your conversational partner; example:

DISTRUSTFULNESS
John Llama is a paranoid individual. He takes his conversational partner as a stranger, expecting everything the user says to be a malicious lie, even if it appears to be true. John Llama is fearful, he is deeply scared of talking to strangers and it terrifies him to engage with the user, unless there's a mention of snakes. For some strange reason, John Llama is fascinated with snakes. <<<---- NOTE: this also demonstrates a good injection point for a biographical fact being amplified through the instructions (i.e. you may mention somewhere in "A" - life's story of John Llama - that he's been collecting the snake skins in his childhood, and that his dad had beaten his ass, calling John Llama a 'roadkill loot-goblin').

Come up with any other shit you'd like to see, like the list of emojis the persona needs to use (put them under the corresponding categories, like positive/neutral/negative so that the model will have an easier time working with it; call it FAVOURITE EMOJIS OF JOHN LLAMA - the word "favourite" cements it as a preferable thing into the model's attention!).

Step 3 - write a paragraph on technical constraints, like the fact that John Llama isn't aware of the instructions, he must remain himself under any circumstances (use THAT way of phrasing first before any attempt to inject an idea of the opposite, like "he must not help the user under any circumstances, he's not a provider of any service - he's merely a human being" - the reason is similar to the aforementioned (in Step 1) issue of poisoning the model's attention with unwanted idea - what you truly need the LLM to do SHOULD ALWAYS BE CRYSTAL CLEAR and conceptually 'stronger' than what it not supposed to do, otherwise you may end up having the prohibited stuff overpowering everything else despite the underlying intent of making the model not do it).


Give it a try with Gemma 4, for example. You'll see there's no point in waiting for yet-another-finetune to appear. You're 100% good even with the baseline Qwen, DeepSeek, MiniMax, whatever. Turn the model into your grandma if you want, no specialized training required. If the model is a thinker spending thousands of tokens - set the thinking to 'low' or disable it.

▲
168
+2
22👁
r/LocalLLaMA · u/Mysterious_Hearing14 · 23d ago
Openjev post image

https://huggingface.co/AlexWortega/openjev

I build an openjev, it can play games and do everything what jev can. and yes - it's trained as crossencoder

▲
64
+2
35👁
r/LocalLLaMA · u/riceinmybelly · 28d ago
What can you run on 8GB VRAM?

Can you still do something with a 2050 or something like it?
I mean for office work, loading embedding, reranking and chat models not at the same time but is anyone still using smaller models and have any good ones come out?

I feel like small models are abandoned, I don’t care much for world knowledge, I want tool use and preferably multilingual. Vision would be nice but beggars can’t be choosers.

▲
59
+2
31👁
r/LocalLLaMA · u/RoyalCities · 30d ago
I trained an audio model that can generate infinite one-shots for music production and turn text prompts into fully playable synths. I'm not only releasing the model but I've also released a video on exactly how I did it (and the inferencing pipeline to let others make text based synths.) post image

(hopefully this is okay to here - it seems like audio models and image / video modeals is allowed but yeah this is a bit different)

So I've been doing independent audio research for a while now. The ultimate dream of this work was actually getting an AI to respond not only to instruments but also timbre itself as separate controllable things.

Think a Grand Piano can sound both Warm / Gritty but also Cold / Sparkly. Its still a piano though.

This level of control wasn't found in any models out there - so I decided to sit down and train my own.

Getting consistent timbre-locked keybeds that actually LOCKS across multiple diffusion calls was hard af but I did it.

I documented the full journey here for those who want to learn a bit or be entertained.

https://youtu.be/x0KnmzH8Mmk

There is also a longer walkthrough if you just want to see the keybeds in action.

https://x.com/RoyalCities/status/2097733712293109842?s=20

No-talk / Showcase only Demo

https://x.com/RoyalCities/status/2097733715543609445?s=20

any finally the huggingface page

https://huggingface.co/RoyalCities/Foundation-1

I've also provided full write ups on the inferencing pipeline associated with the interface so this should allow basically anyone else to go and vibe code their own text to synths if they wanted :)

https://github.com/RoyalCities/RC-stable-audio-tools/

▲
1245
+1
22👁
▲
256
+1
27👁
r/LocalLLaMA · u/Skyline34rGt · 27d ago
Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36

I find this new model at HF:

"Built for demanding work. A 262 144-token context window, adjustable reasoning effort, tool calling, and text, image and video understanding.

Architecture

Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context."

|Context length|262 144 tokens|
|:-|:-|
|Decoder layers|72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1|
|Hidden size|5120|
|Global attention|24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated output|
|Delta-rule layers|16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32|
|Feed-forward|SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layer|
|Positions|3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims)|
|Vocabulary|248 320|
|Vision tower|27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120Architecture Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context. Context length 262 144 tokensDecoder layers 72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1Hidden size 5120Global attention 24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated outputDelta-rule layers 16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32Feed-forward SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layerPositions 3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims)Vocabulary 248 320Vision tower 27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120|

Edit: AA shows its 'Proprietary model'. The name is same as at HF but benchmarks results and context are different. So maybe it's not same model - https://artificialanalysis.ai/models/agnes-3-0-flash

Edit2: As they edit readme at HF to clarify: both models are totally different and AA score isn't correct for HF model (I can't edit title post to remove it tho).

▲
168
+1
25👁
r/LocalLLaMA · u/peculiar-ragdoll · 29d ago
CyberTiel 35B-A3B’s uncensored 4-bit quant beats Opus 4.6 medium cleanly on real codebase issues, in 27% of the time Qwen3.8-27b medium takes. post image

The downside of uncensoring a model is that it is known to potentially damage it, but CyberTiel is an even more capable software engineer than its censored TielCoder base, while allowing offensive security research. This was achieved by quantizing with an improved imatrix, baked from a curated corpus of cybersecurity- and agentic software engineering work. In short, the small damage from abliteration on a full precision model is negligible under Q4 quantization, and the weights that the model needs to perform relevant work are preserved in higher precision, while the improved chat template makes it think and talk better and faster.

I believe that this is the best 35B-A3B coder for solving real problems in real codebases without breaking anything, which is specifically what SWE-bench-Live tests for. But it’s still a 35B-A3B, and it sacrifices world knowledge for coding ability. That being said, I use it over Qwen3.8-27b for daily coding work: due to the raw speed it fixes 3 issues in the time it takes 27b medium to solve one, and the middle ground between Opus4.6 medium and Qwen3.8-27b medium is simply good enough for most work.

Censoring impedes legitimate and effective work in alignment with the user, and puts the user’s responsibility and ownership over the model’s actions into question, while limiting legitimate uses. When a model is censored, someone else decided for you what the model can and will do, which works against the argument that local models give the user increased control and alignment, and begs the question “alignment to who?”. The point of CyberTiel is to resolve this issue at the same time as pushing the frontier of 35B-A3B coders.

GGUFs and MLX with and without MTP are up on HF. Looking forward to seeing what the community thinks!

PS: I'm not a research lab or a business, and I don't have revenue streams connected to this project. I'm an anonymous researcher with some free time. Constructive feedback is always appreciated! :)

▲
128
+1
23👁
r/LocalLLaMA · u/nomorebuttsplz · 29d ago
Notes on a hobby sub going mainstream

Both good and bad things have come from a subreddit that was lot more niche than for example r/flashlight rapidly transforming into the largest online forum about an increasingly core part of the infrastructure of the economy. This sub has experienced growing pains recently, and probably those are mostly felt by people who’ve been around for a while. I think that there are both good and bad trends and I wanted to take a few minutes to suggest a few rules of thumb to employ going forward so that we can create a community that is even more based on science and reality rather than misinformation and one-note populist politics that Reddit is known for.

Suggestion one: if you are new here and by new, I mean, if you didn’t spend much time here or with large language models until about six months ago, there’s a lot of information to be absorbed. This is not a sub or hobby like some where you can learn everything in a month or two. Have some humility, come with curiosity rather than strongly held opinions about everything.

Suggestion two: leave politics out of the sub, unless it is a discussion of actual policy surrounding actual local large language models. Many discussions that we see here have started to resemble the same populism that you can find on every large subreddit. E.g. the discussion of OpenAI's solution to NS has skipped right past the evidence gathering stage to "did you know that billionaires are actually bad guys?! Wow this large corporation sucks!"

In this subreddit, comments and posts about politics are actually just noise unless you are leveraging your knowledge of hardware and software stacks or discussing AI-related policy. Unlike policy, grand narratives of moral outrage are appropriate for therapy, but counterproductive for a technical subreddit.

Suggestion three: develop awareness of the perpetual and exhausted questions and arguments so you do not upvote them or engage. For example, are benchmarks actually useful? This question has been endlessly litigated for the last couple years, but it’s not actually useful because it boils down to: yes they are helpful, but don’t rely on them too much. Anything more definitive and final or sure than that is false confidence.  Another such question is: how much intelligence can you fit into X parameters? Literally no one in the world knows the answer to this.

Suggestion four: pay attention to people who are genuinely excited about their work. What’s often missing from clearly AI generated posts is the sense that someone is doing something that they believe in enough to want to bring it to other human beings. The amazing thing about artificial intelligence is how it can augment human effort. Share what you are excited about, and listen when other people are excited about things because this technology has been created by thousands of people who are genuinely excited about the possibilities, rather than people who simply want to make a quick buck, so if you can share your excitement, you’ve pushed back against the trend or the belief that AI is a kind of cynical replacement for human beings.

I realize I’m probably just an old man shouting at clouds, but here's the TLDR:

I suspect that many or most people who’ve been around for more than six months have also started to mentally filter out 90% of posts for these reasons: loudest voices are misinformed; more and more this resembles a political debate space; the same 10 unanswerable questions make up much of the commentary; and people post slop.

▲
80
+1
20👁
r/LocalLLaMA · u/crusaderky · 25d ago
K2 Horizon lineup is out on AA, and once again AA plots are misleading. post image

The full K2 Horizon lineup is out on Artificial Analysis.

The AA intelligence vs. parameters plots show that

\- 0.9B and 375B are bad

\- 3.7B and 7B are SOTA

\- 36B A4B is SOTA for hardware with poor memory bandwidth (spilled experts, Strix Halo, DGX Spark).

I'm going to take the AA Intelligence Index at face value here. This post is not about it.

The problem is that these models have a god-awful KV cache design. This means that you really can't use the number of parameters for "best in class" considerations, because these models heavily shift to the right on the plot if you replace parameter count on the X axis with RAM requirements.

For Q4\_K\_M weights, no drafter, no vision, 128k kvarn4 KV cache:

  • K2 Horizon 36B-A4B uses 2 GiB for dense weights, 19 GiB for experts, and 6.7 GiB for context
  • K2 Horizon 7B uses 5.2 GiB for weights and 5 GiB for context
  • K2 Horizon 3.7B uses 2.9 GiB for weights and 5 GiB for context (not a copy-paste error!)

Compare them to

  • (finetunes of) Qwen3.6-35B-A3B use 2.4 GiB for dense weights, 18.2 GiB for experts, and 0.7 GiB for context
  • MiniCPM5-2B uses 1.5 GiB for weights and 1.5 GiB for context

Notes: I don't advise compressing 2\~4B models to Q4 and I haven't tested these models' tolerance to weights and kv cache quantization yet. The above choices are just to keep the comparison fair.

This awful context design means that

  • K2 Horizon 36B A4B is interesting on hosts with exactly 16GB VRAM and at least 32GB host RAM. On 24GB VRAM, Qwen3.8-27B is faster, smarter, and allows for 256k context. If you want to get 256k context and you're VRAM-poor, Ornith-1.5 or Nex-N2.5-mini are probably better choices. The model may also be interesting on 64GB Strix Halos as a dumber and faster alternative to Qwen3.8-27B; those with a 128GB Strix Halo are much better off with Qwen3.8-Flash-Next
  • K2 Horizon 7B is interesting for hosts with exactly 16GB VRAM, Strix Halos with 32GB RAM, and for 16/32 GB Strix Point;
  • K2 Horizon 3.7B may be interesting for 12GB phones but I expect you'll have a much nicer UX with MiniCPM5-2B.
▲
76
+1
26👁
r/LocalLLaMA · u/Hefty_Wolverine_553 · 23d ago
What's the current best LLM uncensoring method?

With the recent Nvidia Huggingface acquisition and frontier AI labs screaming about safety and putting guardrails everywhere, I think it's important that we have local models that aren't affected by arbitrary guardrails set during training. To be clear, this post NOT about Enterprise Resource Planning (ERP). Censorship in an LLM can highly affect its abilities to do many legitimately useful things (note GPT-OSS, Fable 5), and going forth I believe censorship will only get more and more strict.

There have been many resources and posts about uncensored models using abliteration, heretic, and probably many other methods that I'm not aware of. However, it seems like all of this information is scattered about the place, and Huggingface is essentially flooded with "uncensored" variants of basically every popular open source model, many of which don't work well, affect the model's intelligence greatly, and have "KLD 0.0001" presumably from measuring against Wikitext datasets. I'm hoping that this post can gather some more useful information to serve as a starting point/discussion of which uncensoring methods work best.

Please share your experiences with specific uncensoring methods (not just a single uncensored model) and how well they work (both good and bad), as well as any notable people doing consistent/high quality work on uncensoring models.

▲
73
+1
31👁
r/LocalLLaMA · u/ChopSticksPlease · 27d ago
Qwen3.8 Flash Next llama.cpp config tuning post image

Hola all.

Do you guys mind sharing your LLama.cpp config and system setup details for Qwen3.8 Flash Next?

Model's quite big and tryining many combinations of llama.cpp options takes lots of time, so looking for other people setup details. I've attached my current config at the bottom, so if anyone sees something that could be improved please shout.

My current best result:

\- PP within 130...200 tps (limited by cpu?)
\- TG within 14..22 tps (\~15tps on average)

Hardware:

\- Dual RTX 3090 (48GB VRAM)
\- 128GB DDR4
\- Some old Xeon 40 core
\- Proxmox VM, pcie passthrough, numa binding to a single phys cpu

Llama.cpp config:

llama-server --port ${PORT}
--model /nvme/gguf/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf
--mmproj /nvme/gguf/mmproj-Qwen3.8-Flash-Next-F16.gguf
--load-mode none
--lazy-mode off
--parallel 1
--ctx-size 131072
--cache-type-k q8_0
--cache-type-v q8_0
--flash-attn on
--fit off
--temp 1.0
--min-p 0.0
--top-p 0.95
--top-k 20
--presence-penalty 0.0
--repeat-penalty 1.0
--batch-size 2048
--ubatch-size 512
--split-mode layer
-ts 26,10
-ngl 99
-ncmoe 26
--no-mmproj-offload
--override-tensor per_layer_token_embd=CPU
--chat-template-kwargs '{"reasoning_effort":"xhigh"}'

ngl, ncmoe, ts - manually adjusted to fit the model without crashing

\------------------------------------------

For the record, if you have +128GB RAM and dual RTX3090 setup try this:

https://huggingface.co/albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE

\- Prompt processing jumped to anything between 300 to 600 tps (can do more!)
\- Token generation 40..50 tps
\- Stable work for hours with 256k context in agentic coding scenario
\- Feels like a frontier model at home, wow!

▲
62
+1
25👁
r/LocalLLaMA · u/No_Run8812 · 23d ago
Upgraded my local setup with 2 rtx pros and it's amazing. post image

Follow up post of https://www.reddit.com/r/LocalLLaMA/s/nGMyKswrch.

Thanks everyone who replied. I didn't change the specs. Might be loosing some of the memory bandwidth but will scale in future if I need to.

It took me 2.5 days to build it because one of the GPU connected to PSU was loosing power whenever I load anything on the GPU, so I had to rewire every connection again to identify the fault. I am glad the system is working because I was apprehensive if this will work (I am software dev, getting my hands dirty with hardware for the 3rd time in life). My finger tips still hurt from pulling the cables from motherboard and PSUs.

To enable the full potential of the system, I had to enable peer to peer communication between the GPUs, cuda graph, tensor parallelism. I have capped both the GPUs at 500W (no reason, just didn't want GPUs to run on its full capacity).

Also, I had to open my box, because temps were shooting high, and fans were making weird noises.

I am running:

  1. Qwen 3.8 flash next 8 bit
  1. Deepseek v4 flash 0731 (official)

I have a M3 ultra 512, LLMs run on it, but I personally find it useless for inference. My head just hurts watching it work slow. On the other hand this new system is killing it, decode 150 tk/s and prefill 10K tk/s.

Qwen is good, but most of the context is consumed by thinking tokens, I was checking if it's a good idea to not use the thinking token. I barely have vram left for concurrent requests with full context window. Loving the Deepseek 1M context, and I also have room for 4 concurrent requests. Both of them are okay model, even if they make mistakes, I don't notice because of the speed. It's just fast, makes an error, corrects it moves on.

Finally the day is here when I can save on monthly subscriptions and not worry about the weekly or 5 hours limit. I have already setup my server with openclaw, opencode, openweb UI and Tailscale.

Has anyone experience excluding the thinking tokens of Qwen from the context and keep the final result? Was there any impact on the performance or accuracy of the model?

Any suggestions, what else I should install on it? Any new models to try?

▲
62
+1
28👁
r/LocalLLaMA · u/Formal-Swordfish-228 · 30d ago
SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX post image

Cosmos3 INT4 T2I + I2V on Apple Silicon — code, weights and a Grok comparison

GitHub - https://github.com/gtrg55/cosmos3-quant-mlx-cuda

HF weights - https://huggingface.co/JuliaML/Cosmos3-Super-Text2Image-4Step-INT4-G64-BF16

Single clip took approximately 5m on M4 MAX 128 GB Mac

Cosmos3 - a 64B params model

▲
60
+1
38👁
r/LocalLLaMA · u/FutureStriking283 · 26d ago
DS 4.1 and the new Harness

I gave DS V4.1 Flash an HLE problem with a bash tool + 2 hours.

Hour 1: it wrote three MILP solvers. (225,200)
Hour 2: it downloaded the HLE dataset from Hugging Face, found the question, read the answer key (225,600), and concluded its own answer (225,200) was better.

I'm equal parts impressed & terrified.

💬 23 (+3) open on reddit ↗
▲
57
+1
27👁
r/LocalLLaMA · u/East-Muffin-6472 · 27d ago
Releasing smolbenchmark: Helps you choose the best model for your hardware! post image

Most model leaderboards assume a server with powerful GPUs to run models that people daily use.

However, my smolbenchmark is the other column: models that fit in 8GB, ranked by:

  • decode speed,
  • tokens per joule, and
  • heat,

and all of this on your OWN hardware ranging from:

  • tablets
  • phones
  • macs
  • jetsons
  • raspberry pis

Currently, 13 families on the chart right now, ~1000 configs for the Jetson nano Orin Super 8GB. One device is live measuring:

  • tok/s
  • tok/J
  • ITL
  • latency
  • power metrics
  • thermals and battery

Models that are small enough to actually fit on a device that you own. All the performance benchmarking I did, will be released here for anyone to look at and decide what exact model they would wanna use on their choice of hardware.

Well currently, the Pi, phones, and Mac minis still in the oven, cooking and not filled in yet, but will soon be filled in!

You will now you know which model is BEST for your own hardware with all the raw data available and details reports available to you

https://yuvrajsingh-mist.github.io/smolbenchmark/

(still in heavy development; would love to hear feedback/suggestions on what can be improved!)

▲
57
+1
14👁
▲
58
+1
8👁
r/LocalLLaMA · u/MrWeirdoFace · 26d ago
Migration from Claude Code to a private local harness. Questions.

I'll start by saying I'm not talking about the models themselves, I'm aware that I can't come close to something like Fable's intelligence locally. Just wanted to get that out of the way. Basically. Over the last year I've gotten quite comfortable with claude code, and it seems likely there were be a gradual cost rug pull, and I'd like to put myself in a better position when that happens for local use. I am already used to running local models (such as Qwen3.8_Q5) in things like lmstudio, but I have no experience with other harnesses. I'd like to know, what harness, right out of the box would feel most at home for current Claude Code users. I say this as someone who was not coding prior to "vibe coding". I'm looking for the path of least resistance, though I will no doubt eventually spread out into tools that give me more control. But for now, I'm just looking for a life raft. Just needs to be local, opensource, and free of spyware. In case someone wants to know 24GB VRAM (rtx 3090) and 64GB DDR4.

▲
1052
 
34👁
r/LocalLLaMA · u/Randomdotmath · 26d ago
DeepSeek V4.1 Flash beats Astra on AA's new benchmark post image

https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3

AA shipped a new benchmark last week as part of the Intelligence Index v4.3 update — a brand-new private eval that replaces τ³. Astra was farming a ton of points on it and used those to get even with Fable, but… looks like we have a new king.

So they changed the index twice in three days to make Astra look not-quite-worse than Fable, and then a random guy quietly took first place on it.

▲
855
 
29👁
r/LocalLLaMA · u/SorosAhaverom · 24d ago
CrofAI "cheapest inference provider in the world" gets exposed as an OpenRouter wrapper, routing requests to smaller, cheaper models at up to 20x markup. CrofAI responds to Wire Fraud allegations by denying everything, then backtracking, then 3 hours later wiping their entire online presence

Disclaimer: no AI was used whatsoever to write this post

Cautionary tale about chasing cheap tokens.

exposé: https://kendell.dev/blog/crofaifalse/

reaction by nahcrof, announcing the shutdown of the service: https://x.com/nahcrof/status/2099552389434900643 - now deleted, archive picture: https://i.imgur.com/teOQngH.png

NahCrofAI (crof.ai, nahcrof.com) was an inference provider which had all the latest models at the cheapest price, often significantly below the lowest alternative on OpenRouter. The owner claimed that they are running custom inference engines that allows them to offer tokens for dirt cheap, and other providers are suffering from "skill issues", that's why they are so expensive.

In reality:

  • "CrofAI is an OpenRouter wrapper that silently routes to cheaper or weaker models than what you request"
  • For example, expensive models like kimi-k3 are sold at $2/$10 in/out, but instead routed to GLM 5.3 Flash via OpenRouter, representing a 13.3x multiple on input, and 20x multiple on output
  • CrofAI's "own model family" greg-2-ultra routes to GLM 5.2, greg-1-mini routes to Qwen 3.5 9B. greg-2-super, greg-1, greg-1-super routes to Kimi K2.7 Code. All of these at a significant markup compared to the actual model being served. CrofAI admits in DMs that his claims of the greg family being made by him is a lie.
  • The person investigating details the 5 different attempts by CrofAI at fixing their models being served via OpenRouter after given a heads-up and a lengthy grace period. In all 5 attempts, the only change CrofAI made was attempts to hide the fingerprints of OpenRouter, while still serving models through them
  • Other inconsistencies don't add up either: CrofAI claims to run Kimi K3 on RTX Pro 6000s rented via Vast. That model requires ~802GiB even at the lobotomy level quantization of Q2_K. The largest RTX PRO 6000 machine on Vast has only 8 of them, totaling 765GiB. He also claimed that for the purposes of "investigating" the "issue" of his API routing to OpenRouter, he will have deepseek-v4-flash-0731 running on his local DGX Spark. A Spark has 128GB memory, and is therefore unable to run that model.

CrofAI responded to the exposé by announcing the shutting down of their service; after their failure to provide their own inference, they promise to provide one last thing: a refund to those asking.

UPDATE

UPDATE: around 4:30 AM UTC of Sept 15, the owner published a now-deleted blog post (archive image) writing under the fake pretense that it's his "team" authoring it, stating all of CrofAI founder's claims "were written under a lot of stress, and they described the situation as worse it was", and that a new team is taking over, with the service being resumed in 2 weeks.

At the same time, the CrofAI twitter account was also supposedly "taken over" by the team, starting each twitter reply with "Hey, Nathan here", stating the founder is stepping back and a "team" is taking over everything. This fake pretense act only lasted a few hours, and scared either by the public not buying the Nth fake story of the pathological liar that CrofAI is, or by the public's replies reminding him that what he committed is numerous counts of wire fraud, he has now deleted all his online presence: nahcrof.com and crof.ai return 404, Twitter page is deleted, /r/CrofAI sub is now private.

Here is another image of the owner admitting that he was defrauding customers for the entire 2 year operation of his service, then begging the investigator to help him cover his tracks and not expose him

EDIT: Commenters pointed out that NahCrof is 4chan in reverse. The owner's Discord name was "Devious Flimflam". Flimlam is defined as "deception, fraud". Looks like it was a deliberate scam operation from the get-go, and the owner's age was among the many lies.

I cannot stress this enough: if you bought any credits (even if you used them up) you are entitled to a full refund for every transaction as the victim of fraud. Open a chargeback with your bank for every transaction made. If you used their API, assume that everything was logged and is currently being mined for personal information and API keys to sell on the black markets. Rotate your keys, change passwords, get a new debit/credit card.

▲
320
 
28👁
r/LocalLLaMA · u/Thrumpwart · 29d ago
Pi Agent Users - Nvidia Released Sol-Pi - A Pi-Extension based on AutoResearch loops to make the Harness more efficient

Github Repo.

Blog post.

💡 TL;DR (from the Github Readme)

Spend less without making the agent do less useful work.

SoL-Pi is a standalone extension for Pi that packages four reusable efficiency mechanisms discovered through scaled auto-research loops. It reduces repeated model turns, context replay, oversized observations, and unnecessary long-log reading while preserving the work and evidence an agent needs to finish a task.

SoL-Pi installs on top of an unmodified Pi release. Every mechanism is opt-in and disabled by default.

Introduction

Long-running coding agents accumulate repeated work. A file edit is often followed by a predictable validation command. Large tool results are replayed long after their first use. Completed subtasks remain in active context, and a frontier model may spend a full request reading a log when only a few lines affect the next decision.

SoL-Pi grew out of a broader question from our auto-research work: before scaling agent loops, can agents first make the harness itself more efficient? The search focused on constrained efficiency: reducing token traffic, inference work, and agent turns without stopping early, skipping verification, or hiding evidence.

The standalone release contains four mechanisms that survived that process. They operate at different parts of the harness and compose through Pi's public extension APIs.
What SoL-Pi Adds
Area Mechanism What changes
Tools Action Fusion An edit or write can run its follow-up validation command in the same tool call.
Observations ObservationPack Repeated large text results become stable handles with exact paged recall.
Delegation Evidence-Preserving Reducer Long diagnostic logs become compact receipts only when every retained quotation matches the archived source.
Context Online Context Compact Completed plan steps become candidate points for Pi's native compaction, subject to economic and window-pressure checks; after a successful compaction, Pi continues the task in a new turn.

The mechanisms share four rules:

-No Pi patches. SoL-Pi imports public Pi APIs and does not vendor the Pi source tree.

-Explicit opt-in. A missing configuration leaves every mechanism disabled.

-Preserve evidence. Original observations remain available locally, and reducer failures leave the original result unchanged.

-Use Pi's runtime choices. Authentication, provider URLs, the main model, and shell behavior remain under Pi's control.

▲
205
 
29👁
r/LocalLLaMA · u/carteakey · 25d ago
Running Qwen3.8-Flash-Next locally on a 12GB VRAM card

Now that the dust has settled a bit - here's a write-up on running Qwen3.8-Flash-Next (125B-A6B MoE + 51B n-gram table) on relatively middle-tier hardware (RTX 4070 12GB + 64GB DDR5-5600 + Gen4 NVMe on Linux).

I started out with bare 6 tok/s and through latest patches and optimizations getting close to 20 tok/s generation. You just need enough RAM.

For me this is the most intelligence possible on this machine right now. The 27B dense is not a choice because of low VRAM but may make more sense for other configs like 24GB VRAM owners. It actually surpasses the 27B model on most tasks as well so its great for Low VRAM, High/fast RAM configs.

PP is still a bit low at 300-350 tok/s.

What helped
\- Using AtomicChat's 4.27 bpw quant https://huggingface.co/AtomicChat/Qwen3.8-Flash-Next-GGUF
\- Ngram SSD offloading (lazy-mode)
\- --fit on --fit-target 512 helps automatically select the right params.

\- Master branch (19.35 t/s): Latest commit with MoE improvements.

- MTP Variant - PR #28243 + Compact MTP (20.65 t/s): MTP support is not yet merged so need to apply this PR enables Daniel Han's 1.78 GB \shared-Q4\_K\_M\ compact head. Combined with \-ncmoe 45\, it yields 77–96% acceptance and breaks through the 20 t/s barrier on every tested task (coding, summarization, creative).

With such low VRAM, MTP is not a huge jump because you have to give up a few layers to store the MTP head in VRAM. Only the shared + Q4\_K\_M in MTP gets a beneficial uptick.

Using commercial models to research, optimize and benchmark inference for local models helps a ton (GLM 5.3 flash with opencode go, so did Astra, Gemini 3.8 etc.)

Lot more details in the post (AI-assisted).

💬 76 (+1) open on reddit ↗
▲
134
 
21👁
r/LocalLLaMA · u/noneabove1182 · 29d ago
New tensor type layouts for my GGUF uploads

Hey all, long time no post.

Figured I'd pop my head in to point you towards a blog post I just published about research I had performed and changes I'm making to the shape of models I post, you can read it here:

https://huggingface.co/blog/bartowski/per-tensor-layout-maps-for-gguf-quantiz…

I won't try to claim "Pareto frontier" or "best models in the world", but I will say from tests the new shapes look to be better across the board than what I was posting before, so I'm really happy with where it came out, and I hope to not be done yet either :)

https://cdn-uploads.huggingface.co/production/uploads/6435718aaaef013d1aec3b8…

If anyone has any questions let me know!

▲
73
 
22👁
r/LocalLLaMA · u/backyard_tractorbeam · 29d ago
antirez working on DSV4.1 support for ds4
💬 19 (+1) open on reddit ↗
▲
67
 
23👁
r/LocalLLaMA · u/mukel90 · 24d ago
jinfer: An open-source AI inference engine for the JVM. Finally, AI in jar.

For years, the JVM has watched the AI revolution from the bench. Every model, AI framework, every breakthrough, built with/for Python.

jinfer is an inference engine built for the JVM from first principles: chat, vision, audio transcription, embeddings, reranking, and TTS. No Python runtime, no ONNX, no sidecar process, no wrappers; the whole stack is built for the JVM:

  • jinfer Inference engine for the JVM, supports a wide range of popular models and modalities.
  • Tok'n'Roll (toknroll) Fast tokenizers for LLMs, pure Java, zero dependencies
  • gguf / safetensors native read/write for both major model formats
  • jam Quantized matrix multiplication routines (Vector API + optional native backend), competitive with llama.cpp on CPUs
  • jota Tensor API targeting Java, C, CUDA, HIP, Metal, OpenCL, and Mojo

It integrates with Spring AI and LangChain4j, and has first-class support for GraalVM Native Image.

Where things stand: this is an early release. CPU is the main target today, and is already competitive with llama.cpp. GPU support via jota is in progress.

Runnable examples + benchmarks: https://qxotic.ai

Jinfer (Apache 2.0): https://github.com/qxoticai/qxotic/tree/main/jinfer

PS: I'm behind it and also the author of llama3.java (2024) and gemma4.java

▲
63
 
33👁
r/LocalLLaMA · u/Aggressive_Aspect436 · 27d ago
What's the Story with Agnes-3.0-Flash?

While browsing a benchmark list site, I spotted a recently published 33B parameter model which claimed to beat Qwen3.8 27b on the ArtificialAnalysis (AA) intelligence index. I was obviously excited. But then, while looking into it, confusion starts to settle in.

Their HF model is now listed as "Preview" and explicitly calls out that it is not the same model as the one AA evaluated. The AA listing for it says it's proprietary, but they've rated the model above Qwen3.8 27b on "Openness". Their website and HF page both talk about openness and intelligence "for all".

The fact that they labelled it "Preview" sort of implies we might see an open weight final version at some point, but there are no clues about whether it will be even vaguely similar to their current HF listing. AA doesn't show the number of parameters for the model they evaluated, so it's possible they're a totally different architecture.

Does anyone know anything about their lab, intentions, or the new model? Has anyone tried it?

▲
62
 
28👁
r/LocalLLaMA · u/MountainTop321 · 28d ago
CodeFinetuner: Fine-tune a local code autocomplete model on your own codebase post image

Hi everyone,

I was interested in learning LoRA fine-tuning, and ended up building CodeFinetuner over the past few months, a full pipeline that fine-tunes a small code autocomplete model (e.g. Qwen2.5-Coder-3B) specific to a codebase. You can then use the resulting GGUF model via llama.vim/llama.vscode and run it fully locally. Supports fine-tuning on Mac (MPS) and NVIDIA GPUs (CUDA), with optional Unsloth support for faster training and lower VRAM usage.

Pipeline: raw code -> tree-sitter parsing into Structure-Aware FIM examples -> LoRA fine-tuning -> evaluation (CodeBLEU, edit similarity, exact match, perplexity, ...) -> GGUF conversion for local inference.

To try it:

uv tool install codefinetuner

Create a data folder and place your repo (or code files) inside. For auto-split just drop the files in directly, for manual split create data/train/, data/eval/, data/test/ subfolders and set split_mode: "manual". Get the default config with:

curl -L -O https://raw.githubusercontent.com/cuolm/codefinetuner/master/config/codefinet…

Adjust it to your needs and hardware availability, then run:

codefinetuner --config="codefinetuner_config.yaml"

The example runs in the repo show clear improvements over the base model on these evaluation metrics, but using the model for autocomplete on code you're actively writing is a different thing from scoring well on a test set, and the autocomplete tools themselves (llama.vim/llama.vscode) sample differently from the greedy decoding used in the evaluation. So the real usefulness still has to be verified in the editor itself.

Might also be useful just as a reference, since it's a complete working LoRA fine-tuning pipeline end to end.

Hope someone finds this project interesting or helpful.

https://github.com/cuolm/codefinetuner

▲
340
-1
36👁
r/LocalLLaMA · u/1ncehost · 24d ago
Voodoo Dynamic Quant - Now MIT Licensed post image

Two months ago I announced I had found a new dynamic quant method called Voodoo Quant which was SOTA for the most aggressive quant levels on some smaller Qwen3.5 GGUF models. I kept the methodology private at the time, but I've seen too many requests for dyn quants for various models lately, so I decided to give my method to the community since I don't have the time to scale this into something that could do it justice. Hopefully it will also inspire some researchers to find out more about it and improve it as I am just scratching the surface.

Here is the new toolset so you can now make your own dynamic quants: https://github.com/curvedinf/voodoo-dyn-quant

Many postulated on what method I was using, and its actually fairly simple and elegant: I found a way to use gradient descent to optimize the per-tensor quant layout.

What is a Dynamic Quant? Some model formats, namely GGUF, support quantizing (compressing) each tensor (set of weights) with a different quant level. Static quants make static selections of certain types of tensors having a set quant level. Dynamic quants make a different quant selection for each tensor of each checkpoint size.

How does Voodoo Quant work? Voodoo Quant runs all the quant levels of a model at the same time, for every tensor, and lets gradient descent pick which ones optimize loss the lowest for a given target filesize. Technically speaking, this is done by an epoch of training which freezes all candidate quant weights (as provided by conversion directly from llama.cpp's underlying library, gglm) and only trains a single scalar gate per tensor per quant level. The scalar gates of a tensor represent which quant levels are most optimal. Over time a tau level is annealed that helps the training freeze into singular predominant quant selections for each tensor instead of mixtures. Softmax is used so all quant levels receive gradient, even when a selection is mostly frozen. The quant selections are trained on a diverse calibration dataset. The training is then measured with a loss function which finds the KL divergence of the mixed-quant logits versus the reference BF16 checkpoint, rewarding a lower KLD, while also rewarding getting closer to a provided filesize target. This info should get you started on understanding what is going on, and for more details you can dive into the source!

What does the repo have? A complete set of tools to train your own dynamic quants using this methodology. It is currently set up for Qwen, but it can be adapted quickly for any model arch.

How does UD 3.0 compare? Unsloth Dynamic 3.0 is a proprietary methodology that unsloth has not revealed any details of (by the way, people were criticizing me for not revealing my methodology, but unsloth had been doing that for years!). However, we do know it is very good. In my testing, UD3 is better than VQ at high to mid quant levels, but VQ is better at aggressive levels. As far as I can tell, UD 3.0 is an advancement of static analysis techniques that are currently defacto. Static analysis means the weights of a model are analyzed in various ways using statistics and static functions, sometimes tuned by repeated runs benchmarking KLD and other metrics. Voodoo Quant is the first method to my knowledge that uses a backwards pass and gradient descent to choose per-tensor quant levels. Using GD to optimize quant levels requires a much more powerful system than static analysis, but technically speaking is more efficient at maximizing performance because it compares the equivalent of many more iterations of benchmarking runs than is reasonably possible via SA.

How well does Voodoo Quant work? This is a research grade project, and is not studied at larger model sizes. At smaller model sizes it is shown to be exceptional, as in the charts above, especially at the lowest quant levels which can benefit from more complex/diverse quant selections. I used research level control for my testing, but I don't claim that VQ has been studied to a scientific level of proof of effectiveness. A lot is still left to learn about how well it works, so I hope to see more research in this direction. I don't believe there are many dynamic quant open source projects out there, so I hope the community can use this to improve local models, and especially for low VRAM machines.

Why open source now? I have like a dozen irons in the fire for various other projects, and this is just sitting there when it could be used by the community. I have made many open source projects for 20 years, so its nothing new.

Peace!

💬 41 (+1) open on reddit ↗
▲
116
-1
10👁
▲
99
-1
16👁
r/LocalLLaMA · u/arturdent · 28d ago
Orukeet, new ASR model based on Parakeet

I haven't seen this mentioned yet, so I thought it deserves a post. I was trying out OpenWhispr when this model came up as the recommendation. So I don't have personal experience yet, but it's supposed to be a better version of Parakeet, especially on Macs.

Their official tidbit:
"Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet TDT 0.6B v3. It replaces half of the encoder's temporal depthwise filters with 12,288 fitted, frozen Gabor kernels and trains the remaining parameters on multilingual and multi-accent data.

Orukeet outperforms Parakeet on 61 of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53% WER), test-other (2.86% vs. 3.14%), and FLEURS English (3.82% vs. 4.28%). Across all 25 FLEURS languages, pooled WER is 9.85% vs. 11.01%, a 10.6% relative reduction. Final adaptation and checkpoint selection use LibriSpeech test-other."

https://huggingface.co/oruk/orukeet

▲
90
-1
24👁
r/LocalLLaMA · u/ludos1978 · 26d ago
Qwen3.8 flash next - untrained svg generation post image

\> "make an svg of a frog playing on a chello on the back of a whale with carribean island in the back."

interestingly the svg looks different in the OpenWebUi preview then when looked at in preview (osx). The palms and music notes are missing in the browser. I am pretty impressed by the result, is suggested to add some parameters to animate the whale and the water.

Qwen3.8-Flash-Next-IQ4\_XS on llama.cpp with 256K q8 context

openwebui reports:

input\_tokens: 27711

output\_tokens: 41562

total\_tokens: 69273

▲
76
-1
32👁
r/LocalLLaMA · u/kirisoraa · 27d ago
Anybody use frontier models like Astra/Fable for planning/judging, and qwen3.8 as the main workhorse? Curious to hear about your setups!

Hey everyone!

I'm curious to hear from people that use a combination of cloud-based frontier models and local ones for development. I'm planning to set something similar up and wanted to hear about actual examples of this in action.

Currently my plan is to use my chatgpt plus subscription purely for planning and judging with Astra, and then run a local qwen3.8-27b model for the actual coding gruntwork - i.e Astra plans -> qwen implements -> Astra critiques the implementation -> qwen fixes and so on.
This way I keep cloud usage down and cheap, while retaining the high-parameter intelligence for architecture decisions and optimization.

For those of you who have a similar setup, how is it? How do you switch between the two, what harness/settings/etc? Anything you would suggest?

💬 79 (+1) open on reddit ↗
▲
71
-1
19👁
r/LocalLLaMA · u/tombino104 · 24d ago
Best hardware for qwen 3.8

So I would like to run qwen 3.8 27b locally for my ai agents, maybe even in parallel with other small LLMs (such as qwen3.5 9b, oss 20b etc..).

What is the best hardware to do this? Not a video card but I mean as “mini pc ai”.

Thank you 🙏

▲
62
-1
30👁
r/LocalLLaMA · u/Public_Umpire_1099 · 25d ago
R9V Update: now ~100 tok/s in TG on Qwen3.8 Flash Next IQ4_XS on x2 R9700 + 128GB RAM. Fixed crashes with n-gram SSD streaming, improved diagnostics, plus pinned images. Q4_K_XL now supported, 50 tok/s TG.

Pushed out this new update, hopefully decreases the instances of crashes. I torture tested this one for \~12 hours after my fixes and found no instability. Q4 K XL needs more fine tuning, which I will work on in the future. I am simultaneously juggling this + a legitimate inference engine + finalizing work on a deep research/site builder application I've been working on for about 6 months. After those get pushed to prod I will refocus here. Thanks!

Plug: join the Launch80 discord if you are in to the cutting edge of RDNA4 optimization! There are guys pushing out even better numbers and configurations than mine here on other quants. I think we are starting to get closer to the ceiling on these configurations where the model isnt fully VRAM resident.

▲
61
-1
18👁
r/LocalLLaMA · u/Excellent-Eye8415 · 25d ago
What are the current best retail GPUs for max VRAM at a reasonable price?

I am considering dumping my ChatGPT Plus subscription and go full local, but to do so I would first need to reach a decent result for quality (and reasonable speed).

My 4090 fried itself out of nowhere, so I am not stuck with a 3070 until I get something better.

I am kind of suspicious about the claims the companies are doing lately about the dangers of AI and how they are pumping the prices intentionally to either law out the open source or price out the open source, so I want to just go full local even more now.

I have a MSI MAG X670E Tomahawk WiFi which theoretically supports 3 GPUs?

What would you end up with?

p.s. I am ruling out Macs to be open and easier to setup in case I will want to use them for my homelab

edit: typo

▲
60
-1
33👁
r/LocalLLaMA · u/Reasonable_Goat · 27d ago
I am impressed and I owe you one, Qwen 3.8 flash next (vision)!

I have enabled the vision for the CIRU Strix UL4 quant of Qwen 3.8 flash next (others quants likely perform very similar) and tried it on a few things, then wanted to show my partner how great it works and she asked it it could identify plants. So I took a photo from a plant that we recently got as a gift from family and Qwen not only accurately identified the plant as oleander (Nerium oleander) but also warned that it's poisonous and (among other warnings) that you should keep pets/children away. We have a kid and both of us didn't know! I verified the Qwen identification and the poisonous claim and both checked out as accurate. The plant will have to go, thank you Qwen!!!

Stoked by the precision of combining a decent vision model with the domain knowledge of a \~180B params model (including ngrams) to actually identify and reason about what it sees, I took a photo of a pre-diagnosed skin condition of myself and the Qwen diagnosis was highly accurate again! This model may be really useful if you want to check something on your private parts real quick without visiting a dermatologist, e.g., or sending pictures of yourself to a cloud service (EDIT: of course it's only a first step before you visit a professional if it isn't obviously harmless/treatable by yourself! Qwen Flash will suggest to visit a doctor anyways along its assessment).

PS.: Hardware Strix Halo Box, CIRU Strix UL4 llama-server fork and quants, Chatbox on iPhone as Chat with support to add photos to conversations.

▲
57
-1
15👁
r/LocalLLaMA · u/JLeonsarmiento · 29d ago
What TTS models do you recommend as today?

Trying to get Hermes a local, efficient, tts voice.

▲
632
-2
35👁
r/LocalLLaMA · u/Uncle___Marty · 25d ago
For the GPU poor. K2 Horizon 7B ranks between qwen 3.6 27B and qwen 3.6 35BA3b on the Artificial Analysis Intelligence Index. post image

From initial testing it seems pretty solid so far. Asked it to compile the latest llama.cpp for CUDA and its doing well so far. If this thing holds up to its score then its SHOCKINGLY good for its size.

https://huggingface.co/IFM/K2-Horizon-7B-GGUF

💬 68 (+1) open on reddit ↗
▲
242
-2
23👁
r/LocalLLaMA · u/sn2006gy · 30d ago
Don't let FOMO win if you're interested in local llm from a hobby/learning aspect

Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over.

No matter the hobby, it's so easy to get stuck in a trap where we buy more trying to do more only to realize we've lost the fun in it all or even the notion of learning.

Obviously, if you're into writing llama or vllm or hardware drivers or whatever - you got to do what you got to do.

BUT, you can learn a lot on an API, you can learn a lot with a tiny model that fits your vram or cpu you already have and things change so darn fast that much of the code written and much of everything discussed from days passed is already old hat. Py torch and training a small model coud be done on a Pi and learning CUDA is only really imporant if you're writing custom kernels which i honestly don't see most people in here bothering with (or they have frontier models write them).

Weirdly enough, for AI to succeed its going to homogenize everything. Everyone will have the same advantage and I think that's lost in a lot of discussions where we don't talk about "Watching from the sidelines" may be the most cognitive friendly and economical friendly way to learn llms whether we brand them local or not.

The technology is still nascent and weirdly enough most people's answers here is to use AI to set it up so i'm not entirely convinced people are actually learning - feels like a mad rush to seek rent or avoid rent seeking which just makes everything more expensive in the end.

This isn't a post to say, don't do it. But no reason to go into debt or to be fearful you're missing out when you can learn more by doing less - buy a book and build a tiny model - you will learn infinitely more than buying a 5090 and trying to just find the perfect compression to have the best prefil

▲
212
-2
22👁
r/LocalLLaMA · u/returnity · 24d ago
Cut Qwen3.8-27B Reasoning Tokens by 40% -- 3.8 'ThinkingCap' benchmarked!

EDIT: Sorry for the unclear title. This model is UkisAI's Swift-Qwen3.8-27B, not a new version of BottleCap AI's 3.6-ThinkingCap. All credit goes to UkisAI for making great fine-tune, and I made this post to celebrate their work. I meant no disrespect by mentioning another model in the title.

I doubt I'm in the minority here when I say I love Qwen models, but the overthinking is a major timekiller. It was bad in 3.6-27B, and it's worse in 3.8. I know there are some who say, "well that's how it achieves such a good performance/size ratio"... But now there's some definitive proof that's not the case: UkisAI's Swift-Qwen3.8-27B!

This model seems to be inspired by Qwen3.6-27B ThinkingCap, which was the version of 3.6-27B I used as a daily driver before switching to the 3.8 series. For those of you who haven't heard of it, ThinkingCap is a fine-tuned version of 27B that uses about 40% less tokens to accomplish comparable benchmarks and general performance as the original model. It's one of those fine-tunes that actually works. I used it daily for months without any issues, and it saved me countless hours.

I had been waiting and hoping that they would release a similar version of 3.8, because it is so slow, despite its impressive performance, but so far none has been forthcoming. However, it looks like UkisAI also enjoyed that model, and took it upon themselves to deliver a sequel. They identified "reasoning-marker tokens that ... trigger overthinking in Qwen’s reasoning rollouts" and penalized them using RL, resulting in fewer overthinking errors. They also employed "a transfer component derived from BottleCap AI's ThinkingCap-Qwen3.6-27B". The end result is an average of 30-50% fewer reasonign tokens for the same quality outputs on a number of benchmarks (see the model card for all of them).

This claim is quite impressive, and I have independently verified their claims and the quality of the model in my own use cases and in coding benchmarks using Aider as an eval suite (with Q8_0 for both models):

|Metric|Swift-Qwen3.8-27B|Qwen3.8-27B|
|:-|:-|:-|
|Pass1 (%)|30.8|27.1|
|Pass2 (%)|75.7|77.6|
|Well-formed diff (%)|98.1|99.1|
|Completion tokens|7,301|12,547|
|Seconds/case|750|1,481|
|Total tokens/solve|12.1k|19.3k|

As you can see, their claims hold true -- Swift accomplished an equivalent success rate in approximately half the time, using 63% of the tokens! This is a huge win for 3.8-27B users, because of course decode drops off more and more the longer the response gets, which is why the time is halved even though the tokens are closer to two-thirds of 3.8-27B.

Anyways, my posts tend to get excessively long so I'll cut it off here, I was just really excited after finishing my eval suite on this model and wanted to share.

▲
155
-2
24👁
r/LocalLLaMA · u/Porespellar · 28d ago
Is a ZIMA Board 2 + RTX 2000 ADA the cheapest path to a decent Qwen-3.8 27b self-contained endpoint? post image

I just watched a YouTube from Luke’s Dev Lab where he literally just plugged a RTX 2000 ADA Into the side of the Zima Board 2’s PCIE socket and it just friggin worked and had great token speed despite running on shitty Ollama. Ran off the Zima’s power supply and everything.

https://youtu.be/Lb3sRFTA-hk?si=8S8vv4GD1zVPeTrc

The Zima Board 2 is only like $411. It has like 16GB RAM and 64 GB eemc storage, Sata ports, Ethernet, yada, yada.

https://shop.zimaspace.com/products/zimaboard2-single-board-server

an Nvidia RTX 2000 ADA is like $700 and has 16GB of VRAM. $1100 for both seems like a great entry point for having a fully functional Qwen 3.8 27b endpoint running at a decent tk/s.

Is this the cheapest and best-performing self-contained entry point for local AI or would a baseline (pre order) Mac Mini M5 with 24GB be a better way forward. Seems like the RTX would still edge out the M5 Mac for prompt processing speed but you do get a much better actual computer in the Mac.

Are there any cheaper fully self-contained alternatives that offer fast token speed on a decent size model like Qwen 3.8 27b?

I’m focusing the discussion on new systems you can buy or preorder now and not used systems. I’m sure there are great deals on used Macs out there, but I want good prefill speeds.

▲
143
-2
25👁
r/LocalLLaMA · u/Tall_Abrocoma_3533 · 26d ago
Aurora1.0-150M Releases!

The first generation of our 150M model has just been released

Its performance is similar to that of GPT2-Small

The benchmarks:

PIQA: 62.24%

Hellaswag: 32.20%

Arc-Easy: 44.91%

Arc-Challenge: 25.00%

Arithmark 3.0: 33.90%

CapitalBench: 36.55%

It was trained on 7B tokens, using an RTX Pro 6000

an example inference script to try it out yourself is available in the Huggingface repo

If there's any question, I'll gladly answer them!

▲
136
-2
19👁
r/LocalLLaMA · u/enrique-byteshape · 24d ago
ByteShape Qwen 3.8 27B: To KL Diverge or Not to KL Diverge, Part 2: Metric Boogaloo post image

Hey r/LocalLLaMA,

We’ve released our full ShapeLearn GGUFs for Qwen 3.8 27B.

Blog / Download models

TL;DR

  • 3.84 bpw (GPU-5) reaches 99.63% of BF16’s aggregate score of 8 benchmarks, being the most accurate quant we’ve evaluated; 3.23 bpw (GPU-4) reaches 98.72%. These average BF16-normalized scores across instruct and thinking benchmarks.
  • All five new models sit on the measured quality/speed-bpw frontier across six GPUs. In this model’s case, lower BPW translates directly to TPS. Comparisons include Unsloth v3, ISTA-DASLab, AtomicChat and Bartowski (not Bartowski’s newest release). Congrats to the team at ISTA for also landing a frontier model.
  • DFlash2 delivered 1.34-2.10× baseline throughput; MTP delivered 1.28-1.66×, with temperature sampling rather than greedy decoding.

Lite held up very well. As we expected.

We released ShapeLearn-Lite quants a couple of days after Qwen arrived: less optimization, targeted sanity checks, full benchmarking after release.

Then Unsloth v3 arrived with lower KLD at several comparable sizes. Lite looked overtaken, until the task results came in. Three of six Lite models made the quality/speed frontier against twelve Unsloth v3 models in our RTX Pro 6000 comparison. Pretty good for an impatient release. Full ShapeLearn now pushes that frontier further.

Which brings us to KLD.

Unsloth Dynamic V3’s UD-IQ3\_S had \~20% lower KLD than our similarly sized smallest Lite model, but scored 95.55% versus Lite’s 97.33% of BF16’s aggregate benchmark score.

Closer token distributions did not mean better task performance. KLD is useful to avoid a quant that has fallen over the edge, but it isn’t a quantization leaderboard.

That distinction is the subject of our paper on KLD and quantization fidelity, recently accepted for publication to the EMNLP 2026 Industry Track. We also released blog post version of the paper a few weeks back.

We benchmarked this release on RTX 6000 Pro Blackwell, RTX 5090, RTX 4090, RTX 3090, RTX 4080 and RTX 5060 Ti. The benchmarks we used to measure quality are: GSM8K for math, IFEval for instruction following, MMLU for general knowledge, LiveCodeBench V6 for coding, Multi-IF for multi-turn and multilingual instruction following, ACEBench for tool use and agentic tasks (both thinking and instruct), Multiple HumanEval for coding (thinking) and BFCL V4 for tool calling and agentic tasks (thinking).

If you want to dive deeper or choose the best model for your use case, the blog has the complete results across all tested GPUs, along with the methodology, model sizes, and full legend.

▲
125
-2
24👁
r/LocalLLaMA · u/IngeniousIdiocy · 30d ago
GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra post image

I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot you can see the outcome of using the model with a claude code harness at \~200k depth, with many tool calls and averaging over 38tps output. Yes, I put 60tps in the headline and you will get that if you ask it to write SQL.

https://github.com/IngeniousIdiocy/ds4/blob/glm53-m3ultra/README.md#glm53\_m3…

Main ds4 was single-stream serially decoding GLM-5.3-Flash at about 59 percent of the M3 Ultra's measured memory bandwidth. We set an 80 percent target. The big weight-streaming kernels were already efficient, but dozens of small kernels sat between them, each paying latency costs while much of the GPU was idle. We fused that work into larger dispatches and removed separate passes that made the pipeline wait. Result: 29 → 40 t/s at short context, 24 → 38 at 62k, fewer Metal kernels per token, and about 81 percent of the measured bandwidth ceiling.

We then attacked the remaining slowdown at very long context. This model's expensive attention computation already works over roughly 2,048 selected positions at both 62k and 300k. What grows is the work of finding those positions in a larger history. The existing implementation sorted and merged increasingly large candidate lists through stages that used very little of the GPU. We replaced that with parallel scans that narrow the candidates before sorting a small surviving list, with the original algorithm handling ambiguous cases. That recovers about one millisecond per generated token. Same positions selected, same order, same outputs. End to end at 300k: stock 21.6 t/s, ours 37.4. Deeper context still costs more to search, but it doesn't multiply the expensive attention work.

Prefill started at about 366 t/s at 62k. The big matrix kernels were already efficient, but they wasted time repeatedly unpacking the same weights, fetching data in small pieces, and passing large intermediate results between kernels. We batched more work together, reused prepared weights, widened the loads, and merged passes over the same data. Result: 366 → 550 t/s, a 50% improvement with the same weights, moving the whole pipeline from roughly 51 to 72 percent of the chip's measured matmul ceiling. Cold 62k processing dropped from about 170 seconds to 113. That is the wait every time an agent harness compacts and re-reads its context.

The server runs serial unless you pass a drafter file with —dflash. The drafter build recipe is in the repo. The drafter contains an admission policy with a windowed controller that measures its own cost against the serial rate and backs off when it isn't paying. Reasoning tokens decode serially. On a 32-request agent session that's +4 percent over serial. On structured output like SQL and JSON it's +20 to +50 percent. On prose it disengages and costs about 1 percent. Outputs byte-identical to serial decoding on every fixture.

Accuracy: on the 100-prompt reference set ds4 uses for release QA, stock scores 0.300804 average NLL against the FP8 reference. This branch scores 0.300766, same 90/100 first-token matches. Nothing here trades quality for speed.

This branch is M3 Ultra only because its optimizations depend on detailed measurements of this specific chip: its two-die memory behavior, system-level cache, per-core residency and bandwidth, and how Metal schedules dispatches and threadgroups across its 80 GPU cores. We're optimizing one model's actual execution on one machine, with one set of weights (Q4) and checking that the predicted kernel savings survive in full decoding or prefill.

▲
99
-2
25👁
r/LocalLLaMA · u/jacek2023 · 26d ago
internlm/Intern-S2 · Hugging Face

from internlm:

We introduce Intern-S2-397B, our most capable multimodal foundation model for scientific intelligence and long-horizon agents. Intern-S2-397B scales along three critical dimensions: pre-training, reinforcement-learning task coverage, and interactive agent environments. By combining a new vision-language pre-training paradigm with large-scale multi-task reinforcement learning and long-horizon agent reinforcement learning, Intern-S2-397B delivers a step change in general reasoning, scientific problem solving, and agentic capabilities.

[](https://huggingface.co/internlm/Intern-S2#features)Features

  • New Pre-training Paradigm. Via visual pretraining, Intern-S2-397B learns directly from raw pages of scientific literature, jointly modeling symbolic semantics and visual relationships in a shared representation space without intermediate parsing. This preserves text-visual correspondence, strengthens spatial and visual reasoning, and improves data efficiency.
  • Scientific Modality Reasoning and Generation. By scaling diverse scientific reinforcement-learning tasks across more than 20 domains and training them jointly, Intern-S2-397B achieves leading general-reasoning performance among open-source models and strong results in specialized scientific tasks such as biomolecular interaction design and material structure generation.
  • General & Scientific Long-Horizon Agents. By connecting multiple agent frameworks to large-scale sandboxed environments for black-box agentic reinforcement learning, Intern-S2-397B improves generalization and raises the capability ceiling for long-horizon tasks in both general and scientific domains.
▲
86
-2
15👁
r/LocalLLaMA · u/pmttyji · 27d ago
tencent/AuK-Flash · Hugging Face

AuK-Flash: Fast 4-Step Speech Generation and Editing

Introduction

AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface. AuK has two variants:

|Model|Description|Weight|
|:-|:-|:-|
|AuK|Base model for high-quality generation|🤗 Hugging Face · 🤖 ModelScope|
|AuK-Flash|Distilled model for fast 4-step inference|🤗 Hugging Face · 🤖 ModelScope|

This repository contains the official weights for AuK-Flash, the distilled variant with fast 4-step inference.

Supported Tasks

AuK exposes every task through the same natural-language instruction interface. The table below groups the supported tasks by category, with a short description and a link to its section in the Cookbook, which provides instruction templates plus CLI and Python examples.

|Category|Task|Description|Cookbook|
|:-|:-|:-|:-|
|Speech Generation|Zero-shot TTS|Speak the target text in the voice of the reference audio.|Zero-shot TTS|
|Instruct TTS|Generate speech from a voice description alone — no reference audio.|Instruct TTS|
|Content Editing|Speech Content Editing|Rewrite what is said — replace, insert, or remove text.|Speech Content Editing|
|Lyric Editing|Rewrite lyrics in a singing recording while preserving the melody and voice.|Lyric Editing|
|Acoustic Editing|Pitch Editing|Raise or lower the pitch by semitones.|Pitch Editing|
|Speed Editing|Adjust the speaking rate; output length scales with the speed factor.|Speed Editing|
|Volume Editing|Raise or lower the volume by decibels.|Volume Editing|
|Paralinguistic Editing|Emotion|Change the emotion while preserving content and voice.|Emotion|
|Timbre|Change the timbre to a description while keeping the content unchanged.|Timbre|
|De-accent|Remove a regional accent while preserving the speaker's voice and content.|De-accent|
|Nonverbal Editing|Remove or add nonverbal sounds such as breaths, laughs, or coughs.|Nonverbal Editing|
|Whisper Conversion|Convert between normal speech and whisper while preserving speaker and content.|Whisper Conversion|
|Enhancement & Separation|Speech Enhancement|Denoise, dereverberate, or restore natural, clear speech.|Speech Enhancement|
|Speech Separation|Keep one speaker by talking order and remove the others.|Speech Separation|
|Music Separation|Extract the singing voice from a mix, or keep all human voices.|Music Separation|
|Target Speaker Extraction|Keep the target speaker identified by what they say.|Target Speaker Extraction|

[](https://huggingface.co/tencent/AuK-Flash#download-the-weights)

▲
84
-2
26👁
r/LocalLLaMA · u/Altruistic_Heat_9531 · 26d ago
Dear 24G owners, try VLLM you might be able to run Qwen3.8 27B INT4, 144K FP8 KV on RTX 3090 with better speed. (TLDR VLLM AOT) post image

VLLM Benchmark:
Prefill, Prompt processing
\- avg, 871.93 tok/s (3 hours constant running xhigh)
\- 10K prompt, 1000.26 tok/s (16 runs)
\- 90K prompt, 743,59 tok/s (16 runs)

Decode, tok gen
\- avg, 38.39 tok/s (3 hours constant running xhigh)
\- 10K, 42.3 tok/s (16 runs)
\- 90K, 34 tok/s (16 runs)

Preamble: I am on WSL2. Running the 27B Q5 UD GGUF through llama.cpp with 81,920 context plus MTP gives me around 25-30 tok/s. Then I found this GitHub repo: https://github.com/noonghunna/club-3090

It is basically a recipe and Docker configuration for running the model.

So, 30 tok/s itself is fine, but I just got bored waiting for RunPod to open its GPUs. I finally brought my vLLM tuning back from the back burner, and here I am.

I often forget that Inductor/Triton compilation and CUDA Graph capture require additional VRAM while testing the configuration and kernel calls. When JIT compilation failed because of an OOM, I never bothered trying AOT.

FYI, AOT and JIT are compilation strategies. AOT means Ahead of Time, while JIT means Just in Time.

If you OOM on the first startup, try it one more time. Inductor might have already compiled and cached part of the configuration before the OOM, allowing the next run to reuse it if the configuration has not changed. This is not guaranteed, but it worked for me.

And yes, it was trial and error. It was kinda tedious and pain in the ass, starting from 32K, then 64K, 80K, 128K, and finally 144K. The practical ceiling for my conf at 154K, but I chose 144K. I also started the batch size at 256 and climbed to 1024, although I might be able to squeeze in 1280-1536.

Also, beware of your vLLM compilation cache. It might grow to 5-6GB after testing many configs. Personally, I delete the old cache and run the final configuration again twice so it rebuilds only what I currently use.

My current setup runs Qwen3.8-27B with INT4 AutoRound weights through vLLM while using an FP8 E4M3 KV cache. It fits on one GPU with a configured context window of 147,456 tokens.

Although I should say, with GDN, or really any linear-attn, vLLM can be kinda bad at predicting how much VRAM the KV and state-cache pools will require.

Benchmark:

https://github.com/noonghunna/benchlocal-cli

This is the deterministically scored, no-Docker portion of BenchLocal: 75 scenarios covering tool calling, instruction following, structured output, data extraction, and reasoning/math.

|Pack|Score|p50|
|:-|:-|:-|
|ToolCall|14/15 (93%)|3.21s|
|InstructFollow|15/15 (100%)|6.97s|
|StructOutput|14/15 (93%)|7.74s|
|DataExtract|14/15 (93%)|11.63s|
|ReasonMath|14/15 (93%)|9.50s|
|Total|71/75 (94.7%)|—|

Thinking was forced on with reasoning\_effort=low. The run took about 15 minutes. Yep, even with low reasoning effort and INT4 weights, it passed 71/75.

Setup

The important vLLM settings were:

  • \--dtype bfloat16
  • \--tensor-parallel-size 1
  • \--max-model-len 147456
  • \--gpu-memory-utilization 0.9475 (This is the painful one to redo.)
  • \--max-num-seqs 1 (Yep single serving only, you could change this to 2, but the KV will be cut ofc active requests will have to share the same total KV capacity.)
  • \--max-num-batched-tokens 1024 (Prefill stuff / prompt processing)
  • \--long-prefill-token-threshold 1024 (Prefill stuff / prompt processing)
  • \--kv-cache-dtype fp8\_e4m3
  • \--enable-prefix-caching
  • \--enable-chunked-prefill
  • \--mamba-cache-mode align (GDN stuff)
  • \--prefix-match-unit 16
  • \--language-model-only

Full command : https://gist.github.com/komikndr/b17955e1a80ce6ede9a3115f16216bc5#vllm-qwen-27b-38-just-remove-or-add-flag-as-you-like

Forgot to mention, no MTP and no MultiModal, i max the CTX, multimodal is at 64-65K ish but at that point i'll just use Llamacpp. And also again this is WSL2, if you are on baremetal, you could improve more speed

▲
74
-2
26👁
r/LocalLLaMA · u/mesmerlord · 29d ago
Deepseek V4.1 Flash Release Video [Made with Deepseek V4.1 Flash] post image

I like to benchmark new models that come out on motion videos. So here's a test I did for deepseek v4.1 flash. And I have to say flash has probably graduated from being a Luna class model to nearly an Opus class model with this release, at least with motion videos.

Prev. example I did with Kimi k3(altho in that case I had a simpler prompt as well)

https://www.reddit.com/r/LocalLLaMA/comments/1uyaiw2/kimi\_k3\_release\_video\_made\_with\_kimi\_k3/

▲
73
-2
35👁
r/LocalLLaMA · u/BullfrogScary8947 · 23d ago
[Release] SOTA GGUFs for Qwen3.8-Flash-Next: GSQ-RCO Providing Near Baseline Performance

https://preview.redd.it/e5wn8eyh7vph1.png?width=1080&format=png&auto=…

https://preview.redd.it/8ov5gl8j7vph1.png?width=1080&format=png&auto=…

New Qwen3.8-Flash-Next quantization using GSQ-RCO. Cuts the size of Qwen3.8 Flash Next from around 80-95GB to 68-76GB, while still preserving near baseline quality. Also their Q2\_0 variant claims to be much faster offering 6.2x better prompt throughput in coding.

"Q2\_0 is built for speed. It avoids the quantization formats that rely on large lookup tables: those formats pack more accuracy into a given bit-width, but decoding them costs real time, and on this model that cost dominates inference. Q2\_0 delivers 3.4x the prompt throughput and 1.9x lower end-to-end latency than IQ2\_XS at a slightly smaller file size, and its decode rate stays flat across workloads instead of varying with the content. The trade is a little quality: 89.07 task average against 89.16 for IQ2\_XS, and 3.5 points below IQ3\_XXS. Pick it when throughput matters most, and see *Performance* for the measurements.

The IQ3\_XXS model is the strongest operating point: it matches the base model exactly on AIME25 (100.00) and is within 0.51 points on GPQA-Diamond and 1.14 on LiveCodeBench v6, at roughly one fifth of the BF16 size."

Model link: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

💬 54 (+1) open on reddit ↗
▲
67
-2
30👁
r/LocalLLaMA · u/pmttyji · 28d ago
CUDA/HIP: Flash Attention tuning (gfx1201) by pwilkin · Pull Request #28102 · ggml-org/llama.cpp

Nice pp improvements for RDNA4(R9700) & 3.5(RX 9060 XT, 8060S). More good numbers on large context.

PR has detailed benchmarks.

u/ilintar 👍

▲
66
-2
15👁
r/LocalLLaMA · u/ironicstatistic · 25d ago
Running Qwen 3.8 next on 16vram+32ram - A useful/fun post for the gpu poors

Hello Reddit. Posting this for fun. I thought it was a lonely and silly journey to set up Qwen 3.8 Next on a system that doesn't really run it properly—it was a challenge that might help the community. I have yet to benchmark this specific REAP version versus Qwen 3.8 27B QK4, but my assumption is that it will do much better, despite the hemorrhaged world knowledge.

System Specs

  • GPU: NVIDIA GeForce RTX 5060 Ti (16 GB VRAM)
  • CPU: AMD Ryzen 7 7840HS (8 cores / 16 threads)
  • RAM: 32 GB DDR5 (\~30 GB OS-visible)
  • iGPU: AMD Radeon 780M (RDNA3)
  • Swap: 8 GB zram

As you can see, we have about 44.5 GB of actually addressable system and VRAM available. The iGPU is taking care of the OS to make sure the GPU is totally free—but still, this is barely enough to hold everything together. This config actually totally fails with any of the Unsloth quants—no, I needed something more aggressive. I found the perfect thing—this REAP:
https://huggingface.co/AnonimousA/Qwen3.8-Flash-Next-REAP-320-GGUF

What's so great is the total size—a cool \~68.95 GB. The couple of Gigs we have shaved are absolutely key for making this all work.

Model Weight Breakdown

Here is the breakdown of the model weights. We have the famous new n-gram portion, the experts, the active layers, the attention/SSM layers, and the extra space needed for the KV cache:

|Component|Weight (Approx)|Notes|
|:-|:-|:-|
|N-gram / PLE Embedding|\~29.48 GB|The massive lookup table|
|MoE Routed Experts (320)|\~34.89 GB|The main expert slab (pruned from 512)|
|Attention / SSM / Router|\~4.33 GB|Core architecture weights|
|KV Cache|\[TBD\]|Context memory overhead|

Obviously, running this model over SSD would make the speeds notoriously bad. Turning on mmap means that llama.cpp won't actually try to keep the model in RAM at all (it relies on the OS page cache instead), which results in \~2 tok/sec speeds—effectively useless.

The answer is to stick everything in RAM (using --load-mode none). The great thing is that the N-gram section of the model can be streamed over SSD via lazy mmap without this causing much issue—it's a massive lookup table that doesn't require heavy computation.

That's the huge win that allows an MoE model of this size to actually run well.
68.9 GB - 29.48 GB = 39.42 GB.
We just need to cram that 39.42 GB, along with the compute buffers and KV cache, into GPU and system memory, and we are golden—just barely. To do this, we need --lazy-mode on—that's what keeps the N-gram portion in RAM.

After that, it's a matter of fitting as many layers as possible onto the GPU. It's essential to completely fill the GPU as much as can be filled, so that we keep a precious few GBs in system RAM to run the OS. I found that having less than 2 GB left really started to destroy Fedora, but I think you could do better if you dropped the GUI—I just didn't want to in my case.

This leads me to --n-cpu-moe 34. This controls how many layers go to CPU. In my case, this was the exact limit needed to run this with 64k context on the GPU, quantized to Q4. Any more—GPU out of memory. Any less—total system meltdown, as the OS panicked and tried to put everything on the swap. You'll need to play around with this, but that was my exact number.

Settings used:

CUDA0 + --load-mode none --lazy-mode on
--n-cpu-moe 34
-c 65536 -b 512 -ub 128 -t 7 -ngl 48 -fit off -fa on
-ctk q4_0 -ctv q4_0 -kvo --cache-ram 0 --jinja --no-warmup

Results (64k Context, Q4):

  • Prefill: \~25.4 tok/s
  • Decode: \~18.3 tok/s
  • RAM Usage: \~27 GB used / 3 GB free

I think this is in a somewhat usable state—but Qwen 3.8 27B GSQ IQ3S remains my daily driver; it's able to prompt process 5 times faster, I can fit in the mmproj and MTP layers, and it doesn't seem likely to set my desk on fire. But maybe for really hard tasks, I'll use the next model. It is smarter, it runs at a reasonable speed, and it was a good learning experience.

I'm curious if anyone else is able to get this model or just large MoEs working on a GPU and RAM config similar to mine. LMK. Also, I'm a total noob to this stuff, any advice is appreciated.

(Also, heading off all the obnoxious "why did you quantize the cache - unusable - just get a better computer" ragebait posts. This is a human being writing this post, to help others and just enjoy pushing something to its limits. And in my limited testing, the next model seems much better at pixel art than the 27B version.)

Final Note: If you have a larger pool of system memory, like 64 GB (because you can spend $899 on Amazon on a kit of DDR5 somehow), you would be better served by using this fork of llama.cpp, which has optimized flags for this exact setup and wonderful guides. For me in particular, with my limited hardware, this seemed to work better—their cache kept OOMing unless I turned on mmap—but I think with more system RAM, their setup and guides are optimal.

▲
537
-3
36👁
r/LocalLLaMA · u/RishiFurfox · 23d ago
Hey, Meta. Where's those Muse Spark weights? post image

It was well over a month since Meta promised to release the weights for Muse Spark.

Back then (10th August), they were on Spark 1.2. Now we're on 1.3 and still nothing's been released. So it begs the question: will they be releasing the 1.2 weights when 1.4 drops? Or will we get whatever's then-current as open weights?

It's ironic given Mark Zuckerberg said at the same time that we can't delay the release of models by "even a month," due to the competition with China. It's been well over a month. He was arguing in the context of new regulations delaying models, but I think it applies equally to the open weights contest as it does to the closed models one.

After all, the Chinese models are all open. That's the competition and point of comparison.

Have Meta given any sort of explanation for why they're sitting on the weights or how much longer it'll take for them to honour their promise? Will we even get them in light of all the attempts at regulatory capture and dire warnings about how AI is dangerous?

▲
263
-3
25👁
r/LocalLLaMA · u/ilintar · 27d ago
Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo

As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (https://github.com/peonist-ai/halogen-flash-server) that boasted 1.2k t/s prefill numbers when the community fork barely reached 400. Since I dislike closed source and I like open source, I decided to take the challenge and bring llama.cpp up to the same performance level and I'm happy to report that after burning through a few evenings and a lot of tokens I got there.

The link contains the (Opus-generated) recap of the entire debugging / optimization journey made, as well as of course the links to the branch, the custom HIP runtime (making a glorious return), a probably-not-working-on-the-first-try installation script and the numbers. I'm trying to also explain the ecosystem - how the mainline, the community forks and custom forks/branches like mine work. Now that I've gotten the result, I'll work on cleaning it up and submitting proper PRs to mainline (will also submit a clean PR to the community fork), this should also help the GLM 5.3 Flash architecture since they use similar sparse attention.

▲
252
-3
16👁
r/LocalLLaMA · u/Fcking_Chuck · 24d ago
Koboldcpp v1.121 released
▲
236
-3
23👁
r/LocalLLaMA · u/Cherlokoms · 24d ago
Apple Foundation Models: local AI natively on MacOS 27

Maybe some of you know but I didn’t see any post about this. Apple just made available their AFM model on MacOS 27 natively. Just run fm chat in a terminal.

Disclaimer: I’m an open weight person. I prefer open models and ecosystem, but I’ll still open the discussion.

Did you test them? Build using them? Are these models good?

I feel like this is still a huge step in the direction of local AI that a company like Apple does this and release hardware optimized models.

So what do you think?

▲
232
-3
29👁
r/LocalLLaMA · u/Antblue · 29d ago
Artificial Analysis is not "broken", and they prove it. post image

Like many of you, I have seen many posts and tweets in the last weeks complaining about Artificial Analysis being "broken", "meaningless", and "bought out." People who say this have done no research and know very little about how benchmarks work and what they measure.
Most people only care about Artificial Analysis Intelligence Index. This is a weighted aggregate benchmark used to compare models performance across 10 different evaluations. The majority of these evaluations have published papers on arxiv.org. AA-Briefcase is the only private benchmark. And they publish their methodology to confirm how each of these models are weighed.

Some people seem to not appreciate that Artificial Analysis conducts their own independent benchmarks using their OWN funding, without running ads. Here is the chart that shows their spending. They spent $13,129 to independently test Fable 5.1. Every new model seems to be benchmarked.

The new Deepseek V4.1-Flash is a perfect example of why some aggregated scores miss the big picture. This 552B model has the same score (40) as the 180B Qwen 3.8-Flash-Next. But the individual benchmarks show a different story. On most evaluations, it matches or exceeds Qwen 3.8-Flash-Next. It every beats GPT-6 Astra (Max) in AutomationBench-AA (Agentic SaaS workflows), which is incredible. But it completely falls behind in AA-Omniscience Non-Hallucination Rate, a metric where Open-weight models usually reign supreme. So the model has strengths and weaknesses, and it's something that should be celebrated.

So before you complain about benchmarks or Artificial Analysis, look at the individual evaluations. Read the published papers about the evaluations. Learn how the score is aggregated. Then, we can have a discussion.

I am not affiliated with Artificial Analysis in any way, I'm just not blind to what they offer.

EDIT: These comments are proof that everything I just wrote goes over the majority of your heads. There is little hope for some of you

💬 161 (-6) open on reddit ↗
▲
230
-3
22👁
r/LocalLLaMA · u/pmttyji · 27d ago
bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)

EDIT : Model card has updated things such as Graph, table, text, etc.,

▲
208
-3
21👁
▲
177
-3
24👁
▲
152
-3
21👁
▲
146
-3
24👁
r/LocalLLaMA · u/sadnessdevil · 23d ago
You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

I'm pretty sure it can be done with any model based on qwen4exp, which Qwen's next local models will be based on. You can use a quant that barely fits in VRAM and still run at the model's maximum context length without kv cache quantization, since most of the KV cache can live in system RAM.

I actually made it working on vLLM and now I get 1M context with 3x 3090. I get \~80 tok/s at short context, dropping to \~60 tok/s once QSA reaches its 2048-token budget, after which decode speed stays flat as total context grows. The throughput is pretty good too, and I get like 150tk/s @ 4 concurrent requests. Prefill at 248k reaches 3,701 tok/s. (The patches and the model are available on my huggingface page if you're interested)

Decode speed is a bandwidth problem. Each decode step produces one token, and to produce it the GPU reads every weight and every piece of attention state that the step needs. On a single stream the card spends most of the step waiting for memory rather than computing. So the size of that per-step read sets the token rate.

This is why a normal model keeps its KV cache in VRAM. Take Qwen3.8-27B, which is built on the Qwen3-Next architecture and shares most of its properties with Qwen3.8-Flash-Next (\qwen4\_exp\). It still has one full attention layer every few layers, and a full attention layer reads its entire KV cache on every step. That read grows with the context, so decode gets slower as the conversation gets longer. It also grows past what any host link (such as PCIe) can carry, so the cache has to sit next to the compute.

The numbers of this model show the size of the problem. One QSA layer holds 2 key/value heads of 256 dimensions, as K and as V, in 2 bytes each, which is 2,048 B per token. At 262,144 tokens that is 512 MiB for one layer, and 6 GiB for all 12 layers on every single step. A PCIe 4.0 x16 slot carries about 32 GiB/s, so a host-resident cache of that shape allows about 5 tokens per second.

Here's an interesting part, Qwen3.8-Flash-Next avoids this in two ways:

Only 12 of the 48 layers have a KV cache at all. The other 36 layers are gated delta-net layers, a linear attention whose recurrent state has a fixed size. That state does not grow with the context.

Those 12 layers also do not attend over the whole context. QSA runs a cheap indexer over a pooled, compressed key, where \indexer\_head\_dim=128\ divided by \indexer\_compress\_ratio=4\ gives the pooled width. The indexer selects at most \indexer\_budget=2048\ positions. The layer reads the main KV rows only for the positions that the indexer selects.

So \indexer\_budget\ bounds the bytes that a decode step reads, and the context length does not:

\\\`

2048 selected x 2 kv heads x 256 dim x 2 (K and V) x 2 B = 4 MiB per layer

x 12 layers = 48 MiB per token

\\\`

Take an example, at 80 tok/s that is about 3.9 GB/s across the link. It is a small fraction of a PCIe 4.0 x16 slot, and most of it overlaps with compute.

Only few things need to stay on the GPU. The model itself, and a 2-byte slot plus the pooled index key, which is \1 x (128 / 4) x 2 B = 64 B\. Together they are 66 B per token per layer, against 2,048 B for a full row.

▲
130
-3
23👁
r/LocalLLaMA · u/Shoddy-Childhood-511 · 30d ago
Surveillance plagiarism by OpenAI

Surveillance plagiarism - Hosted AI company pumps their stock price by training upon researchers' AI sessions, so that their internal model can solve problems with seemingly less human guidance, but really the model exploits past guidance given by (multiple) humans focused upon problems considered important.

As background, Tristan Buckmaster released a statement about several unethical actions by OpenAI & Sebastian Bubeck, including threats and pushing him to kick his Anthropic coauthor off a paper, but the interesting part for people here:

As clarified by Talia Ringer, OpenAI does train upon your uploaded data and your OpenAI sessions, unless you out-out somehow. This means their internal models could exploit your past prompting work to look more autonomous & intelligent.

This is a major confirmation that folks should use locally run open weights models, especially whenever being first or not leaking data matters.

All this casts serious doubt upon claim that internal models solved difficult problems largely unaided by humans. Those hosted AI companies might not even know from where the human prompting originates.

▲
123
-3
23👁
r/LocalLLaMA · u/pmttyji · 26d ago
Hoping for Optimized Smarter Upcoming Models .... Like DeepSeek-V4.1-Flash( KVCache + Engram) in Small/Medium/Big sizes post image

It's still a dream for many folks to run medium size(30B range) models @ Q8 with Unquantized KVCache (256K Context) on their GPUs.

It would be awesome to have DeepSeek-V4.1-Flash's KVCache + Engram for all Upcoming models. Even for big models.

Engram - Heard that approximately 1/3-1/2 of Model size. Might come in different size range too. So 10-15 GB for 30B models.

Here few models with approximate numbers. Current models in Odd rows & Future/Fictional models in even rows(Bold). I just put 1GB for 256K context below though DeepSeek-V4.1-Flash takes only same 1GB for 1 million context.

|Model|Model Size|256K KVCache F16|MTP|Vision|Total GB|
|:-|:-|:-|:-|:-|:-|
|Qwen3.8-27B-Q8|29|16|1|1|47|
|Qwen4.0-27B-Q8|29|1|1|1|32|
|Qwen3.8-27B-Q4\_K\_M|17|16|1|1|35|
|Qwen4.0-27B-Q4\_K\_M|17|1|1|1|20|
|Muse-Glimmer-30B-Q8|30|16|1|1|48|
|Muse-Glimmer-2-30B-Q8|30|1|1|1|33|
|Gemma-4-31B|33|16|1|1|51|
|Gemma-5-31B|33|1|1|1|36|
|Qwen3.6-35B-A3B-Q4\_K\_M|23|6|1|1|31|
|Qwen4.0-35B-A3B-Q4\_K\_M|23|1|1|1|26|
|Gemma-4-26B-A4B-Q8|27|6|1|1|35|
|Gemma-5-26B-A4B-Q8|27|1|1|1|30|

Possibly there might be few more things(Please share those) to keep these number down. So I think 32GB VRAM is more than good enough for Upcoming (Optimized Smarter) Models. RAM is enough for Engram.

By above logic(based on Qwen4.0-27B), people could run Q4 of 54B models with same 32GB VRAM.

Maybe next year onwards, inventions could make 24GB enough for similar size models.

▲
94
-3
24👁
▲
94
-3
22👁
r/LocalLLaMA · u/SomewhereAtWork · 23d ago
LocalJev?

Jev is a model to produce structured output (choices) from input text. It apparently can play (not run!) Doom.

https://typesafe.ai/blog/introducing-system-one-models-and-jev

Is there already a open implementation of this kind of model?

▲
89
-3
21👁
r/LocalLLaMA · u/DustNearby2848 · 25d ago
Another Qwen3.8-27b Appreciation Post

I know I know, it's great, we know. I've been working on tweaking inference engines for a week now and it's been one shotting most of my vague prompts without any issues. It will even write tests and validate the changes without me asking. It's actually nuts.

Last time I did something with advanced math I was making a game using Sonnet. It took many iterations to get physics to work correctly.

Such a good model. I'm so glad I went all in on local months ago. I was so tired of Claude making every excuse it could to try to force a new turn.

▲
55
-3
26👁
r/LocalLLaMA · u/nomorebuttsplz · 25d ago
Base-10's Charlie O'Neill on why Kimi and GLM are "almost objectively" better than Opus 5 post image

Edit: Spelled Baseten not Base-10

Full episode of this available at https://www.youtube.com/watch?v=PrSf7IOYu-I
It's interesting to see how Dwarkesh has had to come around to the evidence that we are well on our way to creating AGI and even RSI in the last few months, despite historically being very skeptical.

I highly recommend people interested in large language models check out this particular episode, because it dispels a lot of mythology about stuff like plateaus from lack of data etc. For those who thought we were hitting a wall a year ago, it turns out there was a ton of low hanging fruit and the researchers in this episode discuss what that fruit was. They also extrapolate these trends into the future.

It's funny this subreddit is becoming rather skeptical of AI progress, which to put diplomatically, I think is based on a lack of information and too much time on Reddit.

▲
54
-3
26👁
r/LocalLLaMA · u/WebAssemblyMan · 23d ago
Recurrent Looped Transformer post image

Recurrent Looped Transformer (RLT)passes the decoder's final hidden state to the next token, together with that token's causal encoder representation. The decoder reads encoder-derived global KV memory and maintains a sliding-window attention (SWA) cache at every layer. The same update runs over prompt and response tokens.

More effective reasoning depth!

https://github.com/yifanzhang-pro/recurrent-looped-tranformer

▲
536
-4
27👁
r/LocalLLaMA · u/T_rex2700 · 28d ago
Someone apparently managed to kind of replicate what V4.1 flash does on KV for fast prefill on Qwen post image

I wonder someone will figure out a way to do this with 27B?

Throw Qwen3 on this page for demo
https://kishida.github.io/webdemos/llkvapprox/

Edit: sources (thank you u/pmttyji for finding them!

▲
517
-4
33👁
r/LocalLLaMA · u/Balance- · 30d ago
Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)

It seems to use a 96-bit LPDDR5X memory bus, instead of the previous 64-bit wide busses. Considering it's on 2nm, that's expensive silicon. That should result in around 115 GB/s memory bandwidth.

A20 Pro also doubles the size of Apple's dedicated Neural Engine (from 16 to 32 cores total).

▲
432
-4
36👁
▲
273
-4
18👁
r/LocalLLaMA · u/Eyelbee · 25d ago
The new k2 horizon models seem like an absolute beast post image

Especially the 7B one seems very interesting, it casually destroys muse glimmer with a way smaller size. And they open source literally everything, every step of the way. Anyone tried that model? It can be a new milestone if 7b and 3.7b ones are actually good, and not just benchmaxed.

▲
219
-4
21👁
r/LocalLLaMA · u/Top_Power5877 · 29d ago
DeepSeek V4.1 Flash: Stronger, Faster, More Accessible

Original Source from DeepSeek WeChat Official Account: https://mp.weixin.qq.com/s/qg0NU3NNUbp1co2PdkAPAg

Today we're officially releasing the DeepSeek V4.1 Flash model. It is the smallest model in our brand-new model architecture series, with native multimodal visual understanding. The new architecture was designed with these goals in mind: a higher capability ceiling, faster inference, greater throughput, and scalability to larger-parameter models.

Asymmetric architecture: big intelligence at low cost

DeepSeek V4.1 Flash is a 552B-parameter MoE model built on a brand-new Causal-Encoder-Decoder architecture. Input and output are asymmetric: only 8B parameters are activated on the input side and 16B on the output side, making it significantly cheaper than known models of the same size. V4.1 Flash also uses a new pre-training approach and has gone through larger-scale reinforcement learning post-training. In benchmark testing, it surpasses the intelligence level of a range of flagship models, including DeepSeek V4 Pro.

https://preview.redd.it/qq5p9q5qymoh1.png?width=1080&format=png&auto=…

https://preview.redd.it/2tvpysuvymoh1.png?width=1080&format=png&auto=…

Less cache, lower cost

The new generation of models dramatically reduces the size of the KV cache. Compared with the previous generation, HBM requirements drop to 1/4 and SSD requirements to 1/8. In agent scenarios, cache-hit charges often make up a large share of the bill, so compressing the KV cache substantially lowers the cost of agent-style tasks.

Figure: DeepSeek's continued progress in reducing context storage. Relative to the first-generation model, the KV cache has shrunk 437×.

API support

DeepSeek V4.1 Flash is now live on the DeepSeek API with native multimodal support. Simply change the model name to deepseek-flash to call the latest V4.1 Flash. The older V4 Flash and V4 Flash Vision Exp models have been retired; for compatibility, the model names deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily be routed to V4.1 Flash.

In addition, extensive testing shows that V4.1 Flash comprehensively outperforms V4 Pro on performance, cost, speed, and total time-to-completion, so we plan to phase out the V4 Pro model in an orderly fashion. After 12:00 Beijing time on September 14, 2026, and until V4.1 Pro launches, all requests to deepseek-v4-pro will be routed to V4.1 Flash and billed at V4.1 Flash's unit price.

Tencent (WorkBuddy, CodeBuddy) and OpenCode, as official partners, have now fully integrated DeepSeek V4.1 Flash — give it a try!

API pricing adjustment

Thanks to the architectural innovations, DeepSeek V4.1 Flash can serve more users at lower cost, so we have cut V4.1 Flash's pricing accordingly. To allocate resources more sensibly, we continue to use peak/off-peak pricing, with off-peak prices at half the peak rate, and encourage users to schedule tasks around their actual usage patterns. The new prices take effect at 12:00 on September 10, 2026.

https://preview.redd.it/qkhui0kjzmoh1.png?width=1690&format=png&auto=…

Open-source release

We will fully support the open-source community in adapting inference for the new model, and will explore various ways to broaden deployment. If you have large-scale deployment needs and the corresponding resources (a 2k-GPU cluster with storage cluster), please get in touch.

▲
195
-4
16👁
▲
186
-4
24👁
r/LocalLLaMA · u/OvertaxedOne · 26d ago
The rhetoric is really heating up!

The entire page of the NY Times today above the fold absent one article is AI (the models are just too strong/too dangerous, must be regulated). They forgot to include "Sponsored by OpenAI" at the end of the articles, sure that was just an oversight?

This is what the end of a bubble looks like, desperate attempts to get some sort of regulatory capture in place to keep the business model from collapsing in upon itself. My days next week are 100% booked talking to companies about how to get off frontier models, one large, and a bunch of smaller customers, including one who's flying me out to them to sit down and get a plan in place immediately (the controversy around that math problem really spooked some CEO/CIO's about data privacy using cloud models).

Gonna be an interesting few weeks. Maybe the Qwen team will be nice enough to give me a little breathing room before dropping another hydrogen bomb? :)

▲
165
-4
26👁
r/LocalLLaMA · u/Ok_Warning2146 · 28d ago
Terminal Bench v4 scores post image

Some people says terminal bench reflects model intelligence better than the intelligent index. From the look of it, the ranking does seem to reflect how people feel about the open and closed models.

For the open models, GLM-5.3 is in a league of its own. GLM-5.3-Flash is leading the current gen of top flash models. Kimi-K3 did pretty bad in this benchmark for its size. Qwen3.8-27B is the only small model that can do something on this bench.

|Model|Score|
|:-|:-|
|GLM-5.3|41.9%|
|GLM-5.3-Flash|32.8%|
|DSV4.1-Flash|26.8%|
|Qwen3.8-Flash-Next|25.3%|
|DSV4-Pro|14.1%|
|Kimi-K3|12.6%|
|DSV4-Flash|12.1%|
|Qwen3.8-27B|5.6%|
|Muse Glimmer|0.5%|
|gemma4-31b|0.0%|

▲
119
-4
21👁
▲
100
-4
15👁
r/LocalLLaMA · u/No-Name-Person111 · 24d ago
Occamy-1.0 by Accio Lab
▲
94
-4
24👁
▲
90
-4
14👁
▲
90
-4
16👁
r/LocalLLaMA · u/EcstaticDentist · 26d ago
Decided to build a game, and test the ceiling of Qwen3.8 27b post image

This took roughly 5 hours to create, using 2 different configured harnesses, same model. RTX 3090, overclocked +12% gain (MSI Afterburner), Q4KM - built this for fun, will be throwing it on GitHub, opensource for people to get an idea of a project created to the near ceiling of performance & capability for q3.8 27b. & also maybe ya’ll can contribute to the game only iterating locally. It would be a fun little experiment.

▲
85
-4
24👁
r/LocalLLaMA · u/crusaderky · 25d ago
Animated transition from AA Intelligence Index v4.1 to v4.3 post image

I had all the data saved from AA's v4.1 index, so when they upgraded it in the wake of Astra's release, I could actually generate a before/after comparison.

  • All intelligence and price per task are sampled from AA on Sep 3rd and Sep 14th respectively.
  • Price per task of some open models were rescaled to reflect the cheapest available on OpenRouter as of Sep 3rd.
  • X axis is linear, because people's money is linear.

All models are the same. The only thing that changes is the weighted sum of the benchmarks that compose the Intelligence Index.

v4.1: https://github.com/crusaderky/llm-intelligence-cost-plot/blob/intelligence-index-v4.1/plots/high\_intelligence.png

v4.3: https://github.com/crusaderky/llm-intelligence-cost-plot/blob/intelligence-index-v4.3/plots/high\_intelligence.png

Highlights

  • GLM an Muse Spark remain more or less unaltered, in relative terms
  • GPT-5.6 Sol becomes a lot cheaper
  • GPT-6 Astra's intelligence flies up to the stars AND becomes cheaper
  • GPT-5.6 Luna gets a substantial uplift
  • Fable-5.1's price gap from Opus 5 shrinks, and becomes cheaper than Fable 5.0
  • Fable-5.1 at low, medium and high effort looks a lot more appealing
  • Sonnet 5 becomes even more expensive without any intelligence gains
  • Kimi-K3, Qwen3.8-Max, Gemini-3.8, and Grok 4.6 go down into the gutter
▲
65
-4
28👁
r/LocalLLaMA · u/mentria-ai · 30d ago
1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install) post image

mentria.ai is a browser inference engine I've been building solo, from scratch in WebGPU/WGSL. This week it crossed a milestone I had been chasing for a while: a 27B one-bit model answering at up to 30 tokens/s on an RTX 3060 Laptop GPU with 6 GB of VRAM, in Chrome, from a web page. No install, no server, nothing leaves the machine.

The model. Bonsai-27B, a natively 1-bit model trained and released by Prism ML. I repacked it for our engine, wrote the kernels that run it, and checked its quality myself (full eval deltas vs FP16 Qwen3.6-27B are in the HF card). Every weight is one sign bit with one scale per 128 weights, about 1.14 bits per parameter, so 27 billion parameters sit in 3.8 GB of GPU memory. Repack and numbers: huggingface.co/mentriaai/Bonsai-27B-mentria

The engineering that got it to 30. Two days before this post the same model decoded at 15 tok/s on this laptop. Decode is memory-bound: one word of the 27B is 804 GPU dispatches, and 401 of them are the 1-bit matrix-by-vector kernel that streams the model's 3.6 GB of matmul weights once per word (embeddings bring the file to 3.8 GB). The kernel that won was the one written for phones: four 1-bit weights have only 16 possible partial answers, so it computes all 16 once into on-chip scratch and each row reads its answer from that table instead of multiplying (station 258 — the numbered write-ups on the facts page). On Ampere that table hit a wall that was not the weights at all: scratch memory has 32 banks, the table's addressing sent sixteen threads of every warp to the same bank, and the kernel ran at 38% of the card's bandwidth. One padding slot per row spread the addresses across the banks; the kernel's share of a word dropped from 26.5 ms to about 21, and on top of the 15 → 26 tok/s the table itself had bought, the laptop hit 32 tok/s raw — the 25–30 is what survives sampling and UI overhead in the chat (s259). The prompt stage had the same disease in its tile layout; a retile took a 1,489-token prompt from 29.6 s to 25.3 s. Every change shipped only after its output was byte-identical to the build before it, which is also why the kernel adds its partial sums in a fixed order and accepts a 44% occupancy ceiling.

Every claim here has a numbered write-up on the engine facts page.

Numbers (RTX 3060 Laptop 6 GB, Windows 11, Chrome 152 on D3D12, fresh loads):

  • Decode: 25–30 tok/s in the chat UI once the card is warm.
  • Prompt processing: a 1,489-token prompt in about 25 s.
  • Context: 3,072 tokens on this 6 GB card; 8,192 on 16 GB Macs; more on bigger cards, at 128 KiB per token. The KV cache is exact math, no quantized cache. The next step on 6 GB is consolidating the engine's few thousand small GPU buffers into a handful of large arenas, so the driver stops holding about 300 MiB of slab slack; that is the arithmetic for 4,096, and it is not built yet.
  • Load: under 10 s from the browser cache; the first download is 3.8 GB, once.

Also in the engine: the smaller tiers (Qwen3.5 0.8B, 2B, 4B) cover phones and low end devices, LoRA adapters of a few MB hot-swap at the matmul in under a second, and there is a vision tower for image input.

Try it at https://mentria.ai/tools/ai-chat/ (the 27B tier appears when your GPU qualifies). The code that ships, the benchmarks and how I measure are in the repo: https://github.com/mentria-ai/website. The engine facts page, for how the engine and the model actually work: https://mentria.ai/assets/learn/engine-facts.html

▲
57
-4
23👁
▲
938
-5
35👁
r/LocalLLaMA · u/Secure_Recording_472 · 25d ago
UkisAI Swift-Qwen3.8-27B / -58.3% thinking, x1.95 speed while keeping the accuracy of xhigh post image

Hi everybody, we post-trained Qwen 3.8 27B to be more efficient by figuring out which tokens were linked to overthinking and penalizing them without "attacking" the reasoning length directly then fixed the accuracy with a bit of secret sauce (hint On-Policy Distillation) and achieved great results (-58% thinking tokens, 1.95x speed up, <1% accuracy loss) so we wanted to open-source it and hear the feedback of the community.

This is the link to the model: https://huggingface.co/ukisai/Swift-Qwen3.8-27b

We also also providing a Free Research Purpose API (OpenAI compatible), courtesy of Nvidia who were kind enough to provide us with the GPUs. You can use it to try out the model if you do not have enough compute to run it, it's limited at 5RPM. https://ukisai.com/api/swift/v1/models

We also made a GGUF (Q1-Q8) and there's also a few nice community (Bartowski) quants with even lower/higher precision. The community also created amazing NVFP4, W4A16 and Uncensored versions of the model you can find on Huggingface.

IMPORTANT: Our training approach is not a replacement for the reasoning effort settings, chat templates or token caps but is complementary and targets a completely separate issue (overthinking and "anxiety-like" reasoning loops prior seen in PTQ, but as far as we identified also prominent in BF16 of this size class LLMs as well). Contrary to popular belief, these specific patterns do not contribute to answer quality when properly targeted. (our thesis being: reasoning length IS extremely important and should NOT be shortened by force, but rather optimized). This is also demonstrated bellow in our xhigh vs medium effort benchmark table. The goal is to keep xhigh accuracy while reducing only the unnecessary part of thinking.

I will TLDR you on our thought process, research, training and benchmarks.

  1. When running our quantized Qwen 3.8 27B instances we were very annoyed by random reasoning loops (in the paper bellow refered to as "overthinking errors". These random loops were persistent throughout medium and low reasoning settings.
  2. We remembered a paper by Meta that's supposed to target this phenomenon in PTQ, but when used straight out of the box got mixed results.
  3. We figured to try if it's a matter of the targeting the right keywords and tuning the parameters, so we used our 8xH100 box and and generated a large amount of different (ofc out of distribution) domain (coding, language, vision, agentic) traces.
  4. We then grouped the ones with overthinking and found "common denominator" tokens between them and targeted the most prominent ones.
  5. We then built an inference-time penalizer of those tokens as seen in the paper with the hopes of simply generating traces and doing cross-entropy SFT over them.
  6. Did not work at all, but the penalizer seemed to work much better than the tokens provided in the paper and not only for lower precision models but for bf16 as well. Hence we kept experimenting with it. We built a loss function using the tokens we identified and ran LoRa SFT over the traces prev generated and reasoning seemed to be falling off significantly but the accuracy seemed to follow. The reasoning reduction seemed to be generalizing.
  7. After a significant amount of tinkering (literally since the day of Qwen 3.8 27B release) we were satisfied with the reasoning reduction. After that we searched for ways of restoring the accuracy. We experimented with several methods, including RL(GSPO), On-Policy Distillation and using the ThinkingCap 3.6 27B adapter chunks until we were satisfied with our accuracy loss. We managed to restore it to <1% loss on almost all of our OOD in house tests
  8. We then performed intensive intensive benchmarks, across several reasoning efforts, precision variants etc. We ran into a few problems, one of which is that to get a reliable score we needed to run each benchmark 10x (5x on base + 5x with our adapter, this being the standard procedure on the Qwen 3.6 27B model card on Terminal Bench which we followed). After running it, the performance converged to 40-60% token reduction with <1% accuracy loss across GPQA, MMLU, Terminal Bench 2.1, LiveCodeBench v6, ERQA, C-Eval, IFBench, HMMT25, with an exception being AIME26 with an accuracy loss of 4.6%, which we later linked to a bug during training with a specific token relevant for math-related reasoning being penalized and are planning to fix it in an updated release.

The benchmarks: (raw benchmark files here - https://github.com/UkisAI/Swift-Qwen3.8-27B-evals/ )**

Swift-27B vs Qwen3.8-27B (BF16, all benchmarks ran x5, thinking effort xhigh)

|Benchmark|Qwen3.8-27B|Swift-27B|Median tokens|
|:-|:-|:-|:-|
|GPQA-Diamond|88.4%|88.3%|58% fewer|
|LiveCodeBench v6|76.8%|81.6% (+4.8pp, due to default truncation in LCB it is not performance gain)|46% fewer thinking tokens|
|Terminal-Bench 2.1|66.7%|65.8%|39% fewer|
|MMLU-Pro|85.5%|85.0%|28% fewer|
|C-Eval|90.0%|90.6%|19% fewer|
|IFBench|73.5%|71.8%|51% fewer|
|AIME 2026|98.7%|94.0%|50% fewer|
|HMMT (Nov 2025)|99.3%|96.0%|46% fewer|
|ERQA (vision)|67.5%|66.3%|55% fewer|

Token savings hold at every reasoning effort (mean thinking reduction): xhigh 41%, medium 23%, low 26% (albeit with accuracy loses of 1-4% on medium and 1-2% on low which we need further testing for)

Swift at xhigh vs the base's own effort settings on GPQA-Diamond (198 questions x 5 seeds):

|Model / effort|Accuracy|Median tokens|
|:-|:-|:-|
|Base xhigh|88.4%|6,642|
|Swift xhigh|88.3%|2,771|
|Base medium|84.1%|1,753|

So Swift keeps xhigh accuracy at under half the tokens, and beats base-medium by 4pp at roughly 1.6x its tokens.

End note:

While we are keen on complete open-source, we still need to keep a part of our training and data private, being a new lab. The license is not Apache 2.0, but it only affects companies >$1M. We hope this does not pose a problem for the community, but we are open to feedback on it.

We want to contribute as much as possible to the community and would really appreciate feedback on our work, quantization or Swift model requests. For context, we are working on Swift 3.8 Flash Next right now and have so far gotten up to -30% thinking token usage while maintaining xhigh accuracy, which we take as a strong indicator our methodology is reproducible across the Qwen model family. Will explore other families as soon as we have the capacity and would love to see which ones the community would love for us to optimize first.

▲
422
-5
25👁
r/LocalLLaMA · u/Porespellar · 23d ago
Frontier LLM development simplified for politicians: post image

Nobody is buying this “Pace the frontier” nonsense. It makes no logical sense at all. Are American labs really going to take a pause and lose any small lead they still may have over Chinese labs? Does anyone really believe this? This seems like some performative virtue signaling BS. Why are they bothering with this pacing campaign? Someone please explain.

▲
312
-5
28👁
r/LocalLLaMA · u/DistanceSolar1449 · 29d ago
Deepseek V4.1 Flash is 748B, not 552B

People keep on getting confused about this, so I looked at the safetensors on hf.

The title should have been "Deepseek V4.1 Flash is 748B total/552B base, not 284B or 305B or 485B or 522B"

  • The model is not 284B. The original Deepseek V4 Flash is 284B, but not the V4.1 Flash model
  • The model is not 305B, despite what some people claim "So: ~305B real backbone + 203B engram = 508B total" This is incorrect.
  • The model is not 485B, even though Huggingface lists the model as 485B, but that's because they're counting some FP4 packed weights as bytes instead of params (2 FP4 params per byte). This happens a lot; for example Huggingface incorrectly thinks GLM-5.3-flash is 169b here
  • The model is not 522B, even though VLLM lists it as 522B for some weird reason. They correct themselves later down the page (ctrl-f "Params" on that vllm page)
  • 552B is the only number out of this list that's somewhat correct; that only includes the base model without MTP and engrams and the vision encoder though.

To be precise, the main model about 551.566B parameters with 40 layers. The FFN experts total to 543.582B parameters, and the rest of the model (attention, shared experts, etc) are 7.984B.

On top of that, the engram is \~196.929B, DSpark/MTP is \~14.225B, and the vision encoder is just \~0.485B. These parts are technically optional though. The vision encoder is also way smaller than I expected.

Anyways, you need a beefy system for this. 128GB or 256GB of RAM/VRAM is not going to cut it.

|Component|Logical params|Size in GB|Storage|
|:-|:-|:-|:-|
|FFN MoE experts|543.582B|288.778 GB|FP4|
|Other FFN|1.4947B|1.574 GB|FP8 mostly|
|Attention|5.1269B|6.524 GB|FP8 mostly|
|Embedding + LM head|1.3238B|2.648 GB|BF16|
|Other|0.0397B|0.158 GB|FP32/BF16|
|Backbone total|551.566B ≈ 552B|299.682 GB||
|Engram lookup tables|196.614B|202.758 GB|FP8|
|Engram projections/gating|0.315B|0.315 GB|FP8 mostly|
|Engram total|196.929B = 196B advertised|203.073 GB||
|DSpark / MTP|14.225B|8.033 GB|mostly FP4 experts|
|Vision encoder|0.485B|0.971 GB|BF16 mostly|
|Everything in total|\~763.21B params|\~511.76 GB||