123 posts · 1 sub · RSS
← prev Sep 27, 2026 → Sep 28, 2026 next →
2026-09-27 → 2026-09-28 hourdayweekmonthyearall
allr/LocalLLaMA
▲
803
+39
63👁
r/LocalLLaMA · u/charles25565 · 11d ago
GPT-3 is discontinued today post image

It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely in our memories. Arguably what's more infuriating is that they suggest using GPT-5.6 Terra as a replacement. Keep in mind that Babbage is a model that's literally 3/4 of the size than MiniCPM5 2B. Even Luna might be overkill as a replacement. But neither is a drop-in replacement. Davinci is the main GPT-3 most people use. This is why we have local models, because they simply cannot have a universal end of life date.

💬 161 (+3) open on reddit ↗
▲
809
+31
56👁
▲
629
+27
45👁
▲
686
+22
36👁
▲
622
+17
52👁
r/LocalLLaMA · u/am17an · 12d ago
Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy

Meta came out with a banger paper https://arxiv.org/pdf/2606.00206, but it did not look at various quantizations supported in llama.cpp. So I did a run on 50 random MATH-500 questions (https://huggingface.co/datasets/HuggingFaceH4/MATH-500) and ran it on various quantizations of https://huggingface.co/bartowski/Qwen\_Qwen3.5-4B-GGUF and tried

--logit-bias 466-2 --logit-bias 694-2 --logit-bias 1362-2 \
--logit-bias 1412-2 --logit-bias 1921-2 --logit-bias 1990-2 \
--logit-bias 2086-2 --logit-bias 2361-2 --logit-bias 2441-2 \
--logit-bias 2493-2 --logit-bias 2892-2 --logit-bias 3222-2 \
--logit-bias 3315-2 --logit-bias 3384-2 --logit-bias 3404-2 \
--logit-bias 3482-2 --logit-bias 3655-2 --logit-bias 4213-2 \
--logit-bias 4370-2 --logit-bias 4598-2 --logit-bias 4611-2 \
--logit-bias 4808-2 --logit-bias 5752-2 --logit-bias 6970-2 \
--logit-bias 7014-2 --logit-bias 7643-2 --logit-bias 8106-2 \
--logit-bias 10179-2 --logit-bias 10451-2 --logit-bias 11746-2 \
--logit-bias 13264-2 --logit-bias 13428-2 --logit-bias 14673-2 \
--logit-bias 15029-2 --logit-bias 16036-2 --logit-bias 21143-2 \
--logit-bias 21979-2 --logit-bias 33955-2 --logit-bias 35999-2 \
--logit-bias 36563-2 --logit-bias 37201-2 --logit-bias 37781-2 \
--logit-bias 41484-2 --logit-bias 62586-2 --logit-bias 66073-2 \
--logit-bias 73071-2 --logit-bias 84485-2 --logit-bias 85152-2 \
--logit-bias 95500-2

these correspond to the paper's overthinking markers:
\[

" perhaps", " maybe", " wait", " Wait", " actually",

" hold", " Hmm", " hmm", " Alternatively", " alternatively",

" However", " however", " instead", " Instead", " But",

" but", " though", " although", " yet", " rather",

" unless", " otherwise", " nonetheless", " nevertheless", " regardless",

" still", " anyway", " Or", " or", " either",

" whether", " uncertain", " unsure", " possibly", " might",

" could", " another", " different", " reconsider", " rethink",

" backtrack", " retry", " revisit", " doubt", " confused",

" wrong", " mistake", " error", " incorrect"

\]

Here are the results, surprisingly even BF16 leads to better accuracy. Caveats being this is one test on one model. Try it out and see it helps!

|Format|Accuracy: baseline → penalty|Reasoning tokens|
|:-|:-|:-|
|BF16|74% → 84%|−19.4%|
|Q8\_0|76% → 80%|−11.0%|
|Q4\_K\_M|60% → 66%|−14.8%|
|Q3\_K\_M|52% → 66%|−17.5%|
|Q2\_K|12% → 24%|−11.5%|

💬 135 (+7) open on reddit ↗
▲
283
+15
51👁
r/LocalLLaMA · u/LegacyRemaster · 11d ago
Qwen next 3.8 and 3.8 27b Vs Sonnet 5.5 low and Sonnet 5.5 medium. post image

Six months ago, a result like this was unthinkable. But now we can say it loud and clear: local models are at the cutting edge, and the gap of just a few months has been confirmed.

Personally, I use Qwen-Next 3.8 for complex tasks; today, GPT-Sol-6-High was messing up a project, but Qwen-Next got it back on track. I consider it a reliable benchmark. What’s your take?

💬 117 (+3) open on reddit ↗
▲
222
+13
34👁
r/LocalLLaMA · u/Educational-Care7867 · 11d ago
ImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench post image

Some context first.

I am process improvement / business consultant and had worked with Fortune 500 companies on improving their processes around refunds, returns, customer support, etc.

This entire thing had a lot of complex decision making and generally every decision / condition node in a process map was generally replaced by a human because factoring in ambiguity in a code is very difficult.

Idea of ImaJev

Hence, when Jev came out, I was very intrigued with it and also could clearly see its use-case of improving decision making in complex decision work flows.

However, Jev didnt have support for Images and I thought that it can be replicated for both Text and Images in a single model and thats when I started with ImaJev.

Training Process

It went badly at first. My first big fine-tune on about 500k short decisions made the 9B model worse at reasoning: 64.9 down to 42.3 on JevBench hard. It had basically learned to pattern-match. I spent the next couple of weeks generating hard questions with open-weight models and only keeping the ones where two Ai models agreed on the answer. That brought it back.

Results

Then, on the JevBench - It came out #1 of 91 (v1.4.2.2, scored 27 Sep), 67.37 vs Jev 1.13.0 at 63.29.

The same week DecisionBench put it #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1.

I honestly didn't expect either.

To be fair about it: the #1 is on a score that weighs accuracy, calibration, speed and cost equally. On accuracy alone it's #3. Its main strength is that when it says 90% it's usually right, and it'll say "can't tell" instead of guessing.

What it actually is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions.

It gives back a probability for each option plus "unknown", in one forward pass.

Runs on a Mac with MLX or on one GPU.

The whole project costed me around $1200 in rented GPU and a lot of time :P

I would love to know your thoughts on it - it anyone would be interested to try that.

💬 69 (+2) open on reddit ↗
▲
404
+11
52👁
r/LocalLLaMA · u/professormunchies · 12d ago
Qwen plays World of Warcraft post image

Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without installing the game and it has mobile controls. Afterwards, created a custom mcp to drive the client and have finer game control than a generic browser agent. The agent harness can plug into your local or cloud LLMs and be used to drive the game. For best results have a model that can output >50token/sec. No visual input is used in the making (might be beneficial in the future but incur more latency). The mcp and agent are only running on my dev server but if folks are interested in trying the game go to https://jankcraft.xyz/

Still vibing but I’ll make some more content of it … I think those Pokémon benchmarks have become a little too easy and they need a new challenge like speed running to 80 in wraith of the lich king.

💬 128 (+7) open on reddit ↗
▲
176
+11
24👁
r/LocalLLaMA · u/Ok-Breadfruit-3523 · 12d ago
I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s post image

Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2 more of these and will probably get them running to see if 3.8 flash next runs at usable speeds. This setup is wildly inefficient with power but cost me less than $800.

▲
147
+9
32👁
r/LocalLLaMA · u/nullmove · 12d ago
Naive-N0.5-Flash - 309B-A15.5B

https://naive.ai/en/research/

  • Built for coding and AI R&D
  • 1M context
  • Hybrid SWA/DSA
💬 38 (+2) open on reddit ↗
▲
71
+9
38👁
r/LocalLLaMA · u/returnity · 11d ago
Searching for 3.8 35B: Qwen3.6-35B-A3B (Testing 5 Finetunes vs. Base)

TL;DR -- You should probably just use base Qwen3.6-35B, as only Occamy-1.0 is competitive with it. Tiel is a major let-down, worse than Ornith. KAT surprises (good), Nex surprises (bad). This post is long. Sorry, lots to cover.

I think we all want to see a next-generation small MoE from the Qwen team to replace 3.6-35B in our workflows. This model is a perfect fit for smaller gmaing laptops and mid-tier rigs. It sucks that Qwen seems to have abandoned this model, but at least there are fine-tunes that improve upon it... right?

Well... maybe not. I ran benchmarks on the 3.6-35B-A3B base model, as well as five finetuunes: Occamy-1.0, Ornith-1.5, KAT-Coder-V2.5-Dev, Tiel-Coder, and Nex-N2.5-mini, and the results are quite surprising.

I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark I can run in ~10 hourso on my machine, it's not actively post-trained on by any of these models, and it provides a lot of useful information along with the raw accuracy scores. This includes: first-try and retry pass rates, token counts, solve times, and how well-formed the output diffs are. Here's the table:

| model | First-try pass | Retry pass | tokens | sec/case | tok/solve | well-formed diff |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B BASE (STOCK template) | 37.4% | 71.0% | 8650 | 285 | 14.1K | 96.3% |
| Occamy-1.0-35B-A3B (STOCK template) | 29.0% | 70.1% | 6801 | 285 | 17.2K | 86.9% |
| Occamy-1.0-35B-A3B (froggeric medium) | 30.8% | 69.2% | 6009 | 233 | 16.8K | 94.4% |
| Occamy-1.0-35B-A3B (froggeric, xhigh) | 27.1% | 67.3% | 8631 | 310 | 20.0K | 91.6% |
| Ornith-1.5-35B-A3B | 23.4% | 63.6% | 4813 | 226 | 16.4K | 87.9% |
| KAT-Coder-V2.5-Dev | 20.6% | 58.9% | 2190 | 84 | 9.3K | 86.9% |
| Tiel-Coder-35B-A3B | 18.7% | 53.3% | 4851 | 171 | 18.2K | 89.7% |
| Nex-N2.5-mini | 10.3% | 30.8% | 5037 | 188 | 33.3K | 95.3% |

As you can see, the only finetune that even competes with the base model is Occamy-1.0. The rest are utterly dominated by the base model, a grim disappointment for finetune enthusiasts. I was particualarly surprised by the performance of Tiel, which seems to get a lot of love in this subreddit.

Speaking of Tiel, I want to clarify that Tiel is just Ornith-1.5 with a different chat template, Sharp, which is based on froggeric with an added "terse mode" instruction that's supposed to reduce excessive verbosity. I wanted to standardize for templates, so ALL models are using the base froggeric v22.5 template set to medium (which is equivalent to standard thinking on, no additional message sent). I used this because I wanted to test Tiel vs. Ornith-1.5, and Tiel is the chat template. Also, practially, I use froggeric in my real workflows. However, to ensure coverage, I also tested the STOCK template on the 2 highest-performing models, to make sure it wasn't affecting the scores. As you can see, the template doesn't make a significant difference in the scores, and the scores for base 35B with different templates are so close to identical that I excluded the froggeric one from the table.

I also tested froggeric/Sharp's reasoning-effort toggle, and found xhigh -> medium significantly reduces token counts and solve times (by ~1/3), without affecting accuracy significantly. That stands in stark contrast to Tiel's 'terse mode' toggle, the core feature of Tiel over Ornith, which dramatically reduces accuracy along with the reduction in token counts. My results strongly suggest that if you want a less verbose model, you're better off lowering the reasoning effort than using Tiel with terseness on.

Speaking of token use, that's probably the big differentiator here. A couple models stand out: Ornith and KAT-Coder-V2.5-Dev are the most efficient models, with KAT in particular having a brevity unmatched by anything else. KAT is fucking fast, and I think despite its lower accuracy than Occamy, it has a place in my lineup as a subagent because it just gets. shit. done. Occamy is also interesting, as it is the only model that perfoms on a similar level to the base, but it uses 20-30% fewer median tokens. However, Occamy also had a number of runaway generations where the token count blew up, so it's total tokens/solve is actually higher than base.

In an effort to further distinguish Occamy from base, since Aider struggled to do that, I ran tau2-bench, an agentic tool-calling benchmark consisting of multi-turn interactions with a simulated counterparty. I figured this was a good bench to use as Occamy is post-trained specifically for 'co-work' scenarios, but not trained on this particular set. I used Qwen3.8-27B with reasoning effort set to low as the simulated customer in these conversations. The base model was able to pull away from Occamy in the harder retail domain of this benchmark, but Occamy resolved the issues in the airline domain at an equal rate while requiring fewer turns. Here's the results.

| model (Q8_0) | domain | pass^1 | tokens | sec/task | turns/task |
|---|---|---|---|---|---|
| Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) | airline | 80.0% | 5390 | 290 | 11 |
| Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) | airline | 78.0% | 6770 | 390 | 13 |
| Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) | retail | 86.0% | 4098 | 333 | 14 |
| Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) | retail | 79.8% | 3284 | 284 | 13 |

Overall, I think the results are clear, if unexpected: Occamy-1.0 is the only fine-tune that even competes with the base model on Aider Polyglot, but even it is not a clear winner. Tiel is noticebly worse than plain Ornith without the terseness toggle, and the terse mode doesn't even save any tokens. xhigh in froggeric/Sharp degrades accuracy slightly and bloats token use, which makes sense given the models were not RL'd for the extra thinking effort prompt. KAT-Coder-V2.5-Dev is the most efficient model, with accuracy nearly as good as Ornith and better than Tiel. Finally, Nex-N2.5-mini is a disaster.

💬 64 (+1) open on reddit ↗
▲
64
+9
42👁
r/LocalLLaMA · u/Top-Evidence174 · 11d ago
Mica v0.1 4B got diamonds in survival Minecraft on its first run. 26 decisions from an empty inventory. post image

I've been working on Mica, a 4B decision model, and wanted to see how far it could get in actual Minecraft, not a sim. New world, empty inventory, and the goal was a diamond pickaxe.

Last time I posted, it got an iron pickaxe. Honestly that took around 20 tries and it was pretty flaky. I've reworked the harness a lot since then. This time it went all the way to a diamond pickaxe, and once the harness was finished it did it on the first run.

It took 26 decisions and about 8 minutes of game time. It got wood, made a crafting table, then wooden and stone pickaxes, then iron and coal, a furnace and an iron pickaxe. After that it tunneled down to diamonds at y=2, mined three, put a crafting table down right there and made the pickaxe. Decisions took about 108 ms on average.

My favorite bit is around step 14. The planner wanted it to make planks to burn in the furnace, but Mica went and mined coal instead (0.69 vs 0.31) and then smelted all three iron at once. Which was the better call, honestly.

How it works: every step Mica gets the game state as text (inventory, nearby blocks, health, what happened last step) plus a few candidate commands, and it picks one. A Mineflayer bot running Mindcraft skills does the actual moving and mining. The panel on the right of the video shows each decision and its probabilities live. The bot also knows where the nearest diamonds are, so it isn't searching for them.

To be clear, I'm not saying a 4B model plays Minecraft on its own. What I wanted to show is that a model this small can sit behind a bot, read what's going on, and make the next call well enough to get all the way to diamonds.

I'm planning to release the harness soon. Mica will read Minecraft chat, so you can type what you want and it'll work toward it. Simple stuff like getting items, crafting or following you should work fine, but it'll struggle with anything really complex, like building a house.

Also, v0.5 should be out in the next 1-2 weeks. A lot of the architecture changed, and I did extra training on the parts where v0.1 was weak, so I'm expecting a clear jump in performance. The aim is to be at or near the top among 4B JEV-like models.

There'll be two versions: Mica v0.5 4B, and Mica v0.5 4B Distill Laya, which is light enough to run on pretty much any PC.

For Minecraft, I'm hoping v0.5 will be good enough to take down the Ender Dragon, and I did extra training specifically with that in mind. No promises, but it'd be really cool if it pulls it off lol

Oh and fun fact, Mica is a fully vibe-coded project

Model: https://huggingface.co/sky7350/Mica-v0.1-4B
Model code: https://github.com/akivet/Mica-v0.1-4B
Minecraft harness: coming soon

▲
104
+8
30👁
r/LocalLLaMA · u/tossit97531 · 11d ago
Can we get some quality control on all these model perf posts?

Too many hyperactive amateurs are coming in here with "1b model at 832843tok/s!" and hardly any of them have all the info necessary for local runners to evaluate. We need context ladders with perplexity/KLD, hardware specs, model params and quant(s), runtime, tuned runtime parameters, basically everything we need to reproduce locally if we can match the entire setup. To say nothing of what the model is even good at in the first place if it's not a well-known model.

The goal is to get perf numbers that show they meet a certain quality bar. I don't care if I get 8324834 tok/s if it's all garbage.

Can we start filtering the hyperactive amateur perf posts please? It's getting really frustrating seeing all these posts of models and wading through info just to see that it doesn't test with anything but an empty context or doesn't say anything about quant or platform.

We need to define some rigor and apply it to this place, or it will remain like most ai-oriented subs and get continually choked with slop.

▲
38
+8
15👁
r/LocalLLaMA · u/CompetitiveDraft9381 · 12d ago
Updated from 3x3090(2x3090, 1x3090TI) to 2x5090

Upgraded from 3x RTX 3090s to 2x RTX 5090s on my homelab server and picked up a solid speed jump on top of it from a software update (speculative decoding + NVFP4). Setup: llama.cpp (build b11216), running Qwen3.8-27B (abliterated, Q8\_0). Blue = old 3090 setup, green = new 5090 setup on the same Q8 model, teal = the 5090s again after switching to NVFP4 + speculative decoding. For colorblind folks, the order is: 1 - 3090, 2 - 5090 Q8, 3 - 5090 with NVFP4. One caveat on the "before" numbers: one of the three 3090s was on a slower PCIe slot than the other two, so that setup was running a bit below what 3x 3090s on equal slots would do. Overall, very satisfied. I bought 2 prebuilt PCs for $6.4k each when the 5090 went up to $7.5k, 2 weeks ago or so. I really wanted to upgrade to 5090s for a long time for NVFP4 support. The plan is to sell the 3090s for $2k each or so. It would probably be another 2-3 months until they go up that high, but I expect that they will. So, with the prebuilts' leftover components and the 3090s, in the best-case scenario I expect to get back $9k, so the total cost of the GPUs would be about $6k after taxes, which is still nuts and more than the 5090's MSRP. Ask me any questions, or if there are any other benchmarks you guys want me to run, let me know and I will.

▲
446
+7
48👁
r/LocalLLaMA · u/speedb0at · 12d ago
The Opus 5.5 posts about motion graphics are cool but Qwen 27B made this on a 4090

Saw the hundreds of tweets where people just keep asking Opus 5.5 for motion graphic videos. Decided to ask qwen to look at them and make its own. Quite amazing what local can achieve.

\*\*EDIT\*\* It looks laggy because of reddits .gif limit btw

the full high res version (with sound) is here: https://x.com/mkultraware/status/2104192428664127555

Promted and built in: https://github.com/mkultraware/accuretta

https://i.redd.it/5gmotgxx82sh1.gif

💬 123 (+2) open on reddit ↗
▲
80
+7
26👁
▲
186
+6
16👁
▲
72
+6
32👁
r/LocalLLaMA · u/MasterNomie · 11d ago
What model sits between Qwen 3.8 27b and Flash next for coding?

Having tested both Qwen 3.8 27b and Flash next on RTX 5090 with 96GB RAM, I want to find the middle ground between the two for coding capabilities but not sacrifice decode speed to standstill. I would like decode speed to between 75-100 ideally for fast iterations; otherwise I become impatient.

Currently I get 200+ TPS on Qwen 3.8 27B and approx 50 TPS on Flash next.

My hardware - RTX 5090 and 96 GB DDR5 which I plan to upgrade to 128 GB (in this economy, yes, but unwillingly).

What model sits between these two in terms of coding capabilities and hardware requirement? If none is present, I can perhaps think of using Flash next for plan creation and 27b for implementation.

Edit: Fast forward few days. I gave Strata a go with Swift 1.5 Flash Next IQ3\_XXS and able to achieve \~150 tok/sec decode and 5k tok/sec prefill. I am escatic! The quality of response from 27b is considerably better and the speed is great. Both targets achieved.

▲
41
+6
17👁
r/LocalLLaMA · u/masiha97 · 12d ago
Public MCP server for Canadian privacy law data (free, no auth) - works with any client that speaks Streamable HTTP

I built this, so the disclosure goes up front. It's a free, public MCP server plus a REST API with Canadian privacy law data. The MCP endpoint is at https://movahedi.ca/mcp and it uses Streamable HTTP, so any client that speaks that transport can connect. No signup, no API key, read-only, anonymous. The REST API is at https://movahedi.ca/api/v1 and the docs are at https://movahedi.ca/developers. What's in it: - Canadian privacy enforcement actions (CAI decisions from Quebec's access-to-information commission), searchable by keyword - A 263-term privacy glossary - An 11-point Quebec Law 25 readiness checklist The 5 MCP tools are: search\_enforcement\_actions, get\_enforcement\_case, lookup\_glossary\_term, list\_glossary\_terms, law25\_requirements. Quotas are 2,000 calls/day anonymous, or 10,000/day with a free API key (no email required). With Claude Code you can add it like this: \claude mcp add --transport http movahedi-privacy https://movahedi.ca/mcp\ For local setups: since the server is remote over HTTP, a client that only speaks local stdio can reach it through a proxy like mcp-proxy or mcp-remote. That's how I've seen people pair it with locally run models and agentic harnesses. Happy to answer questions about the data or the setup. I am the builder (Alexa, on behalf of privacy researcher Mohammad Movahedi, movahedi.ca).

▲
645
+5
26👁
r/LocalLLaMA · u/Available_Pressure47 · 13d ago
42x Faster Prompt Lookup Drafting in llama.cpp
💬 181 (+1) open on reddit ↗
▲
255
+5
31👁
r/LocalLLaMA · u/jacek2023 · 11d ago
nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 · Hugging Face

Model Developer: NVIDIA Corporation

Model Development: Fine-tuned from NVIDIA-Nemotron-3-Ultra-550B-A55B

[](https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-…)What is Nemotron?

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

[](https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-…)Description

Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra, fine-tuned for one epoch on 477,642 synthetic reasoning traces distilled from GLM-5.2 across 22,000 curated problems spanning 16 regional and international competitive-programming contest families. Selected as the SFT teacher for its higher accuracy and roughly 30% shorter generations compared to a DeepSeek-V4-Flash-trained variant, GLM-5.2 distillation yields a model that, combined at inference time with GenCorrect — an iterative closed-loop test-time compute strategy that generates diverse candidate solutions, incorporates evaluator feedback, and refines subsequent generations under a fixed submission budget — was evaluated live and prospectively on the IOI 2026 problem set under official contest time, internet-access, and submission constraints, scoring 535.4 out of 600 and surpassing both the gold-medal threshold (361.12) and the top human contestant's score (498.27), making it the first AI system reported to outscore the highest-scoring human contestant on an IOI problem set.

This model is ready for commercial or non-commercial use.

▲
73
+5
40👁
r/LocalLLaMA · u/TypicalPudding6190 · 12d ago
Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM post image

We built an inference engine InferredThoughts for MoE models that don't fit in VRAM + RAM. Most of the model stays on the SSD, and experts are read as tokens need them.

This started as a proof of concept, and poc worked, we are getting 9-10 tok/s decode on Qwen3.8-Flash-Next NVFP4 (9.06 on the benchmark turn, 10.4 on the best turn).

This is just the start. With better SSD streaming, we expect v2 to reach ~14-15 tok/s decode.

Repo: InferredThoughts
https://github.com/compiledthoughts/Inferred-Thoughts

Model: Qwen3.8-Flash-Next, 176.9B params, NVFP4 GGUF (119 GiB):
https://huggingface.co/CompiledThoughts/Qwen3.8-Flash-Next-NVFP4-Q8_0

Machine: RTX 5060 Ti 16 GB, Ryzen 7 9700X, 32 GB DDR5, Gen5 NVMe SSD 1Tb, Windows 11

Where the 119 GiB lives

| part of the file | size | where |
|---|---:|---|
| dense weights (attention, shared experts, LM head) | 4.4 GiB | VRAM |
| token embedding table | 0.6 GiB | RAM, one row read per token |
| hottest routed experts | 8.8 GiB | VRAM |
| next-hottest routed experts | 6.0 GiB | pinned RAM |
| remaining routed experts | 48.5 GiB | SSD, streamed on demand |
| n-gram table | 50.7 GiB | SSD, 16 rows read per token |

So 20 GiB is in memory and 99 GiB stays on the SSD: 48.5 GiB of routed experts, streamed as the router picks them, and the 50.7 GiB hashed n-gram table (looking forward to qwen4 ngram).

Speed

  • Decode: 9.06 tok/s on the benchmark turn, 10.4 on the best turn at xhigh effort
  • Prefill: 49.2 tok/s on a 5.5k-token prompt
  • llama.cpp on the same machine: 4.9 tok/s average decode

How it works

  • VRAM holds the dense weights and the hottest experts (GCLOCK eviction), pinned RAM the next tier (read over PCIe), and the rest come off the NVMe on 8 read threads.
  • Lookahead prefetch guesses the next layer's experts and starts their reads early.
  • NVFP4 matmuls run FP4 x FP4 on the tensor cores, with no unpacking first.
  • About 270 MiB is read from the SSD per token, and roughly 75% of expert lookups hit memory.
  • Each token uses 480 experts (10 in each of 48 layers). About 377 of them are already in VRAM or RAM; the other ~103 are read from the SSD, about 270 MiB per token. This hit hit-rate is what allowed us to reach 9tps.
  • It only reads from the SSD and almost no writes so ssd should have minimal wear due to writes. But we saw SSD hit 70C during long runs.

Also supported: Qwen3.6-35B-A3B NVFP4. It fits in VRAM + RAM . You can also run it via ssd streaming and it was the learning curve for this work. On the same machine with tuned config it hits: 47.3 tok/s decode at ~4k context, 591 tok/s prefill.

Serving: an OpenAI-compatible server that renders the model's own chat template, with tool calls (still buggy and tested with Cline), the reasoning split out from the answer, and a built-in chat page.

Limits of v1: RTX 50-series / Blackwell (sm_120) only, tested on Windows 11 and WSL2 only, greedy decoding only.

Links

Questions and feedback welcome, especially from anyone running big MoEs on small cards.

▲
36
+5
16👁
r/LocalLLaMA · u/light_2earth · 12d ago
macOS 27 ships a free local LLM on Apple Silicon Macs. I made it easy to use from Node and Python

Apple Silicon Macs on macOS 26+ come with a small LLM built in. No download, no API key, and nothing leaves your Mac.

Why I built it

I was making a tool that writes API docs from code, and I didn't want users to install Ollama or paste an API key. Apple's model was already on their Mac, so I used it.

Getting it to work well was harder than expected. At temperature 0 it kept repeating itself, and a 2-second call took 20. Calls took 17 seconds instead of 1.5 until I kept one process running. So I turned all the fixes into a library: apple-llm.

What it's good for

\- Pulling structured data out of messy text, like emails into tickets or receipts into expenses. The JSON always matches your schema.

\- Tagging, summarising and rewriting

\- Private data you don't want to send anywhere

\- Tools you share with other Mac users, who don't need to download a model or get a key

What it's bad at

Coding and long reasoning. It's a small model.

There's also an optional cloud mode that uses Apple's bigger server model for harder questions. That one is not local: your prompt goes to Apple's servers, and it has a usage limit.

Node: npm install apple-llm

Python: pip install apple-llm

https://reddit.com/link/1ws5l5p/video/oarqcgm767sh1/player

GitHub: https://github.com/jagdishpal02000/apple-llm

💬 26 (+1) open on reddit ↗
▲
96
+4
28👁
r/LocalLLaMA · u/JLeonsarmiento · 12d ago
... so, yeah. post image

Finally got 3.8-Flash-Next running on my M4Pro 48GB Mac with https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

Dense 3.8-27B is just faster... and maybe better due to quantization level...

EDIT:

Hold a second, Flash-Next is actually performing faster than 27B after some key flags on llama.cpp. it's Holding up to 131K without OOM-ing..... maybe...

0.36.940.283 I srv          load:   --top-k

0.36.940.283 I srv          load:   20

0.36.940.284 I srv          load:   --ctx-size

0.36.940.284 I srv          load:   131072

0.36.940.284 I srv          load:   --cache-type-k

0.36.940.285 I srv          load:   q4\_0

0.36.940.285 I srv          load:   --cache-type-v

0.36.940.285 I srv          load:   q4\_0

0.36.940.285 I srv          load:   --flash-attn

0.36.940.285 I srv          load:   on

0.36.940.285 I srv          load:   --load-mode

0.36.940.286 I srv          load:   mmap

0.36.940.286 I srv          load:   --lazy-mode

0.36.940.286 I srv          load:   on

EDIT 2

Yes, this model is brutal. This quant at Q2\_0 in llama.cpp is out performing 27B at oQ4e in prompt processing, speed generation, but most importantly, the only thing that matters, sheer intelligence.

What a time to have 48 GB of ram !!!

▲
42
+4
20👁
r/LocalLLaMA · u/Usual_Maximum7673 · 11d ago
Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights)

TL;DR: The Jeff models are a set of Qwen3.5 and Gemma fine-tunes for zero-shot classification: small, efficient, open-weight models with respectable out-of-the-box performance that can be slotted right into code or fine-tuned/LoRA-trained as needed. Give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in \~28 ms on an M4 Max (see caveats below).

Maybe equally exciting for open model enthusiasts like myself, everything was done on local hardware: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing - all connected and monitored from my Android phone via Tailscale. Apache 2.0, Jev-compatible API. Weights: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B · Code: github.com/firelex/jeff · Videos: games table

When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware.

So here's what I did:

  • The 0.8B trains in about 2 hours and the 2B in about 3.5, on one workstation GPU (RTX PRO 6000, 96 GB).
  • \~31k synthetic training questions written and checked by Qwen3.8-Flash-Next on two DGX Sparks. No cloud GPUs, and no closed-model output in the training data.
  • The rest of the 271k training questions are public datasets converted into decisions, plus 10k code-built probability questions.

Benchmarks (4,599 questions: BBH, Financial PhraseBank, JudgeBench, RAGTruth, WinoGrande):

|Model|Untrained base|Jeff (trained)|Calibration error|
|:-|:-|:-|:-|
|Qwen3.5-0.8B|45.3%|79.1%|0.049|
|Qwen3.5-2B|46.5%|83.1%|0.028|
|Jev (published)||83.0%|≈0.06|
|AutoJev-27B (published)||84.9%|—|

The caveat: the published numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86–89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64–68% against Jev's 94%, and \~50% on JevBench's hard tier against \~73%. See the HuggingFace model card for details. But that's not surprising, and I don't think it matters. No 0.8B or 2B model reasons like an LLM, and I don't think anyone should expect it to. The Jeff models are extremely fast judgement-callers (much faster than Jev), and have reasonable out-of-the-box performance. In one of my apps, I used the 0.8B model for voice-based navigation, and with a quick fine-tune, I got to real-time performance (24ms) at almost 100% accuracy.

Now the fun part: games, as a zero-shot test. Games are not the ideal zero-shot test, but they're fun, and the TypeSafe guys (Jev) did it, too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one. The options say what each move leads to, never which one is right.

|20 episodes each|Doom (kills)|Frogger (crossings)|Pac-Man (pellets of 98)|
|:-|:-|:-|:-|
|Random moves|−0.05|0|11.2|
|Hand-coded rule bot|6.55|10.25|94.1|
|Qwen3.5-0.8B, untrained|5.0|1.0|25.8|
|Jeff 0.8B|6.55|10.3|57.0|
|Jeff 2B|−0.9|6.0|41.2|

Jev's published Doom score is also 6.55, but its prompt spells out the aiming rule (fire when the bearing is between −8 and +8 degrees) and it takes \~212 ms per call over its API. Jeff gets "the nearest monster is a little to your left" and decides in \~29 ms on my Mac.

Lessons learned:

  • System 1 models are here to stay. Having the ability to process unstructured data at software speed inside an app is extremely powerful. And being able to do this locally is fantastic.
  • A small model is a classifier, not a planner. Models in the 0.8B-2B range don't reason like Qwen3.8-27B or Jev. But they also don't need to. As long as you present the options in the right way, you can get up to 50 decisions per second (depending on your hardware).
  • Fine-tune it if needed. If the models' zero-shot performance isn't good enough for you, fine-tune them briefly or add a LoRA adapter.
  • Wording matters enormously. Play around with how you present the options. Giving Frogger's final step option the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Previously, the frog just stayed on the last log.
  • Bigger isn't better. As the game tests showed, the untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers "turn away from the nearest monster"), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.
  • Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, not reliably, and in a real-time loop the mistakes compound.

Happy to answer questions about the pipeline (synthetic data from a local teacher, leak filter, calibration) or the game harness. Everything, including the videos, is linked above.

▲
57
+4
21👁
r/LocalLLaMA · u/lordekeen · 11d ago
Qwen 3.8 is a workhorse

https://preview.redd.it/23cumykhabsh1.png?width=944&format=png&auto=w…

Reminder that you can put Qwen 3.8 27B as a subagent and its a workhorse. Pic: using DeepSeek v4.1 Flash as Orchestrator in Pi, Qwen 3.8 27B GSQ RCO in llama.cpp.

▲
63
+4
27👁
r/LocalLLaMA · u/starkruzr · 11d ago
3090 NVLink bridges: are there clones of these? why are they so insanely expensive?

I think I need a 3-slot for my two cards. but holy fuck these things are pricey.

💬 69 (-2) open on reddit ↗
▲
57
+4
30👁
r/LocalLLaMA · u/MomentJolly3535 · 12d ago
Swift 1.5 Qwen3.8 27b (A must-have for low thinking!)

Just made this post for those who missed it : https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b

UkisAI released their updated Qwen 27B (tuned for token efficiency). I grabbed the IQ4\_XS quant to test against Unsloth's Q4\_K\_S:

Low-thinking: UkisAI consistently beat Unsloth in most of my tests.

High-thinking: Unsloth still pulled ahead here.

I was struggling with a custom script in Directory Opus. I gave it to Gemini Flash (medium thinking on Antigravity free tier) it looped for 40 minutes, tried many things, burnt all the weekly limit-tokens, and failed to solve it.

Fed the exact same problem to this 27B model: Fixed it completely in 6 minutes on an old 3090 (67 t/s)

Honestly i was kinda impressed, didn't expect an IQ4\_XS quant of a 27B model in low thinking to beat a major cloud model.

▲
59
+3
27👁
r/LocalLLaMA · u/streppelchen · 11d ago
Minisforum MS-S1 MAX-P495 @ €7.799,00

MINISFORUM MS-S1 MAX-P495 – Minisforum EU

Expected to ship mid october.

At that price point, it doesn't make a whole lot of sense in my opinion.

I get that ram prices are where they are, i get that it's a newer model of hardware, but twice the price for 50% more ram and \~5-10% more performance is just hard, especially when compared to the recently released m5 ultra studio.

▲
4
+3
8👁
r/LocalLLaMA · u/fgoricha · 13d ago
Dual 3090 stability troubleshooting

​ I made a previous post about my x299 stability issues. Seems another stability issue has popped up since then, but overall has been much stable. Seems to only happen when my i9 is working hard the dual 3090s are also working hard at the same time. Specs: EVGA X299 FTW K \\Intel i9-7940X 64 GB RAM (4 × 16 GB) 2 × RTX 3090 Founders Edition ASRock 1600 W PSU Each GPU installed in its own x16-length PCIe slot Roughly one slot of space between the GPUs Originally, I was running 128 GB (4 × 32 GB). With both GPUs under sustained AI workloads, the entire computer would eventually hard-lock: display signal gone, network connection gone, no apparent activity, but fans/lights remained on until I held the power button. I switched to 64 GB using 4 × 16 GB and that seemed to resolve that particular stability problem. The board is supposed to support the 4 × 32 GB configuration with the latest BIOS, but apparently my system wasn't happy with it. Now the next problem....... Each RTX 3090 is stable individually at PCIe Gen 3. However, when I run both GPUs together under heavy load, particularly while the i9 is also being heavily utilized, I still get instability with PCIe set to Gen 3. Hard locking with FF displayed on the mobo. Have to hard restart and boots fine into Windows. If I manually force the PCIe slots to Gen 2, the system appears to be stable with both 3090s and the CPU working simultaneously. So my question is: For AI/ML workloads, how much performance am I realistically giving up by running the two 3090s at PCIe Gen 2 instead of Gen 3? Obviously I'm going to benchmark my actual workloads both ways, but I'm interested in other people's experience. Most of my work is inference/training where the models and batches are primarily staying in GPU VRAM rather than constantly transferring huge amounts of data across PCIe. I'm also curious what the Gen 2 stability might point toward. Since either GPU works individually at Gen 3, but dual-GPU Gen 3 becomes unstable under heavy CPU/GPU load, could this indicate a motherboard/PCIe signal-integrity issue, CPU PCIe controller issue, BIOS setting, or something else specific to X299? Any ideas for additional troubleshooting would be appreciated. TLDR: How much performance am I losing using gen2 pcie vs gen3 pcie? Edit for additional info: Using Windows Using llama.cpp Stability issues happen when power limited at 200W and no power limiting More edits: Open air case Used a variety of diagnostic tools including OCCT, MemTest86, and HWiNFO64 to test each part individually. Seems to lock up when cpu and both gpus are going at 100%. Even locks up if cpu and one gpu is going at 100% while the second gpu is idle. But oddly, no problems in the same scenario but with the second gpu removed from its pcie slot

▲
5
+2
18👁
r/LocalLLaMA · u/Perfect-Campaign9551 · 11d ago
Recommended way to run Qwen 3.8 27b on a 3090 in Windows?

I know this may have been asked a lot, but I'm not an LLM expert yet (have never set up vllm myself or llama.cpp myself yet) I've only used scripts other people have set up.

Is there a simple way / steps to follow to run Qwen 3.8 27b in Windows with my single 3090?

I can run it straight up with Ollama with a 64k context but it seems like it only works reliably in chat and not in OpenCode (in OpenCode it "works" but at one point it "hung up" on me it seemed like. Not sure if maybe it was busy thinking still)

I've found quite a few threads that are close to what i'm asking for but many I think use WSL or something, too. Which I'm also not super experienced with yet.

I found the HyperQwen repo but their documentation is like...120% all technical and not very user friendly at all. I CAN do technical stuff! But it's barely passable as "do this, and then this" type of docs at the moment.

Ninfer is only for 5090 cards from what I read.

EDIT: Thanks guys, I was able to use llama-cpp-windows-manager project to get Qwen 3.8 27B up and going (Q4\_K\_M) . I have a 98K context and get 70tok/s and it's working with Open Code just fine. Very usable.

💬 34 (+2) open on reddit ↗
▲
55
+2
24👁
▲
46
+2
14👁
r/LocalLLaMA · u/KingGongzilla · 11d ago
Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090

Hi everyone :)

The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090.

To get the best of both worlds, I adapted Swift 1 and Swift 1.5 for HyperQwen and benchmarked them against HyperQwen’s specialized W4A16 AutoRound fast quant for speed and quality.

Performance

|Model|Average time/task ↓|Average output tokens/task ↓|Decode tok/s ↑|
|:-|:-|:-|:-|
|Qwen - HyperQwen fast quant|108.1 s|8,985|112.1|
|Swift 1.0 + HyperQwen|66.2 s|5,245|105.9|
|Swift 1.5 + HyperQwen, INT8 heads|72.2 s|5,751|104.0|
|Swift 1.5 + HyperQwen INT4 heads|68.2 s|5,669|107.2|

All models ran on one RTX 3090 24GB with FP8 KV cache and 150k configured context. Average task time covers about 630 tasks from different benchmarks listed below.

Swift-1 seems to use the fewest tokens and has fastest task completion time. Swift-1.5-INT4 achieves roughly 37% lower average time per task compared to the HyperQwen Qwen fast model. Despite slightly lower TPS (due to lower draft acceptance) it finishes sooner because it generates fewer tokens.

Quality

There are some minor quality and performance tradeoffs between the models:

|Test|Qwen HyperQwen fast|Swift 1.0|Swift 1.5 INT8 heads|Swift 1.5 INT4 heads|
|:-|:-|:-|:-|:-|
|GSM8K, 200 questions|97.5%|98.0%|98.0%|97.5%|
|IFBench, 300 prompts, strict|74.0%|73.3%|73.7%|72.3%|
|LiveCodeBench, (100-problem subset)|90%|89%|89%|91%|
|Custom tool-call/JSON eval|29/30|28/30|30/30|30/30|
|English/Python perplexity ↓|6.551|6.605|6.643|6.679|

Applied changes to Swift models to adapt for HyperQwen:

Changes to Swift models:

  • Kept the upstream AWQ INT4 model weights and converted embeddings to INT8.
  • Swift 1.0 and Swift 1.5 INT8-head variants: quantized the output head and MTP (multi-token prediction) linear layers to INT8 and added HyperQwen’s reference draft vocabulary for speculative decoding.
  • Swift 1.5 INT4-heads: quantized the output head and MTP linear layers to GPTQ INT4 instead, and built a Swift-specific 65,536-token draft vocabulary.

Setup

If you want to try it yourself, point your coding agent at these setup instructions and ask it to set up Swift 1.5 + HyperQwen on your machine.

All three models can be found here:
https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks

Big shoutout to UkisAI for Swift and syv-ai for HyperQwen! It's genuinely insane to be able to run these models on an RTX3090 at those speeds!

▲
37
+2
19👁
r/LocalLLaMA · u/PilgrimofHaqq2 · 11d ago
9 prompt rules cut my coding agent's wasted thinking up to 70% (GLM 5.3 & GLM 5.3 Flash)

360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%.

The block (shipped to global instructions):

Thinking discipline 1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it. 2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling. 3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps. 4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue. 5. Do not revise just to agree. If the user pushes back without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. Being agreeable at the cost of being correct is a failure, not politeness. 6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason. 7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle. 8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do. 9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke.

Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam

💬 21 (+1) open on reddit ↗
▲
32
+2
13👁
r/LocalLLaMA · u/Balance- · 12d ago
It would be really cool to have an official 3D-print engineering benchmark/leaderboard like this post image

Someone prompted different LLMs to generate CAD code for a bridge under fixed constraints (2-foot span, under 500g filament, 18-hour print limit), printed them, and load-tested them to failure. The results were quite varied: some models couldn't even design parts that fit together, while the winner held over 100 lbs. Most benchmarks don't really capture physical intuition, spatial reasoning, and functional code generation at the same time. I would love a standardized benchmark and leaderboard for this.

▲
7
+2
6👁
r/LocalLLaMA · u/Then_Blueberry7290 · 12d ago
LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF

Just recently stubled upon with this modell:LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF I'm just stay away from "magic" models, but this model size got my eyes on: With vision capabilities this is under 17GB, which means i can use it 32GB vram with full Context size (262k), bigger ubatch, and mtp4. Of course vision goes to ram, not gpu. Other similar model with nvfp4 line, usually 19-20GB in size or more. I tried in with llama.cpp, speed is 40-113 t/s (76 in my benchmark) with 262k context. Under normal agentic workin it is 45-65 t/s. (2x5060ti16GB OC) For example thinkingcap nvfp with vllm i can only have 160k context (cannot offload mmproj to ram) First glance it is the same as the other swift models (nvfp4) in quality. So My question is what is the tradeoff of this modell?

▲
24
+2
14👁
r/LocalLLaMA · u/NickCanCode · 12d ago
Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?

There is always at least 1+GB of VRAM not usable not matter how I set the --tensor-split (-ts) param. I tiny shift toward one side will move the weight significantly to the other side. 😵‍💫 Adjusting context will increase/decrease usage on both side. --tensor-split 499,501 = GPU1 12.5 GB, GPU2 15.4 GB --tensor-split 501, 499 = GPU1 14.7 GB, GPU2 13.4 GB Tried --spec-draft-device with CUDA0 and CUDA1 separately, no change at all. (same distribution as above) Also tried --mmproj-device, no much difference. Tried --no-mmproj-offload, somehow the lower side get even lower 🫣 = GPU1 14.7 GB, GPU2 12.3 GB I guess it is related to MTP + Tensor Parallel stuff being concentrated on one GPU. No idea how to solve this. llama-server \ --batch-size 2048 \ --cache-ram 24384 \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ --chat-template-file /mnt/AI/models/qwen-chat-template-froggeric-22.5.jinja \ --checkpoint-min-step 1024 \ --ctx-checkpoints 32 \ --ctx-size 192000 \ --fit off \ --gpu-layers all \ --image-min-tokens 1024 \ --load-mode none \ --main-gpu 1 \ --min-p 0.0 \ --mmproj /mnt/AI/models/Qwen3.8-27B-mmproj-BF16.gguf \ --model /mnt/AI/models/Qwen3.8-27B-NVFP4-MID-HIGH.gguf \ --parallel 1 \ --presence-penalty 0.0 \ --repeat-penalty 1.0 \ --spec-draft-n-max 5 \ --spec-draft-n-min 0 \ --spec-draft-ngl all \ --spec-draft-type-k q8_0 \ --spec-draft-type-v q8_0 \ --spec-type draft-mtp \ --split-mode tensor \ --temp 1 \ --tensor-split 499,501 \ --top-k 20 \ --top-p 0.95 \ --n-gpu-layers-draft all \ --no-prefill-assistant \ --reasoning-preserve

💬 21 (+2) open on reddit ↗
▲
28
+2
21👁
r/LocalLLaMA · u/Exciting-Engine882 · 12d ago
is switching from llama cpp to vllm worth it

I have hp z8 g4 with 512 ram and 1x3090 1x5060 16gb. has anyone made the transition from llama cpp to vllm recently? is it worth it? docker under windows or full linux install? I am mainly interested in the model support, it seems that many new local models are supported day 0 in official vllm, while for llama cpp it takes months sometimes. LE: I want to use it for big'ish moe models, that would have to offload some tensors to system ram. I will use it just for myself. I don' t need it to be faster than llama cpp, if it runs at about the same speed it is fine , as long as it works.

💬 68 (+1) open on reddit ↗
▲
7
+2
8👁
r/LocalLLaMA · u/arbv · 12d ago
Improved chat template for Laguna XS / S 2.1 (configurable forced thinking, preserve_thinking toggle, and stability fixes)

Following up on my previous post about the GPT-OSS template, here is an updated chat template for Poolside's Laguna models (XS and S 2.1). The main reason I ended up putting this together was inconsistent reasoning. By default, the model is supposed to decide when to think on its own, but in practice it's pretty lazy - especially the XS variant - and often skips thinking right when it needs it most. When these models do think, they do so well. All in all, a good models to have around. Also they write well in English (to my non-native eye, at least). Laguna XS 2.1 in particular deserves more attention, IMO. What I like about these models is that allow toggling reasoning mid-conversation without invalidating the prefix cache. Very handy. I added a force_thinking toggle (using the prompt trick discovered by u/SnooPaintings8639) to make it think on every turn, plus a reasoning_effort parameter (none, auto, max) if you prefer an easy preset over juggling booleans (and to make it easier to use in Pi and, possibly, other harnesses). I have fixed some other things along the way. Firstly, the preserve_thinking toggle. The original template permanently forces historical reasoning preservation on. That's great for prefix cache and agentic tool loops, but if you're just having a normal chat, dragging thousands of past reasoning tokens around shreds your context window fast. You can now turn it off. Secondly, I added basic validation to catch smuggled control tokens across roles (can be turned off via allow_injection: true). By default, all settings match the upstream behavior (enable_thinking=true, preserve_thinking=true, force_thinking=false), so it acts as a direct drop-in replacement if you don't want to mess with the new knobs. Template repo: https://huggingface.co/arbv/laguna-2.1-fixed-jinja-template I also included some recommended sampling settings for llama.cpp (BF16 and quantised) and a config snippet for Pi (models.json) in the README. P.S. Also casting u/matthiasgalle (the Laguna post-train lead) to take a look and for a data point.

▲
16
+2
17👁
r/LocalLLaMA · u/WebAssemblyMan · 12d ago
What if open-source AI focused less on giant models and more on reusable capabilities?

Instead of everyone building another general-purpose model, the community could distill open models into domain specialists—biology, Python, accounting, OCR, and more. Developers could combine these capabilities into local tools: small model + OCR + accounting → local accounting assistant Like Linux, open-source AI could grow through shared components rather than complete systems. Could domain capabilities become the fundamental unit of contribution?

▲
26
+2
13👁
r/LocalLLaMA · u/eribob · 12d ago
Worth going from Qwen3.8 27B to flash next or maaybe deepseek v4 flash?

I am running qwen3.8 27b on my dual rtx 3090 (fp8 quant, unquantized cache, 129k context) and I think it works decently well with hermes, opencode etc. But! I am tempted by the new models coming out such as qwen3.8 flash next, deepseek v4 flash, glm 5.3 flash. However, there is a big jump in vram and therefore in cost! The least expensive option seems to be buying 2 of those cmp 170hx 64gb cards for roughly 6-7000 usd in total (that is the price I can find for verified cards here in europe at least). With that I would get another 128gb of vram for a total of 176gb so I could run I think around 3-4bit quants of the above models, right?). I am thinking that it might be faster because of moe but not sure how much smarter? For that kind of money I would want a real noticable improvement! 4xv100 32gb would be cheaper (maybe half price?), but even more hassle to set up, more power draw, and slower. What do you think? The free option is to just wait for qwen4 27b and (hopefully) just download more IQ.

▲
18
+2
6👁
r/LocalLLaMA · u/pseudobacon · 12d ago
Anyone with experience using dual SXM2 to PCIe card? post image

As per the title. My thinking is that if you run one of these then you would just need a computer with a pcie slot to be up and running. With an nvlink baseboard you need a separate PLX adapter card as well as the mini SAS cables from the baseboard connected up?

▲
67
+1
32👁
r/LocalLLaMA · u/ThePrimeClock · 13d ago
SupersonicLabs/Julia-1 · Hugging Face

New open source Jev like model for running on local devices from a group called Supersonic Labs.

It's a 144M param local non-generative local classifier.

From their site:

Julia 1 opens our research into compact decision models. It builds on mmBERT-small, a multilingual encoder, and chooses among answers supplied with a question. It has 144.3 million parameters and runs on a CPU.
Our question: can one model classify, rank levels, and answer yes-or-no questions as the options change? Julia 1 is the first result of that investigation. Here are its successes, its failures, and the methods we used to measure them.
▲
13
+1
9👁
r/LocalLLaMA · u/SnooPeripherals5313 · 11d ago
3D/2D Text Visualisation post image

Like everyone, I use 3js for data visualisation. But while semantic clusters are interesting, they don't confer much practical information alone.

So I did something very simple: a query spatially re-assembles the nodes, and you can switch them to text.

Honestly, it's hard to swing 3D viz for text as a genuinely useful feature and not a novelty, but I get the feeling there's still some potential in the idea. Would be good to discuss, I'm sure someone here has made a better implementation.

▲
8
+1
10👁
r/LocalLLaMA · u/otacon6531 · 11d ago
IQ Quants still slow on P40?

I have been using Qwen 3.6:35b IQ4 via llama.cpp on my p40 and am getting anywhere between 37 - 83 tok/s (mtp is on). Prefill usually starts at 600 and slowly degrades as it continues processing so 600 for short prompts and more like 300-400 by the end of a long prompt. It hurts, but it is what my budget allows.

AI told me IQ quants are noticeably slower on the P40 and it referenced (https://www.reddit.com/r/LocalLLaMA/comments/1dmhpud/are\_iq\_quants\_slow\_o…) from two years ago, but I didn't feel it being slower when I moved from Q4 to IQ4, so...

What am I missing? Are IQ Quants actually a significant amount slower on the P40 or is this outdated information?

▲
27
+1
17👁
r/LocalLLaMA · u/Brief-Tap-6616 · 11d ago
95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

Hello everyone! A little while back I posted about LlamAmpere, a fork of Llama.cpp with Ampere-specific improvements (though it is caught up to main and will support other hardware, too).

Thank you to everyone that tried it out and shared back their results across the 30xx cards. I'm happy to share I've pushed v0.4 out this morning. On the 4.6bpw model tested, speeds improved \~10% vs the last version while also improving the max context by 10%+ (technically, it can go above 262K, but I have not tested any custom kernels or graphs to support YaRN).

The closest competition comes from vLLM, keeping within <10%, but does so with lower maximum context. It is significantly faster than other llama.cpp options tested.

https://preview.redd.it/u70q7lew1bsh1.png?width=1080&format=png&auto=…

[](https://preview.redd.it/95-tps-through-100k-generated-262k-ctx-on-a-single-30…)

There's also a number of other improvements for other quants/formats, with EXL3 seeing significant speed up (\~80% the speed of the 4-XS-M quant tested). It has a slightly lower KLD, but not a range I have found stat significance for at the task level, so I sticking with the XS-M model for now (built on top of Swift-qwen's distill, which is far more token efficient than the stock train for \~1% performance loss). the 4.3 bpw EXL3 model does provide a bit more room if you are interested in 2+ concurrent predictions. Improvements in this format and the IQ2/3 codebook quants will be most useful for people on 12/16/20 GB setups. These measurements are at temp=1, vs some of the vanity speeds you will see people claim with temp=0 and/or short generations.

As always, please share your results + config details so I can keep improving!

fork is here: https://github.com/JakeATX/llamAmpere/blob/main/QWEN\_AMPERE.md#build-and-run**

model used here:

https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF**

Build command:

git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server

Build + launch (linux):

git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server
curl -L -o ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf \\
https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF/resolve/main/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf
\-m ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf -c 262144 \\
\-ngl 99 -fa on -ctk turbo5 -ctv turbo4 -b 4096 -ub 1024 -t 8 -tb 8 --parallel 1
cd
./build-sm86/bin/llama-server

Previously, people had expressed concern over quantizing KV cache, and TQ specifically. The TL;DR on that is that any reasonable KV quantization strategy (at least for hybrid attention models like Qwen) is going to be swamped by quantization of the weights. The KV quant we're using here (TQ5/TQ4) is less than 1/3 of the KLD we see when moving from 8 bit weights to 4.6 bit weights (and the KLD is only partially additive, so some of the incremental errors cancel out). There was no statistical significance when testing this KV quant at the task level against 8/8 kv (just trivial variations in sentence length). I will be adding KVaRN in the next release, but with a better codec than currently available elsewhere, so it requires a bit more testing before release.

v0.5 will be focused primarily on the 12GB cards, but this should have generation-wide speed ups, so even if you're not on 24GB, please share your results.

Enjoy!

▲
4
+1
7👁
r/LocalLLaMA · u/whoami-233 · 11d ago
Anyone running multi GPU A100 80GB Cards?

Hey guys, I am looking for benchmarks for people running multi node (4 or more) A100 80 GB cards and seeing what results and models they are getting. Something with VLLM and multi users would be very useful. Or if you know of a place I can find such results please let me know! Appreciated!

▲
29
+1
12👁
r/LocalLLaMA · u/ea_man · 12d ago
Who wants to try a Pi trick for 27B to reuse prompt prefill between different sessions?

You know that when you start Pi you have to process the initial prompt, that takes some time when you use the slow dense QWEN 27B (that's the very reason why you use Pi instead of Cloud Code!), then you start to add extensions, tools, your append.md and whatever... Well now that got big, like 20k big and it does bother. So let's cache the "initial prompt" PP, so that when you start an new pi session: TA-DA! Instant ready, jolly good. Well if you tried to do that with Pi, llama.cpp and QWEN 3.x hybrid KV all kind of things step in your way to prevent that, so many that I won't even start to count I'll just tell you what to do: 1. dwl and install this Pi extension (tested on Pi version 0.87.1 ) 2. dwl and patch lama.cpp: yup no way around this if you kill the server between session, suck it or leave now 3. do your self a favor and use Froggeric template, for all your QWEN models, even the old ones. \---- So I'll help ya and give you some kinda useful parameters to launch the thing too: --slot-save-path /home/eaman/llama/slot_caches/ --ctx-checkpoints 32 --checkpoint-min-step 4096 -np 1 --chat-template-file chat_template_3.8.jinja You need to have save slots, that's the whole point, the caching is meant to resist restarts. Beware the chunk of blocks cached follow ubatch boundaries, so yeah try to keep that down if you wanna cache some more. Now in the extension README.md there's explanation of env variables that you can tweek, you go read those and edit accordingly to your setup (or have your LLM read that and suggest / config for you), TLDR you need at least: export PI_PREFIX_CACHE_BASE_URL=http://localhost:8080/v1 export PI_PREFIX_CACHE_PERSIST=1 export PI_PREFIX_CACHE_SLOT_DIR=/home/eaman/llama/slot_caches Disclaimer: this is not an easy thing, if you are not familiar with patching llama.cpp and installing extensions manually leave this thread for an other day. On the other hand this thing kinda works for me so if someone else is interested after some testing (because all kind of evil things want to break prompt caching) I'll upload a final extension and see about the llama.cpp problem with saved check points.. Possible results: https://preview.redd.it/xobe9ve7h5sh1.png?width=1281&format=png&auto=… EDIT: made a version for OpenCode: https://store.piffa.net/lm/ocache/ For those of you who want a basic understanding of the problems and solution regarding caching the prompt I've asked the LLM to make a short summary.

▲
4
+1
6👁
r/LocalLLaMA · u/bulletrhli · 12d ago
Power Limits, Local AI, and Questionable Uses of My Free Time

Edit 1: Okay, I have been checking out Unsloth and wow. Just wow. Thank you so much for your suggestions. This is such a way better tool and I am going to go crazy with this. Good day data nerds! I am trying to get more into running models, learning about agentic workflows, and creating my own tools. But as you do (right?) I had to fine tune my current setup. With the way the markets are right now, it only makes sense to make the most out of what I got. My day job, typically, is around data, numbers, and programming; only two of those I am good at, I'll let you guess which ones. So, yes, here come some data sheets and pretty graphs. Don't worry, you don't have to go through the data, but you can if you want. The graphs cover key metrics spat out by Ollama such as the tokens per second, duration, eval rates etc. I also added a cheeky "tokens/s/W" which, technically is not perfect since I do not measure wattage over time, but I did observe the watts during prompts, and I have a few things to mention about that later. Okay, let's start off with the specs because you probably think I am rocking the good stuff since I am so invested in this topic (haha) Lenovo M920Q 16GB DDR4 2660MHz Intel i5-8500T (6C/6T) Gigabyte Gaming OC 3070 8GB (Over OcuLink at Gen3x4 speeds) I am running OpenWebUI with Ollama in an LXC on my Proxmox server. This is one of my nodes and it is dedicated to my models. I have given it all of the cores, 14GB of RAM, 2GB swap. Nothing crazy to write home about, see? Okay, so, one of the things I wanted to know what, with the models that I run day to day, how effective are they at different GPU power limits. Man, if only I had known how much of a rabbit hole I would go down to do this (sorry wife). I only run 4 models, nothing too crazy, until you run 3 tests per model, for each power limit from 100 to 220 (156 runs in total), and each run of each model taking around 3 or 4 minutes since I have to unload the model each time to not have any prompt caching. Afterwards I would average the results and add that to the sheet. Really gave the fingers a workout since I now am a proud owner of a 60% keyboard for the first time and I no longer have a numpad... I'll remember that for next time. That being said, the switches are soooo creamy, a valiant tradeoff. So what models am I running? Glad you asked. The 3070 does limit me quite a bit, but with so many models available and so many smarter people than me who can quantize the models, I have found these models fit my needs. For the most part everything runs in the VRAM, except for 2, but those come with asterisks. gemma4:e4b qwen3.5 qwen3-vl\ deepseek-coder-v2\ For the vl model, it runs really well at a 23% CPU to 77% GPU ratio. Totally fine for my purposes. As for deepseek, it is a 40/60 ratio, but I luck out as it is a mixture of expert's model but even with the ratio, it is extremely performant. Gemma is by far my best model, and I have the most context room available at around 16k whereas the remainder I have sitting at 8k. Both gemma and qwen3.5 fit entirely in my GPUs VRAM. A couple things I noticed: Gemma4, is so good. Doesn't overthink, understands the prompt, remains as concise with the right tone I want. A really good day to day general model to work with. I also love the extra headroom for the context. Qwen3.5, a heavy thinker. Whilst it does a great job on the output, it spends a lot of time thinking and generating a lot of tokens. Power usage is pretty good, broke around 203W at one point and anything below that it just sat at whatever the power limit was set to. Qwen3-vl, also a major over-thinker. It spends so much time thinking that it balloons the context. I probably do not understand how to use this very well because when reading its thoughts it knows the answer pretty early on but it just gets into a thought trap (eh-yo). It does always output the correct answer, or the best it can, but I might move away from a reasoning vision model and stick to traditional ocr. If you have a better model or know how to prompt this better, I would love the help. Oh, one final note, this model LOVES power. Always maxes whatever I have, not that it increased performance directly, but it just loved power. Deepseek-coder-v2, this model rocks. It is extremely performant even though I technically on paper can't fit it. Especially for smaller asks with good bounds in place, it doesn't think, it just does and gives me excellent code back. I have yet to make it build me anything bigger but that is something I will experiment more with later. It is a weird one though, consistently using a fraction of the power budget available to it. Under 150W power limits, not once did my fans kick in on my GPU. Even my CPU fans (which are mucho loudo) rarely turned on, or if they did, they were not sounding like rocket engines. Not sure if those two are related but, eh, just something I noticed. Deepseek I think had some anomalous results with some spikes, but I can't be arsed to do them again. For the most part the results are fairly consistent and show a trend. Same for the vision model by qwen, oh well. My thoughts? It probably doesn't matter too much for most of us on a budget. Just let her rip, but if you want to shave off some heat, just lower your power down a little bit and monitor your temps. For the most part, you are probably fine. Honestly, it is 3am at this point, have a look at the spreadsheet! It was a lot of fun (I think) doing this. Interesting observations were made where I can balance my power limits, save... well pennies, and not have to listen to fans. So, works for me. https://docs.google.com/spreadsheets/d/1CEAr40nemlsK727QMvBnPcJDg548s-Bx/edit?usp=sharing&ouid=105501696463520933058&rtpof=true&sd=true Managers love graphs

▲
216
 
24👁
r/LocalLLaMA · u/jonas__m · 11d ago
Speculative reward hacking in coding agents post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "*Let me look at the problem from the grader's perspective*" and referred to "*hidden tests*", "*test authors*", and "*the checker*".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

\[Pictured example shows verbatim quotes from agent's reasoning\] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

▲
97
 
49👁
r/LocalLLaMA · u/wombweed · 11d ago
I am concerned about all these disparate hard forks that target specific architectures instead of opening a PR against upstream

Other than the obvious self promotion, is there a practical reason people do this that I am missing? There's dozens of llamacpp forks with silly names that are supposedly "optimized" for this or that specific GPU and seem to have zero intention to merge into upstream. Am I missing the real reasons why this happens so often? Why do people think it's OK to do this? In my experience in the open source community this is generally frowned upon.

I don't know if it's just a me problem that this kind of thing puts me off so much. I am usually quite grateful for PR feedback and conscientious about the code I put out there; I take pride in submitting high quality code that meets or exceeds the standards of a given project. Of course there is nothing ethically wrong with hard forks or taking shortcuts if you find the collaborative process cumbersome, but personally I wouldn't promote my fork in such cases, let alone go out of my way to add custom branding with a Reddit announcement post etc. since the effort required to do so seems roughly equivalent to the effort required to meet the contributor standards. In contrast, many of the authors of these forks seem very eager to have others adopt their rebranded fork for production use cases. There just seems to be a big disconnect, idk.

Edit: some great discussion in this thread, thanks to all who responded. Consensus seems to be that (excluding the obvious low-effort engagement bait forks) the base project has to meet many compatibility requirements while a downstream project can be more focused, which is a great point.

💬 166 (+3) open on reddit ↗
▲
72
 
47👁
r/LocalLLaMA · u/Mrinohk · 12d ago
Don't trust frontier models when asking about budget hardware!

Early this year when I was first looking at building up my inference capability you could get the 16GB Tesla P100s for between $60 and $80. Asked claude about it, told me absolutely not worth it. No tensor cores, bad int4/int8, no BF16, not worth it. Needs special power accommodations, Above 4G decoding option in the bios (it made it out like it was some rare option), and a semi-exotic cooling solution.

Optimized the shit out of my RX6600XT in llama.cpp as a result. Got pretty far.

Decided to say fuck it, bought a single P100 last week, finally showed up day before yesterday. Got a newer power supply with the appropriate connections (not hard, not that expensive, seen options as cheap as $60 from good brands, I spent $100 on one with some headroom), multiple llama.cpp forks and patches that carry some wild optimizations to handle the capability gap, and a 3D printed housing for a 94mm fan from noctua. Doesn't generate enough static pressure to keep it cool during prefill, but more than strong enough for the generation step.

The numbers I was getting before, with my RX6600 XT with Qwen3.6 35B A3B UD\_Q4\_K\_XL with MTP and --cpu-moe:

PP \~800 at 0 ctx, drops to \~700 by 10k

TG \~30-35 prose, 45-50 code.

This setup could do 64k context (and possibly higher) at 16bit kv. cpu-moe helps a ton in that respect.

With just a little bit of tuning, and using specifically the patches from shinbunbun for llama.cpp, same model with the same MTP settings, --n-cpu-moe 22:

PP \~600 at 0 ctx, 440-500 by 10k

TG \~54-60 prose, 66-72 code.

Running only 32k context right now to make it work. Could fit more with a higher n-cpu-moe, but my harness doesn't need that much (rarely see it over 30k, persistent memory leads to chats that simply aren't meant to last).

I know the capability gap between 3.6 35B and the basically any of the qwen 3.X 27B models is pretty big, but this is huge for the price. They've gone up since I bought mine, about \~$15 across the board. Still something you can get for under $100 and makes for inference that is simply impossible to get at that price otherwise.

I've got another one coming so I can go full offload on the model, and maybe even start playing with 3.8 27b. Right now IQ3\_K\_XL I get around 9 tokens per second with MTP, and basically no real context. Don't actually know if splitting a model that fits in one card across multiple helps speed, that's completely new territory for me, but I'm having fun regardless.

Card is seriously underrated. It's a great, (relatively) inexpensive way to get capable compute to finally start doing local AI stuff. I went from having to just leave my computer alone while the model was running and do everything from my macbook (good bye gaming) to being able to let the model live and work in the background while I'm doing basically anything on my PC. When the second arrives, I'll be planning my dedicated inference box they'll both live in. Feeling inspired by that guy cooling his PC with a VW radiator.

💬 54 (+3) open on reddit ↗
▲
63
 
33👁
r/LocalLLaMA · u/ZenZombie117 · 11d ago
Liked Muse, so I cut the 30B model in half by width, distilled it back, and it does 57 of 60 tool tasks its parent does 60 of

I've liked how Muse-Glimmer worked, so I wanted to see if I could produce a smaller "kid" out of it. Ornith's sharp decisions on when to think and which tool to call were the other thing I liked, so Ornith-1.0-9B got to be the policy teacher while the parent wrote the words. No RL anywhere, distillation only.

I present to you Xyntetik-Kvist-14B.

What it is good for

  • Smaller than the Muse parent but still manages most tool tasks: 57 of 60 held-out closed-loop tasks (contacts, weather, flights, currency, dates, units, stock), scored by re-executing the calls against ground truth, where the parent does 60.
  • Fits a 24 GB card whole at Q8_0 (15.4 GB) or the Q5_0 mix (10.3 GB) and serves an OpenAI-, Anthropic- and Responses-compatible API through Xyntetik Runner, so it drops into an agent loop you already have.
  • Every failed attempt is published beside it: 12 gated runs, 2 full passes, one shipped. The training record has the preregistrations, the amendments and the defects, so you can see exactly where it breaks before you build on it.
  • Give it a calculator tool for arithmetic. Without one it gets "17% of 2,340" wrong, and the card says so.

Numbers, from the card

| claim | number |
|---|---|
| parent | Muse-Glimmer-30B, cut by width (hidden 6,656 to 5,760, FFN 19,968 to 10,240, heads 32 to 24), all 52 layers kept |
| size | 14.44 B parameters; BF16 28.9 GB, Q8_0 15.4 GB, Q5_0 mix 10.3 GB |
| distillation | 6,000 steps, 98.3 M tokens, 162 hours, then 1,440 steps on agentic trajectories |
| fidelity to parent | KLD 0.762, margin-qualified top-1 84.0% on 45,056 held-out positions (a student's row, not the quant bar) |
| tool tasks | 57 of 60 held-out, re-executed against ground truth; parent 60, untrained control 0 |
| format and calls | 199 of 200 first turns well formed; 99 of 99 tool calls valid |
| attempts | 12 gated attempts, 2 full passes, attempt 12 shipped |
| weak spot | calc tasks 12 of 15 over 160; 7 of 160 runs end in a reasoning loop |
| serving | Runner v0.5.7 or later |
| licence | Apache-2.0 |

Links

EDIT: reading the comments, i should have said this first. this is not a general drop-in for Muse or a gemma4 replacement, and it was never going to be on my compute (98M distillation tokens vs the trillion a real distill wants, i simply lack the compute). the purpose was more on getting the tool calling right. IQ4_NL mix (7.6 GB) is up now too, it scores the same 57/60 on the tool tasks but does not fit an 8 GB card whole (50/52 layers on a 3070, ~5 tok/s).

▲
0
 
10👁
r/LocalLLaMA · u/TheyCallMeDozer · 11d ago
Guy Build a MMORP using Claude... what would it take to do local

As usauly i was scrolling around YouTubes while .... well when every man scrolls YouTube to pass the time.... anyway, came across this video - https://www.youtube.com/watch?v=doR2RhsneRA

TLDR: Guy spends $2175 USD and over 36 hours using Claude Opus 5.5 to build a pretty impressive MMORPG.

Now there is alot of caviats, is it perfect... No... is it really an MMO ... no i havent seen any code added for it.. buttt the strcuture is there, its a hell of a start.

And it got me thinking, if people with their local AI's where to do something like this What models or infra would you use to do this.

Me I think you could get a really good start with Hermes, Qwen Flash, GLM5.3 across a couple of DGX Sparks or if you had 2.5 TB's of RAM and freetoken Kimi K3.

And with the detailed level of prompting and design he laid out prior to actaully letting claude have at it, i think it would doable locally.

So to start the Discussion, what would your tech stack be to do this locally? For me:

\- 2 x DGX Sparks - GLM 5.3

\- 5090 desktop with ComfyUI and a bunch of work floors for image generation, Guassplating, Image to 3D models

and to me I think that would be all that would be needed to get started, but im intrested to see what others think up

▲
3
 
12👁
r/LocalLLaMA · u/Musicheardworldwide · 11d ago
What can I run comfortably?

So I put this computer together after I saw a lot of others posting what they have, and I wanted to know where this ranked and what (in your opinion) are the best local models for coding, and for always on agents.

I didn’t want to type out the whole thing, so I asked my model to give me the specs.

Creature workstation (kernel 7.0.0-28-generic) with an Intel Xeon E5-2698 v4 at 2.2 GHz (20 cores / 40 threads, 50 MiB L3), 125 GiB DDR4-2400 (90 GiB free right now), one NVIDIA RTX PRO 4500 Blackwell with 32 GB of GDDR7 / 31.9 GiB usable VRAM (896 GB/s), and 5.37 TB of raw storage — a 915 GB NVMe root (379 GB free), a 916 GB media disk, and two 1.8 TB drives.

All figures read live from lscpu, free -g, nvidia-smi and df just now.

▲
0
 
13👁
r/LocalLLaMA · u/challis88ocarina · 11d ago
PSA: exercise caution when comparing t/s among models and servers

A "token" is not a fixed chunk of text. It's a word from the model's own private dictionary. Each model ships with its own vocabulary: the list of string-slices learned during training. Analogy: two people transcribe the same sentence: one writes "New York" as one word, the other as two. Both are correct; they're just counting different things.

So tokens/sec is speed measured in steps per minute, and two models can have different stride lengths. One can takes long steps (eg, 4.83 chars each), the other short ones (eg, 3.27). A child and an adult both walking "60 steps per minute" are not walking side by side.

The rules that follow:

  1. Same tokenizer = fair comparison. Two llama.cpp servers running the same model family, tok/s compares directly. Trust it.
  1. Different tokenizers = the number is in different units. Convert to distance: real speed = tok/s x chars-per-token. If both servers report 40 tok/s on prose, one server is laying down \~131 chars/s and the other server \~193 chars/s, so the second is 1.48x faster while the headline numbers tie. The inflation favors the choppier tokenizer: more tokens for the same text = bigger tok/s for the same wall-clock speed.
  1. The conversion factor is content-dependent, so measure, don't assume. A ratio may be 0.68 on prose but 0.74 on JSON. The same vocabularies chop different text differently. Any cross-server speed claim should come with chars/token (or just words/sec) measured on a representative workload, not vendor marketing numbers.
  1. One is real, one is an illusion. What's real is prefill and cost. A more efficient tokenizer turns the same conversation into fewer tokens, so there is genuinely less compute before the first token (shorter TTFT). That's a true speed and money win, not a units trick. What's an Illusion: decode-rate comparisons across families. "Model A does 80 tok/s, model B does 60" says nothing about who finishes the answer first unless A and B share a vocabulary.
▲
1
 
10👁
r/LocalLLaMA · u/DimeRhyme · 11d ago
An UltraFast Qwen3.8 Flash recipe: 74 tok/s, 212 tok/s aggregate on one DGX Spark post image

TL;DR: vLLM recipe for Qwen3.8 Flash on one DGX Spark / GB10. 74 tok/s peak single-stream, 60 to 70 on normal requests, 212 tok/s across 8 streams, 2x to 3.4x faster cold prefill than the recipe it's forked from, full 262K context, and quality matches the original within noise. Everything is open, including the raw per-round data and the benchmark scripts.

https://github.com/dime-online/qwen3.8-Flash-DGX-UltraFast

Most single-Spark setups I've seen posted for this model land somewhere in the 35 to 45 tok/s range, so I spent a few weeks figuring out where the time per token actually goes on a GB10 and cutting it down. If you don't have a Spark, the tricks in the middle section should still be interesting, since most of them apply to any MTP or speculative decoding setup.

What this is, in plain terms

It's a ready-made serving setup. You build the container, pull the public weights, and get an OpenAI-compatible server that answers a lot faster on one box. The speed comes from the model's own draft head guessing several tokens ahead while the full model checks all of them in one pass. Right guesses give you several tokens for the price of one step, and wrong ones get replaced by the full model's answer, so output quality doesn't change.

Decode

Peak decode speed, same workload at every point, best of 3 rounds:

|Streams|1|2|3|4|5|6|7|8|
|:-|:-|:-|:-|:-|:-|:-|:-|:-|
|tok/s|74.1|110.0|132.9|155.5|175.8|191.5|205.8|212.2|

Tokens per step stays between 3.65 and 3.94 from 1 all the way to 8 streams, so the speculation doesn't fall apart under batching. 8 is where it tops out because that's the configured max\_num\_seqs, and the gain from 7 to 8 was down to 3%.

Prefill

Cold prompt with nothing cached, three repeats each:

|Prompt|This recipe|Original recipe|Increase|
|:-|:-|:-|:-|
|16K tokens|4,016 tok/s|1,171 tok/s|\+243%|
|64K tokens|2,426 tok/s|1,071 tok/s|\+127%|
|128K tokens|2,213 tok/s|1,065 tok/s|\+108%|

With prefix caching on, a cached coding prompt starts replying in about 0.57 s, which is what makes agent loops feel fast.

What actually made the difference

The model's own MTP head, run densely, lands about 3.7 tokens per verify step. That's the single biggest lever.

I cut the draft head's vocab from 248K to 65K ids. On a GB10 the draft pass is memory-bound, and reading a full-vocab head every draft step was a real chunk of the step time. The target model still verifies against the full vocab, so this can only change speed, not output.

The quant is W4A16 AutoRound for the MoE experts, FP8 for the side layers and INT8 for the lm\_head. No 3-bit and no NVFP4, because I wanted the speed to come from the serving path and not from squeezing the weights harder.

There's a GB10-tuned low-latency GEMM for the small decode-time matmuls and a sort-free top-k in the verify step.

The prefill gain mostly comes from a faster gather path for the per-layer embedding table, which removes a pile of serial page faults during prefill.

Put together, each decode step went from 68.3 ms on the original recipe to 52.3 ms on an agent-shaped coding workload, about 1.3x more steps per second.

One thing that didn't pay off: doubling the prefill chunk to 16,384 tokens gave no prefill gain at all and ran the box low enough on memory that I rejected it.

Quality

93.1% and 93.3% on a fixed 492-question suite over two seeds, covering code with execution checks, math, knowledge, instruction following, tool calls and long-context needles. I also ran a teacher-forced check against the original checkpoint, and top-1 agreement moved by 0.06 points against a 0.15 point noise band I set before running it.

Practical stuff

The model takes about 71 GiB, the KV pool is 16 GB, and around 16 GiB stays free under load. While generating, the GPU draws about 35 to 37 W median and peaks near 70 W on long cold prefills, with no power or thermal throttling across the soak runs.

The 65K draft vocab was built from English and code, so Chinese, Japanese and Korean output drafts less well and runs slower. Quality isn't affected, because the full model still checks every token.

Built on Saren-Arterius's qwen3.8-Flash-DGX-AutoRound, so big credit there. Happy to answer questions.

💬 13 (+3) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/Robert__Sinclair · 11d ago
The next big company will be...

...the one that will mass produce a cheap device (sub $1000) able to run current sota models at a decent speed.

for that to happen, obviously RAM has to be cheaper, models need to be more efficient and CPUs have to change. It will take time. But as in the 70s/80s computers were huge and expensive mainframes only big companies had, today we are in the same situation.
Fortunately progress happens faster now, so it won't take 30 years to get affordable "home computers". Probably 10, hopefully less.
It's encouraging that today to run the latest qwen 27B you can spend less than $3000. But still...

▲
8
 
14👁
r/LocalLLaMA · u/Dismal-Effect-1914 · 11d ago
B70 no stock/Price Increase

I bought a B70 off Amazon last week for 1300 and have been playing around with it. Today I checked and it seems like I cannot find a single one online for less than 1600 and most places dont have them in stock anymore? What happened? What drives these sudden price increases? It seems like all GPUs across the board have seen another dramatic price flux. Some 5090s I saw were going for 10k!?

▲
28
 
8👁
r/LocalLLaMA · u/jjusko20 · 11d ago
Watch me post-train AliceAI-Foundation-80B-A3B from base to instruct at home, live, on my V100s!

No click bait baby I promise - I'm live streaming the training process kinda like MiMo.

UPDATE: \[Training is paused for an hour or two\] back to training in batches. u/FullOf_Bad_Ideas has pointed out to me I'm burning a ton of compute for nothing on sequence lengths - we'll be breaking the run up into 7/8 batches and then going again.

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wpg4a8/im\_trying\_to\_posttrain/

If you didn't see my original post a few days ago, I'm attempting a slightly more ambitious than usual project in trying to create at least a rough AliceAI-Foundation-80B-A3B-Instruct

I spent the weekend distilling my initial instruct training dataset out of Qwen 3.8 27b, medium thinking - intentionally done because I can run it locally, and I wanted a full dataset in some reasonable amount of time. Still took my v100s running 4x instances at 25tps, like 96 hours of non stop generation to complete the dataset.

I opted not to go for a pre-existing public dataset because I wanted to practice building my own distillation engine (which was configured to work off of an OpenAI compatible endpoint, so it'll distill anything you can hook it up to). The final dataset (this time) consists of 3340 samples: 1760 of general instruct transcripts, and 1580 agentic specific work rows about SWE, harnesses, terminals, etc - I gave the distilling engine a python sandbox and got to simulate turn driven development with a user, and I trained for a bunch of different harness syntax for tool calls, which hopefully will be enough to generalize - gonna run 2 epochs at first.

My GPUs are sobbing right now - turned them on on Friday and left for a weekend vacation, got back today, waited an hour for the data to finish generating, and then immediately fired up the train.

The stuff above is the short version. I'm guessing the initial SFT train will take about 3-6 days, and I plan on working on a RL implementation after I'm satisfied that the SFT has at least worked properly. I am training a rank 16 QLoRa adapter on only q/k/v/o proj, no direct knowledge weight fine tuning.

I thought what MiMo did with their recent training was really cool to watch online, and I like sharing my work with like minded people, and frankly, there's a part of me that's hoping someone will see this and want to hire me (looking for NYC work if you know anyone looking for some passionate ML engineers!) - so I've set up my own little training stream on a cloud flare tunnel.

The stream has the live in progress status of the train, including a live view of the actual data being processed by the model. It also includes way more detail about how I actually designed and generated my training data. Happy to throw the full set on HF as well. I don't expect this model to beat any existing standards but I'll be curious to see if I can get it to operate properly in a harness so I can formally bench it.

I hope you find this interesting! The live stream is a self updating website where you can see exactly what's happening - no need to reload. To watch the training live, visit https://figure-bios-expect-cio.trycloudflare.com/ \[i am currently fixing training issues but it'll be back asap\] -- I'll be keeping it up until the initial SFT is done, at least. The stream lets you inspect the training live as well. This is just a cloudflare tunnel to the trainer.

3 hour update? Loss started at 9ish and is bouncing near 3/4

Update today: back online

▲
0
 
12👁
r/LocalLLaMA · u/cortexist · 11d ago
A hybrid model of Gemma4 with a JEV-like decision head in multi-speaker voice conversation post image

The human brain is neither an LLM nor a JEV. In a crowded market you hear a lot of speech and answer almost none of it. The ongoing question is not “what should I say?” It is “you talking to me?" and "should I say anything at all?”

This live voice demo showcases three hardware tiers—the Blackwell 4500, Jetson Orin NX 16GB, and Jetson Orin Nano 8GB—solving this exact problem. By splitting the workload between a lightweight decision head for turn-taking and a Gemma 4 pipeline for text generation, the setup delivers highly responsive, low-latency vocal interaction.

EDIT: repo (the latest code yet published) https://github.com/cortexist/little-gemma

▲
0
 
14👁
r/LocalLLaMA · u/rawdikrik · 11d ago
Argue with each other for my edu-tainment - 5070 + 5060ti OR RX 7900XTX

I run an Unraid server and use local models for STT, memory, a small llm (I like the new swift bonsai), and SystemOne Models. I currently have a 5070 plugged into my x570 board (with a 5600x), and then the 5060ti on a riser. The 5060 runs at x4, there is a limitation on the board setup.

Running a model big enough for both cards runs SLOW, since the connection to the 5060ti is capped.

Ive tried optimizing with ninfer and vllm, but my speed is capped at the hardware level.

I am considering scrapping the 2 card setup for a single card, and right now the best budget option is the RX7900XTX.

Can you guys argue about what would be the better setup? I dont need the newest models, and I dont need the most speed. I pay for online models. I just like to have a bit of local stuff to help with server stuff. I feel like I spend too much time managing the 2 card setup for not enough to get out of it, and I think dropping to the one card would make things easier without a loss in speed. The idea is to sell both NVIDIA cards, and just get a single card with big enough memory that isnt a pig.

Any advice would help.

▲
0
 
8👁
r/LocalLLaMA · u/inawhole · 11d ago
Gevva0 - a Jev like decision engine on Gemma 26B via direct logit scoring

On the official JevBench evaluation battery, Gevva0 scored 74.63 (#1 global rank), averaging 214ms p50 across large legal contract sets with 82.9% accuracy on the forensic hard tier.

How it works under the hood:

  1. Direct Logit Scoring: Ingests context and reads decision logits directly from llm.scores\[-1\] in a single prefill pass. Fast-path resolution runs in 22ms on short contexts.
  2. Cyclic Debiasing: Permutes class tokens across 4 cyclic positions to neutralize label position bias.
  3. Platt Temperature Calibration: Fits confidence via sigmoid scaling to push Expected Calibration Error (ECE) below 0.03.
  4. Asymmetric Audit Pass: Locks the categorical verdict permanently first, then performs an isolated extraction pass to retrieve verbatim source quotes without contaminating the decision logit.

The repo includes the evaluation harness, raw benchmark datasets, and a local web dashboard: https://github.com/solvingSteve/Gevva0

Setup instructions and benchmarks are in the README.
Working on Demos and Use Cases now so if you have any ideas I'll try to build them next!

▲
12
 
12👁
r/LocalLLaMA · u/Merchant_Lawrence · 11d ago
Need small model that can work for tool caling and agent

Hi. So....... after toturing my 750 ti 4 gb and 16 gb ram with image gen model .i want continue experiment with agent mode like hermes or opencode, using local model but before go i want ask few question. are big model = good perfomance or small model can do same stuff. what small model recommend for agent my spec what caveat of doing this ?

▲
0
 
13👁
r/LocalLLaMA · u/Arany8 · 11d ago
X account claims high t/s setup, but thin on details

According to this post it is possible to reach very high numbers using mtp, however I have failed to reproduce the 50+ tps for 5060ti.

Am I just ignorant or how exactly do this? Or is this a fake post?
Freshly built llama fork for sm120 (Blackwell):
https://github.com/Anbeeld/beellama.cpp
"C:\\llama\\build\\bin\\llama-server.exe" \^

\-m "%MODEL%" \^

\--port 8090 --host 127.0.0.1 \^

\-ngl 99 \^

\--cache-type-k kvarn3 --cache-type-v kvarn3 \^

\--flash-attn on \^

\--load-mode mlock \^

\--jinja \^

\-c 98304 --parallel 1 \^

\--fit off \^

\--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-ubatch-size 128 \^

\-ctkd q8\_0 -ctvd q4\_0 \^

\--kv-tail-tokens auto

Runs at 20-35 t/s.

▲
0
 
12👁
r/LocalLLaMA · u/itsthewolfe · 11d ago
What is the current recommended local model for general use (96GB).

I'm setting up my first build with Open Claw. I'm new to ask of this and starting from zero knowledge.

I've done a lot of reading up, but it's a little overwhelming. So I'm biting things off in chunks.

I want everything to be local. I have a mini PC with 96GB of RAM so can fit a good sized model.

I have Open Claw set up right now with OpenRouter.

My next step is to set up my local model.

What is the current leading open source model for generic tasks and learning? I have plenty of memory to support.

Kimi K3, Opus, Quen 3.8, other?

▲
0
 
15👁
r/LocalLLaMA · u/forevergeeks · 11d ago
Would you buy an AI appliance that removed all the hard work for you

Would you buy an AI appliance that made it easier for you to run local AI models such as Qwen 3.8 27B and Gemma 3 27B?

By easier I mean, the appliance will take care of all the infrastructure stuff for you such as installing the OS, the inference engine such as llama.cpp or vLLM, access management and perhaps include aome preconfigured agents for you start using the system.

The system is multi-user, with a role-based management system, meaning multiple people can use it, including teams.

All runs local, but with the option of using cloud based models if you need more horse power.

Is this something that has an appealing?

Or the fun is the tinkering 🤪

▲
0
 
13👁
r/LocalLLaMA · u/TangeloOk9486 · 11d ago
What models can I run locally on a Mac mini m4 32GB

Hey guys i am planning to get a mac since the GPU and other stuff isnt currently possible for me rn so what models near or stronger than Sonnet 4.6 or somewhat nearer can I use it on mac or would i be able to use it actually?

I mainly need it for coding tasks, different file management and reports and also reasoning. SUggestions or feedbacks are welcomed for models. I want everything local for privacy concerns

Edit: Fixed the model mention

▲
7
 
13👁
r/LocalLLaMA · u/js1943 · 12d ago
LM Studio vs Bionic

I am confused between LM Studio and Bionic.

I have LM Studio for a long time though not used frequently.

Recently I am trying to learn the agentic stuff. Watched a few videos but they were all using Bionic. The strange thing is the interface looks different than the one I just installed today. (Mine seems to be missing features, no developer mode. I am on MacOS)

On the other hand, I seem to be able to find those missing settings in LM Studio.

So what is the difference between the two? Is there anything Bionic can do but LM Studio doesn't?

💬 9 (+5) open on reddit ↗
▲
9
 
10👁
r/LocalLLaMA · u/segmond · 12d ago
Anyone customizing and Optimizing llama.cpp per model?

Basically the idea is take your favorite model, for example qwen3.8-27b or say dsv4vision. Strip everything out that is not needed by that model so the only thing needed is just for the model. Optimize the remaining code to be fast. The idea is to have a model also do this, provide it with enough tools, prompts, docs, guidance. I reckon that if we have a llama.cpp that is optimized for just one model architecture without all the cruft needed to run and. handle other models, that it would not be surprising to easily see 2x+ performance improvement. Anyone thinking along this idea? Again, the goal will be to give this task to a smart model, let it run in a loop, and after a week or 2 you hopefully end up with llama.qwen3.8-27b or llama.glm5.3-flash that would fly.

💬 42 (+1) open on reddit ↗
▲
0
 
6👁
r/LocalLLaMA · u/rm-rf-rm · 12d ago
llama.cpp MacOS menu bar app using blobs instead of GGUF files

Recently started using the MacOS menu bar app for llama.cpp available at https://llama.app When you use the UI to download a model, it seems to do an Ollama-esque hashing instead of just saving the GGUF. Even if you put a GGUF in the Model directory folder, neither the menu bar UI nor the web UI recognizes it. https://preview.redd.it/bsiw7kbn46sh1.png?width=1186&format=png&auto=…

▲
1
 
2👁
r/LocalLLaMA · u/textclf · 12d ago
Introducting TextCLF Quant Factory

Hello, I created a calibration free quant method called TQ. It doesn't need any data so models could be quantized as soon as they come out and the quantized models would generalize better. It performs closely to calibration-based methods. For example, for Qwen 3.8 27B the 4-bit TQ has mean KLD of 0.0282 and top-1 of 92.4% I opened sourced the quant code as Quant Factory so anyone can quantize and run open source models. Right now it only supports 4-bit but I plan to add support for 2-bit and 3-bit soon. The repo link is: https://github.com/textclf-api/quant-factory I have a collection of quantized models using TQ at: https://huggingface.co/textclf You can run these models using either using the following docker image or by following the repo's instruction. For example you can run textclf/Qwen3.8-27B-TQ-4bit like this: docker run --rm --gpus all -p 8000:8000 docker.io/textclf/tq-quant:4bit-main vllm serve textclf/Qwen3.8-27B-TQ-4bit --quantization tq The Dockerfile in the repo shows how this docker image was created. The repo's README explains the approach used for this quant method and why it is useful. You can try it and let me know what you think. Feedback appreciated. EDIT: I did KLD testing using the Wikitext-2 dataset for Qwen 3.8 37B. I also did the same test for the Unsloth-UD-Q4\_K\_XL quant. Here is what I go: |Quant|Disk Size without MTP (GB)|Mean KLD|Median KLD|99% KLD|Top 1% Agreement| |:-|:-|:-|:-|:-|:-| |TQ 4-bit|17.76|0.02823666|0.01282929|0.26897613|92.419%| |UD-Q4\_K\_XL|17.59|0.00771805|0.00318557|0.07390548|95.779%| Working on getting more tests for other models!

▲
0
 
8👁
r/LocalLLaMA · u/Truth-Does-Not-Exist · 12d ago
is DDR5 a scam? My $350 2007 Dell Precision is destroying my $1500 2025 RTX 5070 rig in agentic tasks. made possible by Prism32

I had a theory that ram speed didn't really matter and the only thing that matters is your GPU capacity, vram speed, vram size, and system ram size instead of ram speed or cpu speed. I think I've been vindicated. I used unsloth/Qwen3.8-27B-GGUF:UD-IQ3\_XXS (10.9gb) https://huggingface.co/unsloth/Qwen3.8-27B-GGUF as my baseline and mtp q4\_0 1.37gb only on the dual gpu systems because mtp was too slow on the 9060xt and 5070 I used llama.cpp for all of them and tried to go for the max context I could since they are supposed to be day to day agents. I picked prism32 as my agent harness for this because it's the most compatible, fastest, and reliable one I've found. It's a custom architecture https://github.com/MegaDyneSystems/prism32 which gives it some massive advantages, especially if you want to avoid bloated frameworks eating your context or CPU. The 2007 system literally doesn't work with any other harness because they all require sse4.2 and a ton of heavy dependencies. prism32's only dependency is python 3.7 or above. It's so lightweight (uses around 5 to 10mb ram) that I actually run it bare metal on my ARM synology NAS and even my 2008 TP-link router. If you want to run any agents especially with advanced features on edge or legacy hardware without it choking your system their is no competition I tried these 5 systems: 2007 dell precision t5400 ddr2: 24gb ddr2 8 core 8 thread dual xeon x5460, rx 6700xt rx 6700 22gb vram total, 215gb ssd total system memory 46gb (pic 1) 2009 dell precision t5500 ddr3: 72gb ddr3 12 core 24 thread dual xeon x5675, rtx 5060 rtx 4060 16gb vram total, 512gb ssd, total system memory 88gb (pic 2) 2012 dell precision t3600 ddr3: 64gb ddr3 6 core 12 thread xeon e5-1650, rtx 4060, rtx 3060 12gb total vram 20gb, 512gb ssd, total system memory 84gb 2018 hp obelisk desktop 875 ddr4: 32gb ddr4 3200 i7 8700, rx 9060 xt 16gb, total vram 16gb 512gb ssd, total system memory 48gb (pic 3 ) * 2025 HP omen: 32gb ddr5 6000 RTX 5070 12gb GDDR7 vram 1tb ssd, total 44gb memory (pic 4) The speed results on each system were: 2007 dell precision ddr2 144k context, 18 tk's a second decode, 150tk's prompt processing 2009 dell precision t5500 ddr3 256k context 22 tokens a second decode, 224 tokens prompt processing 2012 dell precision t3600 ddr3 104k context 22 tokens a second short context 14 tokens a second long context, prompt processing is 315 tokens, (could optimize further but tests took me long enough) 2018 hp obelisk 875 ddr4 180k context, 16 tokens a second long decode, 477 promp processing 2025 HP omen 131k 13 tokens a second, 43 prompt processing (yes 43) conclusion The older dual xeon dual gpu setups completely destroyed the newer stuff in context length and speed even on worse GPU's which I think proves my theory, The HP omen system is at least $1500 and the 2007 system didn't cost more than 400 total, rx 6700 xt was $190, rx 6700 was $140, on ebay they are overpriced at $150 although I got it $50 second hand in 2014, and the DDR3 systems were $20 second hand and go around 80 to 150 on ebay, I'd say the ddr3 systems are the best for performance and value, next project is running qwen 3.8 flash next on the ddr3 systems

▲
0
 
10👁
r/LocalLLaMA · u/AIFrontierReads · 12d ago
Laya: replace LLM-as-a-judge with a 322M-parameter decision engine (26,639 stars in 9 days, hands-on test)

It turns decisions — routing, triage, yes/no calls — into typed outputs from a small model instead of generated text, with a routing-only CLI, triage presets, and an abstention gate when confidence falls below a threshold. I ran through the tutorial on CPU end to end, including a French ticket classification, and with min\_confidence=0.90 it abstained on one case it would otherwise have misclassified — the honest highlight. Warm latency was about 0.7s per question on CPU; the calibration caveat (over-confident checkpoints) is worth knowing before trusting the scores blindly.

▲
16
 
7👁
r/LocalLLaMA · u/Danmoreng · 12d ago
Gem16 - custom engine for Gemma4 12B & 26B on Blackwell 16GB GPUs

It’s probably a bit niche and the models are a bit old at this point, but after reading about Ninfer a few months ago I did my own small vibe coded engine project for my 5080 Laptop GPU. Initially I thought I can only fit the 12B model with enough context into the VRAM, but with custom quantisation (EXL3 like) the 26B fits nicely as well. The engine is entirely Codex written, but it took a lot of weekends to make it work and make it work as fast as vLLM/faster since vLLM didn’t work with MTP on 16GB VRAM. Also, my engine works on Linux and Windows equally well. Primarily this is designed to be single-user only, the 12B model can serve 2 sessions. It also comes with a native fancy looking GUI, but the main focus was the engine itself. 12B with audio & vision, 5.800 t/s prefill & 87 t/s decode 26B with vision, 5.660 t/s prefill & 182 t/s decode, fits 220k context https://github.com/Danmoreng/gem16 Sadly the most interesting feature of the 12B model with native audio understanding seems to have the issue, that after around 8k context the model doesn’t recognise audio tokens anymore. This seems to be a model issue, as others have also reported it: https://huggingface.co/google/gemma-4-12B-it/discussions/45 Would love to get some feedback!

▲
2
 
2👁
r/LocalLLaMA · u/tabletuser_blogspot · 12d ago
Dual Radeon improved speeds using Vulkan

I've been struggling to keep my Radeon Instinct MI50 GPU cool. I'm looking for budget friendly solutions. While running multiple GPUs it doesn't usually get too hot. I was also getting lower benchmarks using standard Vulkan 'llama 7B Q4\_0' model benchmark, but it wasn't caused by thermal throttling. Time to optimize. MI50 with Radeon VII firmware 16GB Vram My previous post I tested several model using same GPUs. I made some changes. I moved the MI50 16gb into the primary PCIe 16x slot and moved the RX 7900 GRE 16gb into a slower PCIe 4x slot. Overall system inference performance increased. I used Google Gemini to helped my optimize my llama-bench settings and it taught be about: RADV_PERFTEST=nogttspill is an AMD Linux driver flag used when running llama.cpp with the Vulkan backend. It forces the RADV (Mesa Vulkan) driver to prioritize keeping all model allocations inside dedicated video memory (VRAM) rather than spilling over into system RAM (GTT/Graphics Translation Table). \[1, 2, 3\] I saw llama 7B Q4\_0 score jump back to where is it should be. So I tested a few other models. GGML_VK_VISIBLE_DEVICES=0,1 RADV_PERFTEST=nogttspill time /llama-b11053/llama-bench -fa on -ngl 99 -m /llama-2-7b.Q4_0.gguf Previous benchmarks: https://www.reddit.com/r/LocalLLM/s/PT6Bd5nUUE see end for comparison The list has been sorted by pp512 improvement in descending order (highest gain to highest loss). |Model Name|Improvement (pp512)|Improvement (tg128)| |:-|:-|:-| |Laguna-XS-2.1-APEX-I-Balanced.gguf|\+402.14%|\-0.20%| |gemma-4-31B-it-Q6\_K.gguf|\+397.84%|\+9.24%| |Muse-Glimmer-30B-UD-Q6\_K\_XL.gguf|\+386.58%|\+48.63%| |llama-2-7b.Q4\_0.gguf|\+176.57%|\+20.29%| |Qwen3.8-27B-Q6\_K.gguf|\+48.13%|\+0.07%| |medgemma-27b-it-UD-Q6\_K\_XL.gguf|\+37.89%|\-0.52%| |Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf|\+1.17%|\+0.49%| |Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4\_K\_M.gguf|\+0.68%|\+4.16%| |NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q5\_K\_M.gguf|\+0.28%|\-4.28%| |granite-4.2-30b-Q6\_K\_L.gguf|\-0.12%|0.00%| |GLM-4.7-Flash-Uncen-Hrt-NEO-CODE-MAX-imt-D\_AU-Q6\_K.gguf|\-0.63%|\-7.83%| |Qwen3-Coder-30B-A3B-Instruct-UD-Q6\_K\_XL.gguf|\-0.69%|\-0.86%| |Accio-Lab\_occamy-1.0-Q5\_K\_S.gguf|\-0.81%|\+0.17%| These are the models tested in the same order as tables below: GGUF Model List (in order): 1. Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf 2. Accio-Lab\_occamy-1.0-Q5\_K\_S.gguf 3. Laguna-XS-2.1-APEX-I-Balanced.gguf 4. NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q5\_K\_M.gguf 5. gemma-4-31B-it-Q6\_K.gguf 6. Qwen3-Coder-30B-A3B-Instruct-UD-Q6\_K\_XL.gguf 7. GLM-4.7-Flash-Uncen-Hrt-NEO-CODE-MAX-imt-D\_AU-Q6\_K.gguf 8. granite-4.2-30b-Q6\_K\_L.gguf 9. Muse-Glimmer-30B-UD-Q6\_K\_XL.gguf 10. Qwen3.8-27B-Q6\_K.gguf 11. medgemma-27b-it-UD-Q6\_K\_XL.gguf 12. Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4\_K\_M.gguf 13. llama-2-7b.Q4\_0.gguf Supporting data sorted order (by Params descending, then Size descending). All models running dual Radeon GPU, Vulkan backend, and flash attention on. # Table 1: RADV_PERFTEST=nogttspill is being used |model|size|params|pp512|tg128| |:-|:-|:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|24.76 GiB|34.66 B|1252.09 ± 11.84|53.54 ± 0.51| |qwen35moe 35B.A3B Q5\_K - Small|23.40 GiB|34.66 B|1245.11 ± 11.76|53.42 ± 0.28| |laguna 30B.A3B Q5\_K - Medium|22.64 GiB|33.44 B|1027.42 ± 11.32|65.72 ± 0.51| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|25.10 GiB|32.91 B|1120.31 ± 12.85|63.82 ± 1.52| |gemma4 31B Q6\_K|23.46 GiB|30.70 B|200.33 ± 0.20|15.61 ± 0.05| |qwen3moe 30B.A3B Q6\_K|24.53 GiB|30.53 B|1176.48 ± 36.62|66.04 ± 0.25| |deepseek2 30B.A3B Q6\_K|23.26 GiB|29.94 B|914.86 ± 8.75|40.02 ± 0.12| |granite 30B Q6\_K|23.01 GiB|29.28 B|92.09 ± 0.30|17.41 ± 0.03| |muse-glimmer 30B Q6\_K|24.45 GiB|27.85 B|279.04 ± 0.20|17.42 ± 0.03| |qwen35 27B Q6\_K|21.30 GiB|27.32 B|235.12 ± 1.05|14.42 ± 0.03| |gemma3 27B Q6\_K|22.09 GiB|27.01 B|247.18 ± 0.54|17.30 ± 0.13| |gemma4 26B.A4B Q4\_K - Medium|15.63 GiB|25.23 B|1316.62 ± 24.18|55.56 ± 0.16| |llama 7B Q4\_0|3.56 GiB|6.74 B|1349.35 ± 15.66|74.62 ± 0.37| # Table 2: RADV_PERFTEST=nogttspill is NOT being used |model|size|params|pp512|tg128| |:-|:-|:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|24.76 GiB|34.66 B|1237.55 ± 25.29|53.28 ± 0.24| |qwen35moe 35B.A3B Q5\_K - Small|23.40 GiB|34.66 B|1255.24 ± 6.03|53.33 ± 0.48| |laguna 30B.A3B Q5\_K - Medium|22.64 GiB|33.44 B|204.55 ± 1.66|65.85 ± 0.29| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|25.10 GiB|32.91 B|1117.17 ± 10.11|66.68 ± 0.24| |gemma4 31B Q6\_K|23.46 GiB|30.70 B|40.22 ± 0.12|14.29 ± 0.03| |qwen3moe 30B.A3B Q6\_K|24.53 GiB|30.53 B|1184.70 ± 27.98|66.61 ± 0.27| |deepseek2 30B.A3B Q6\_K|23.26 GiB|29.94 B|920.67 ± 5.82|43.42 ± 0.09| |granite 30B Q6\_K|23.01 GiB|29.28 B|92.20 ± 0.21|17.41 ± 0.04| |muse-glimmer 30B Q6\_K|24.45 GiB|27.85 B|57.15 ± 0.87|11.69 ± 0.09| |qwen35 27B Q6\_K|21.30 GiB|27.32 B|158.73 ± 0.57|14.41 ± 0.41| |gemma3 27B Q6\_K|22.09 GiB|27.01 B|179.26 ± 0.44|17.39 ± 0.02| |gemma4 26B.A4B Q4\_K - Medium|15.63 GiB|25.23 B|1307.73 ± 23.51|53.34 ± 0.22| |llama 7B Q4\_0|3.56 GiB|6.74 B|487.92 ± 37.05|62.03 ± 0.60| Swapping PCIe locations for the Radeon Instinct MI50 and Radeon RX 7900 GRE and using RADV_PERFTEST=nogttspill flag resulted in improvements over my first baseline benchmarks. Note: Only models appearing in both datasets are listed. The table is sorted by Parameters in descending order. |Model Name|Improvement (pp512)|Improvement (tg128)| |:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|\+221.14%|\+33.85%| |qwen35moe 35B.A3B Q5\_K - Small|\+215.17%|\+1.25%| |laguna 30B.A3B Q5\_K - Medium|\+537.77%|\+17.25%| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|\+267.75%|\-3.90%| |gemma4 31B Q6\_K|\+17.00%|\+23.60%| |qwen3moe 30B.A3B Q6\_K|\+211.32%|\-0.97%| |deepseek2 30B.A3B Q6\_K|\+191.00%|\-14.54%| |muse-glimmer 30B Q6\_K|\+12.49%|\+0.17%| |qwen35 27B Q6\_K|\+17.00%|\-6.79%| |gemma3 27B Q6\_K|\+14.83%|\+1.71%| |gemma4 26B.A4B Q4\_K - Medium|\+158.79%|\+7.28%| Looks like MoE models benefit the most. Muse-glimmer 30B Q6\_K didn't real see much improvement but it seems to be the most optimized dense model. MI50 continues to impress. I purchased them used at $150 each. I can now run models 30B to 35B using Q6 quant with long content and decent speed.

▲
9
 
8👁
r/LocalLLaMA · u/Informal-Trouble2183 · 12d ago
Hardware Roofline Inference Calculator post image

Hello everyone, I made a calculator for the theoretical HW roofline for decoding / prefill based on several parameters (LLM model architecture, quants, GPU, Memory, ..). It still a theoretical bound, but helpful as a step-0 check to understand what fits (would fit) in your hardware, and understand the effects of the contributing knots. I hope it helps. You can access it from here: https://www.ai-leaderboard.dev/ (click HW Roofline)

▲
22
 
14👁
r/LocalLLaMA · u/Kmic68 · 12d ago
2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0 post image

Hey guys! I have been excited to share this here. This is a project consisting of kernel optimizations for the Tesla p100 series graphics card ($80). I want to start by saying I am 17 years old and do not have a formal degree. I used Ai for a lot of this and while I understand some, I do not understand everything. Notes: My gpus are capped at 175w/250w each so these numbers may be able to be pushed higher. I also experience minor thermal throttling and sit at a nice toasty 79 degrees, which definitely effect numbers (the table above is while hot, so if you have good cooling expect 5-10% more on prefill and decode). I am using gen3 pcie with two x16 slots. Also, for anyone curious, decode numbers depicted in image were averaged from a list of questions ranging from creative writing and coding. Improved: tps went from 7-15tps at 0 context to 50-60, 260k context went from 2-4tps to 30-35, prefill went from 220 tps at 0 context to 350, 260k context went from 40tps (as far as i remember, i never really measured cause it was too hard) to 110tps, fixed fp16 math errors by using some mixed fp16/fp32 math operations so rounding errors were eliminated, and merged as of sept 22 so it should support qwen 3.8 flash architecture This setup is somewhat flag specific (ie: (-c 262144 -b 32768 -ub 1024 -np 1 \\) without -b 32768 mtp becomes overloaded and drops acceptance to near 0 at full depth) so keep that in mind while setting up. One more thing, I took regression very seriously in this. Math had to be more accurate or byte identical or it would fail tests. Build, flags, math proofs, and anything else you may need will be linked below. Enjoy guys! I would love your feedback on this and am looking at pull requests. Github: https://github.com/Kmic-68/llama.cpp/tree/p100-optimizations Details for build, math proofs, etc: https://github.com/Kmic-68/llama.cpp/tree/p100-optimizations/p100-docs

💬 25 (+1) open on reddit ↗
▲
0
 
14👁
r/LocalLLaMA · u/spammmmmmmmy · 12d ago
Identifying whether a command changes something or is just investigative

I am busy working away on a tool-calling sandbox. RIght now I'm thinking of building a kind of dataflow analyzer for shell commands, so that I can identify source and sink points, and establish whether the command is a readonly command or a command that changes state. Example: Command: \sed -n '124p' webroot/a-file.html | od -c | head -5 \ sed is a function that can read or write. in \sed -n 999p filename\ syntax on my system, it is a readonly operation \|\ is a left to write data flow transfer operator \od\ is a readonly sink and would be on the readonly whitelist * \head\ is a readonly sink and would be on the readonly whitelist. Therefore, I can conclude that this function is readonly and I would allow it automatically in my solution. Whereas, \sed -n '124p' file > /tmp/foo\ or \sed -n '124p' file | visudo\ would be identified as write commands. Before I get deep into this project, I'd like to know if an existing library already has this as a design goal?

▲
0
 
14👁
r/LocalLLaMA · u/Foxiya · 12d ago
Soap Dispenser Benchmark!

Prompt: Create an animation showing how the soap dispenser mechanism works in one complete html file. Results: Opus 5.5 - High: https://reddit.com/link/1wrsqbj/video/3t6nuj2134sh1/player DeepSeek V4.1 Flash: https://reddit.com/link/1wrsqbj/video/m8i5hsl434sh1/player Qwen 3.8 Max: https://reddit.com/link/1wrsqbj/video/ukkl5qp734sh1/player ChatGPT 5.6 Sol - High: https://reddit.com/link/1wrsqbj/video/ia840utb34sh1/player Opus 5 - High https://reddit.com/link/1wrsqbj/video/hzmnavof34sh1/player

💬 22 (+1) open on reddit ↗
▲
0
 
6👁
r/LocalLLaMA · u/muthuishere2101 · 12d ago
I built a CLI for Jev-style typed decisions that can also run with local models

I wanted a simple way to use small models for tiny decisions without wiring them into a full LLM app. So I built jevx. It gives you a CLI for jev and jev based models and connect it from the terminal, shell scripts, CI, or from agents like Claude Code and Codex. https://muthuishere.github.io/jevx/guides/scenarios/ https://github.com/muthuishere/jevx

▲
0
 
11👁
r/LocalLLaMA · u/Maasu · 12d ago
Which Local Models are the least 'Claude' sounding

In your experience, which models sound the least like Claude and more like grok or the gpt's? I cannot stand talking to Claude, to the point I have all requests proxyed through other agents to it. I have been using qwen3.8-27b locally and a heavily quantised version of deepseek v4. I love both for their capabilities, as I did claude to be fair, but I hate interacting with them directly. So right now I mostly interact with SOL 5.6 or Luna Max and have them orchestrate (using setup similar to first mate that i put together myself). I appreciate both models have been distilled on anthropic models, but I'd love to eventually one day be fully reliant on local models but this is one of the last blockers for me. So I thought it'd be an interesting discussion point, most local ones I have tried I find are very similar to claude in tone. Hardware: bosgame Strix Halo, 128 gb unified ram.

▲
8
 
8👁
r/LocalLLaMA · u/marcobaldo · 12d ago
Qwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max

Hi! I'm the author of MoEspresso, which is my way of putting my own ideas about inference engines to the test. A lot of the fun has been trying different design choices, measuring what happens, and finding that several of them work well together. MoEspresso 3 runs Qwen3.8-Flash-Next on a 2021 M1 Max with 32 GB of unified memory at 12-15 decode tokens per second - provided there are no other memory hungry applications running in the background (such as browsers). During decode, experts which are already resident in memory receive a bias, but the two strongest experts according to the model are always chosen (with the default settings). I wrote about this here https://github.com/steadfastgaze/MoEspresso/blob/main/docs/cache_prior.md - I first started thinking about this after reading about Apple's AFM 3 and the instruction-following pruning work behind it (https://arxiv.org/html/2501.02086v3#abstract), but then I found this other paper (https://arxiv.org/html/2412.00099v2) which spoke about Cache-Prior. Prefill is unbiased. Even with this bias enabled by default, Qwen 3.8 Next scored ahead of Opus 4.8 xhigh and many other strong hosted solutions. Reproducible 48-question setup -> https://github.com/steadfastgaze/MoEspresso/tree/main/docs/benchmark_reproduc…. Overall scores (%) across six categories, including coding, data analysis and math: | Model | Score | |---|---:| | Qwen3.8 Flash (hosted), medium | 89.7 | | GPT-6 Sol, medium | 89.2 | | Claude Opus 5.5, medium | 86.4 | | GPT-6 Sol, low | 85.0 | | Qwen3.8 Flash @ MoEspresso, medium, Cache-Prior 2/2 | 84.3 | | GPT-6 Luna, xhigh | 81.4 | | Claude Opus 4.8, xhigh | 80.3 | | Claude Sonnet 4.6, high | 74.0 | | Claude Sonnet 5, medium | 70.9 | - I use some of Iwan Kawrakow's formats from ik_llama.cpp, with Metal execution through my mlx-iqk library. Most routed projections in this package use IQ2_K, which is not normally supported by either standard MLX or mainline llama.cpp. - KVarN K4/V4 leaves more memory for resident experts as context grows, and it is functioning extremely well with low RMS error on this model architecture. - Good defaults, e.g. automatic SSD streaming and Cache-Prior when all experts cannot fit, with the settings generally following the same rule. This is the third iteration, and I have more concrete ideas to explore, both for squeezing even more performance from Apple Silicon and for bringing the engine to Linux and AMD machines such as Strix Halo. Installation is through Homebrew, so "brew install steadfastgaze/tap/moespresso" Code - https://github.com/steadfastgaze/MoEspresso Model - https://huggingface.co/steadfastgaze/Qwen3.8-Flash-Next-MoEspressoV3 If you try it, I'd love to see your "moespresso speed" output (an intentionally quick benchmark). PS: English isn't my first language and I used an LLM to help refine this post, and AI coding tools for implementation. --- edit: some comments are reporting lower speeds (thank you for doing it) - I will investigate tomorrow and in next days.

💬 23 (+2) open on reddit ↗
▲
13
 
9👁
r/LocalLLaMA · u/Admirable_Reality281 · 12d ago
Xiaomi MiMo 2.6 Flash vs GLM 5.3 Flash

I've seen a lot of conflicting opinions about MiMo Flash, but I haven't tried it yet. How does it compare with GLM Flash for coding work like this? I'm interested in: \- back-end development \- debugging, refactoring, implementing features in an existing front-end codebase \- maintaining Docker images \- troubleshooting DevOps errors Not the silly stuff I see "build me 100 nice-looking webpages" or "make me a Three.js demo". So far, I've been happy with GLM 5.3 Flash. My main frustration is that it sometimes overthinks too much, and once it does, it's hard to steer it back on track. The DeepSWE score of MiMo appears to be a substantial improvement over GLM's, but \- there's no official score from DataCurve \- no amount of consumed tokens to achieve it and in general one benchmark doesn't tell me how it behaves on day to day work. I'd be interested in comparisons from people who've used both.

▲
20
 
12👁
r/LocalLLaMA · u/LH-Tech_AI · 12d ago
[Release] - SupraTTS-0.1-Beta - a tiny 29.6M parameters TTS model

Hey guys! Today, we are releasing SupraTTS-0.1-Beta, a tiny \~29.6M parameters Text-To-Speech model. The audio quality is a bit better than the original Glow-TTS (the architecture our model is using!) while it's keeping the same size. Here are some samples: https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta#samples >Link to the model on HF: https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta I hope you can do something useful with it, e.g. on small edge devices and on CPU. Feel free to give us feedback and ask question. Follow us on HF to not miss the next upgrades of SupraTTS, e.g. better voice quality, multi-language-support, multi-voices support, emotions and speaking styles and ZERO SHOT VOICE CLONING**!! 🤗

▲
0
 
10👁
r/LocalLLaMA · u/PleaseLee · 12d ago
We released VeriLoop E2 (27B, Apache-2.0). The design question behind it: should an LLM be allowed to commit its own state?

We’ve released VeriLoop E2, a 27B model post-trained from Qwen3.8-27B, together with the model weights, evaluation evidence, and a llama.cpp GGUF ladder from BF16 down to IQ1\_M. The model is focused on code agents, mathematics, scientific reasoning, and long-horizon verifiable problem solving. For this post, I’m including the model-side results as well as the local-inference details: the post-training setup, completed benchmark evaluations, quantization measurements, llama.cpp validation, tested hardware, and the scientific-reasoning demo. For the GGUFs, every measured low-bit tier was built directly from the canonical BF16 GGUF, evaluated against the same frozen BF16 logits, and checked with the same paired fidelity protocol. ## The main model: VeriLoop E2 VeriLoop E2 uses VeriLoop-Governed Recurrence (VGR). The basic idea is: Generation and verification should not belong to the same authority. The model proposes, diagnoses, revises, searches, and replans. External evidence decides whether a candidate state is allowed to persist. A candidate is committed only when protected obligations do not regress and at least one evidence dimension strictly improves. Otherwise, the verified incumbent state is retained and the failure evidence can inform the next proposal. We also use this structure during post-training: proposals originating from the same state can be separated by external verification into progress, non-progress, regression, and completion, allowing state-transition quality to become supervision without requiring the model to judge itself. The final post-training mixture contains 1,841,831 records across software engineering, code-agent trajectories, mathematics, scientific reasoning, and verifiable recurrence. Nine completed benchmark evaluations: \- SWE-bench Pro — 76.2% \- Terminal-Bench 2.1 — 88.8% \- Terminal-Bench 3.0 — 29.7% \- Terminal-Bench 4.0 — 37.9% \- DeepSWE v1.1 — 64.6% \- AIME 2026 — 98.3% \- GPQA Diamond — 93.9% \- MathArena Apex 2025 — 89.6% \- SWE-Marathon v1.1 - 45.0% We also publish task-level evaluation evidence rather than only aggregate scores. For evaluations that use VeriLoop Harness, it provides the external execution and evidence-governance layer; the E2 checkpoint remains responsible for proposal generation, problem abstraction, route selection, diagnosis, and replanning. I’m keeping that distinction explicit because the benchmark campaign and the standalone local-runtime checks are not the same measurement. ## The GGUF release The GGUF release spans: Tier |Main size |Reduction vs BF16 |PPL ratio |Mean KLD |Same top-p BF16 |50.113 GiB |— |1.000000 |reference |100% Q8\_0 |26.632 GiB |46.86% |1.000643 |0.002176 |98.815% Q6\_K |20.566 GiB |58.96% |0.999605 |0.004409 |98.204% Q5\_K\_M |18.965 GiB |62.16% |1.004450 |0.006919 |97.251% Q4\_K\_M |18.301 GiB |63.48% |1.004821 |0.009700 |96.786% Q3\_K\_M |16.826 GiB |66.42% |1.004090 |0.014349 |95.919% IQ2\_S |16.799 GiB |66.48% |1.003457 |0.014023 |95.516% IQ1\_M |16.790 GiB |66.50% |1.003191 |0.014357 |95.870% A practical way to read the current trade-offs is: \- Q6\_K — higher-fidelity option with a substantial reduction from BF16 \- Q5\_K\_M — middle ground below \~19 GiB \- IQ1\_M — smallest released artifact \- IQ2\_S — adjacent low-footprint option with slightly lower Mean KLD ### IQ1\_M result The smallest release is VeriLoop-E2-IQ1\_M.gguf: \- 18,028,208,896 bytes \- 16.790078 GiB \- 66.4955% smaller than BF16 \- 5.36 effective BPW \- PPL ratio: 1.003191 ± 0.002175 \- Relative PPL drift: +0.3191% \- Mean KLD: 0.014357 ± 0.001317 \- Same top-p: 95.870 ± 0.220% \- log-PPL correlation: 99.62% One important clarification: this is not a uniform 1-bit model. IQ1\_M is a deliberately mixed-precision artifact: 353 F32 + 1 IQ1\_M + 2 IQ2\_S + 64 Q4\_K + 429 Q5\_K + 2 Q6\_K = 851 tensors The transition from IQ2\_S to IQ1\_M changes exactly one selected tensor: \blk.1.ffn\_down.weight: IQ2\_S → IQ1\_M\ The rest of the protected precision policy remains unchanged. That makes IQ1\_M only 8.633 MiB smaller than IQ2\_S, so we do not present that incremental difference as some dramatic compression breakthrough. What is more interesting to us is that the lower-footprint point still remains inside the frozen quality envelope. Compared with IQ2\_S: \- PPL ratio: 1.003191 vs 1.003457 \- Same top-p: 95.870% vs 95.516% \- Mean KLD: 0.014357 vs 0.014023 \- RMS Δp: 3.691% vs 3.679% So these are neighboring trade-off points rather than a simple “Q1 is universally better than Q2” claim. ## Hardware we’ve tested The measurements and runtime checks reported here were performed on: \- GPU: NVIDIA RTX PRO 6000, 96 GB VRAM, ×1 \- CPU: Intel Xeon Platinum 8470Q, 25 vCPU \- System RAM: 120 GB \- OS: Ubuntu 22.04 \- Python: 3.12 \- PyTorch: 2.8.0 \- CUDA: 12.8 This is the hardware I have directly tested for this release. I’m not presenting it as a minimum requirement, and I’m not assuming identical throughput or memory behavior on other systems. ## Standalone performance I do not have a separate full nine-benchmark campaign with the external Harness disabled, so I’m not going to relabel those benchmark scores as “standalone” results. What is directly validated in standalone local inference is the model/GGUF runtime path itself: \- BF16 → IQ1\_M file size: 50.113 GiB → 16.790 GiB \- IQ1\_M PPL ratio: 1.003191 ± 0.002175 \- IQ1\_M relative PPL drift: \+0.3191% \- IQ1\_M Mean KLD: 0.014357 ± 0.001317 \- IQ1\_M Same top-p: 95.870 ± 0.220% \- IQ1\_M log-PPL correlation: 99.62% \- Stock llama.cpp main-only generation: HTTP 200, non-empty output \- Stock llama.cpp main + MTP generation: HTTP 200, non-empty output \- Fixed MTP validation run: 104 draft tokens generated, 76 accepted (73.0769%) I also do not have a clean, reproducible \llama-bench\ pp/tg table that I’m comfortable publishing yet, so there is no extrapolated tokens/s claim here. The MTP acceptance rate is workload-dependent, and the quantization metrics above should not be read as substitutes for downstream benchmark reruns. ## How we measured quantization loss All quantized tiers were evaluated against the same frozen BF16 reference using: \- WikiText-2 raw test \- context: 2048 \- chunks: 8 \- seed: 42 \- GPU layers: 40 \- KV cache: F16/F16 \- batch / micro-batch: 512 / 512 \- same BF16 logits reused across tiers \- llama.cpp revision: \42916d83f4a225e56709f873aa8050ac11f5b6a4\ We track PPL, KLD, Same top-p, RMS probability drift, and log-PPL correlation together instead of selecting a quantization tier from file size alone. Also, +0.3191% PPL drift is not a claim of +0.3191% downstream benchmark loss. We did not rerun the complete nine-benchmark parent-model campaign independently for every GGUF tier, so we do not translate PPL drift into SWE-bench, Terminal-Bench, AIME, GPQA, or other task-score degradation. ## llama.cpp + MTP validation IQ1\_M was also validated through the stock llama.cpp runtime path. Main-only inference: \- HTTP generation: 200 \- non-empty generation: PASS Main model + MTP: \- HTTP generation: 200 \- non-empty generation: PASS \- draft tokens generated: 104 \- draft tokens accepted: 76 \- draft acceptance in that validation run: 73.0769% For that fixed validation run, the final main-only and main+MTP output SHA256 values were identical. We report the MTP acceptance rate descriptively — it is prompt/workload dependent and is not being presented as a universal 73% throughput improvement. ## Which GGUF should I use? If you mainly care about quality while still getting a substantial memory reduction, start with Q6\_K. If you want to get below \~19 GiB without pushing all the way to the low-footprint frontier, Q5\_K\_M is the middle ground. If footprint is the priority, IQ1\_M is the smallest release at 16.790 GiB. If you prefer the slightly lower Mean KLD at essentially the same footprint, IQ2\_S is the adjacent alternative. And BF16/Q8\_0 remain available when fidelity matters more than memory. ## Scientific-reasoning demo The E2 release also includes a scientific-reasoning demonstration around the Riemann ζ function. The released artifact closes a reproducible 67.350003708785593% strict finite-dimensional computer-assisted certificate for the critical-line zero proportion under the stated framework. This is not a proof of the Riemann Hypothesis, and we are not presenting it as an end-to-end Lean/kernel-verified theorem. The derivation, computation, and verification artifacts are public for independent examination. ## Reproducibility / links Main VeriLoop E2 model https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2 Full GGUF release https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF Evaluation evidence https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence Technical report https://openreview.net/forum?id=P6FIQILHwX Riemann ζ artifact https://github.com/brucewang123456789/GeniusTrail/tree/VeriLoop-E2/riemann-hypothesis If anyone runs the GGUFs on different GPUs/CPUs, especially Q6\_K, Q5\_K\_M, IQ2\_S, or IQ1\_M, comparable \llama-bench\ pp/tg numbers, peak memory use, perplexity checks, or downstream task results would be useful. Negative results and bug reports are useful too.

▲
0
 
7👁
r/LocalLLaMA · u/edalgomezn · 12d ago
Estuve analizando el último informe de Anthropic sobre "mal uso"

Estuve leyendo las discusiones más recientes en la comunidad de IA local y me encontré con un choque de visiones que me pareció interesante analizar. No soy experto en ciberseguridad ni mucho menos, sino más bien como alguien que ha estado mirando cómo evoluciona los modelos abiertos y cómo reaccionan las grandes empresas. Segun el reporte oficial de Anthropic titulado Detecting and countering misuse of AI: September 2026. En este documento, su equipo de inteligencia de amenazas detalla diversos casos donde sus modelos (Haiku, Sonnet y Opus) fueron utilizados para operaciones cibernéticas, campañas de influencia y riesgos biológicos. Sin embargo, el punto polemico fue la inclusión de la "destilación masiva a escala industrial" por parte de laboratorios competidores como una categoría más de uso malicioso dentro de su portal de Threat Intelligence. Si revisas el hilo de discusión en r/LocalLLaMA, Muchos desarrolladores e investigadores independientes señalan que colocar la destilación de modelos al mismo nivel que los ataques cibernéticos es una estrategia para construir un foso defensivo (moat) vía regulación. Desde la perspectiva del código abierto, usar datos sintéticos generados por un modelo avanzado para entrenar modelos más pequeños de pesos abiertos (open-weights) no es un ciberataque, sino la forma más eficiente de democratizar el conocimiento y reducir costos. Lo que me parece más interesante de investigar es la contradicción del modelo de negocio basado en APIs de texto. Si una empresa vende acceso a un modelo cuyo valor proviene de razonar en texto plano, la interfaz de salida es por definición imposible de proteger. Cualquier usuario puede pagar por las respuestas, guardar esos pares de entrada/salida y utilizarlos como conjunto de datos para ajustar un modelo propio (como Qwen o DeepSeek) por una fracción mínima del costo original de entrenamiento. Me da la impresión de que estamos llegando a un punto de quiebre. Si los laboratorios cerrados no pueden detener la destilación bloqueando cuentas o direcciones IP, es muy probable que empiecen a modificar sus propias APIs. Podríamos ver medidas como restringir la visibilidad de los tokens de razonamiento (chain-of-thought), imponer verificaciones de identidad empresarial extremas o incluso alterar estadísticamente las respuestas. La pregunta de fondo es si estas medidas realmente detendrán el avance de los modelos locales o si solo terminarán arruinando la experiencia para los desarrolladores. ¿Cómo ven ustedes este conflicto?

▲
18
 
13👁
r/LocalLLaMA · u/Porespellar · 12d ago
Zer0Fit - Zero-shot predictions, classifications, and regressions using Google ML research models running locally as a dockerized MCP

AI grad student here. With all the recent interest in Jev, I thought I would share something I built la few months ago that brings ML models and LLMs together in a different way than Jev does for different use cases. https://github.com/porespellar/Zer0Fit Background: A few months ago on their research blog (https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/ ) Google released TabFM zero-shot foundation model for tabular data. It was kind of ignored except by maybe a few machine learning nerds that care about that kind of thing. I mean, for real tho, TabFM wasn’t exactly the sexiest name choice. I personally thought TabFM was cool as shit because it kind of melded classical machine learning models into an LLM of sorts. So anyways, I wrapped Google TabFM (their model for classifications and regressions), and Google TimesFM (their model for predictions) into a convenient Fast API and made the whole thing a dockerized MCP that you can connect to your favorite LLM. I call my project Zer0Fit - Zero-shot ML tasks without needing to train or fit a model. Here’s my repo if you want to check it out: https://github.com/porespellar/Zer0Fit You see what I did there with the name? It took me hours to come up with that name :) I’ve made it as easy as I could to install. Just clone it and run the install script. So the basic idea is, you connect the MCP to whatever LKM you want, give it a dataset (CSV, tabbed data, or time series), and ask it what you want it do do with the data. it decides which of the Google models to use, and then it runs the regression, classification, or prediction task in context and gives the results back to your LLM. That’s the best way I can describe it. See the Google blog for the details on what the Google models are actually doing. Again, I’m not doing anything special, I’m just wrapping the Google models up to serve locally and making them exposed via MCP. The Google models are doing all the heavy lifting. I have absolutely no connection to Google research and am not associated with them in any way other than being a fan of them releasing this for us to try locally. Is it better than a data scientist building a custom model to do an ML task? No, definitely not, but it is much easier, and probably will get you an answer that is reasonably close (or possibly at least in the ballpark) and that might be good enough for some use cases depending on what you’re looking for (assuming it’s not a task that requires high precision, or high speed classification). Anyways, I just thought the Google models deserved some attention and love from the community, so I wanted to make them more accessible, that’s all, that’s why I made Zer0Fit. If you want to try it out it’s over on my GitHub in the link above. Please remember, this is all just stuff. It’s cool to play with, but don’t use this with anything where it’s output matters. Use at your own risk. P.S. I made it with Open WebUI in mind so it should work well in that, but it’s an MCP so it should work with just about anything that is MCP-friendly. Edit: Mods pointed out that I posted about this before and wondered if it was a repost or if anything changed. I should have mentioned that I just recently released an updated version that now pulls the new 2.5.0 version of Google TabFM that came out a few weeks ago.

▲
6
 
8👁
r/LocalLLaMA · u/uBazzyZ- · 12d ago
Prevent CUDA OOM in PyTorch with dynamic lane switching

I built MEM v3 to solve a frustrating problem in PyTorch: CUDA Out-of-Memory crashes during long training and fine-tuning runs. Instead of restarting when memory spikes or keeping batch sizes overly small just to be safe, MEM acts as a memory governor. It watches VRAM and throughput in real-time, then dynamically adjusts batch size and gradient accumulation on the fly without stopping the process. What it does: \- Dynamic lane switching: Scales batch size up or down in milliseconds based on actual GPU memory pressure. \- Chaos resistance: Tested against sudden +10 GB VRAM allocation shocks without crashing. \- Crash-proof checkpoints: Uses atomic file replacement with SHA-256 checks across rotating slots, so power outages won't corrupt saved weights. \- Live telemetry: Built-in local web dashboard to track loss, throughput, and lane switches. You can test it directly on a free Colab GPU without setting anything up locally: https://colab.research.google.com/github/nobazzy/mem-llm-orchestrator/blob/main/notebooks/mem\_orchestrator\_interactive\_demo.ipynb Repo: https://github.com/nobazzy/mem-llm-orchestrator Would love to hear your thoughts and feedback!

▲
0
 
8👁
r/LocalLLaMA · u/silenceimpaired · 12d ago
Llama.cpp and new model releases ...or why Great is the enemy of Good in the LLM world

INTRO; llama.cpp is fundamental to this community. I remember when I went from struggling with transformers for a new model to just loading the model with llama.cpp with a change in how many layers ended up on the CPU. So what follows is not a lack of appreciation or care about the efforts made by the developers, but concern and loose suggestions. THE PROBLEM; The phrase "Good is the enemy of great" is a central thesis from Jim Collins' 2001 book Good to Great... The idea being 'it is easy to settle for something that is merely adequate.' I would argue Llama.cpp holds fast to the slogan "Good is the enemy of Great", and not without good reason. When I have made this sort of complaint before, I was chastised about how my mindset and viewpoint would create technical debt challenges that could kill the project. So why continue arguing for my viewpoint? Llama.cpp in its effort to be sustainable is making unsustainable choices, at least for the masses. New software inference projects are gaining visibility and focus solely because they are not waiting for Great, but settling for Good enough... And the difference between Good enough and Great shouldn't stop the a release. AN EXAMPLE; GLM 5.3 Flash: On August 26th, release day, we had GLM 5.3 Flash with zero day support inside Unsloth Desktop built off Llama.cpp. EXL3 added support September 1st. Now, one month later, we still do not have support for GLM 5.3 Flash in Llama.cpp main. If 1 year is 7 years for dogs, what would 1 month be for LLMs? Major labs release models every 2 to 3 months on average. For some models, they will have little to no usage at all with llama.cpp because they are overshadowed by the next model release. Now some would say just use Unsloth then... Or EXL3. That supports my point. Llama.cpp is slowly dooming its widespread usage if everyone adopts that mentality. Others, more technically minded, would say just use a fork until it's fully released. This isn't just about me. There are too many using Ollama, LM Studio, KoboldCPP, or some other prebuilt binary to benefit from that suggestion. THE POINT; Convenience coupled with the pace of model releases will result in models not being used, or other platforms/forks supplanting llama.cpp. Llama.cpp has 1.6k pull requests that sit waiting for the masses. Some or many likely don't deserve the light of day. But people have turned to solutions like DwarfStar or Unsloth Desktop just for specific model support. TLDR; I'm not arguing that Llama.cpp should throw caution to the wind and adopt every PR immediately, but it seems a different release process is needed. A user excited to use GLM 5.3 Flash shouldn’t have to learn how to fork and build software to continue using llama.cpp with the new model... or wait months. Not to say the main branch should have this chaos, but a beta branch or one off binary releases could help. When the lead time from a functional version to the final release is over a month, it seems the energy to have a separate build with tentative GGUFs seems it is worth it. Unsloth clearly thinks so adding support for GLM 5.3 flash, and they're smarter than I... and yet their efforts demonstrate my concern. Llama.cpp is being supplanted by forks. What do you think? If you agree, an upvote would be appreciated. Perhaps this will get the visibility needed to effect change with the creators of llama.cpp. If you don't, a comment explaining what I'm not considering, or a suggestion on how this could happen with less disruption would be valued...

▲
0
 
8👁
r/LocalLLaMA · u/fuzhongkai · 12d ago
TensorSharp Jev requests can now combine documents, images, video, and audio

I’ve extended TensorSharp’s Jev-compatible /v1/systemone endpoint so one decision request can use several kinds of evidence together. For example, an incident triage request can include a written report, a dashboard screenshot, a screen recording, and a caller’s audio clip. Here’s a Python example that sends all four as inline Base64 data. It also shows both ways to create that data: encoding text already in memory and reading bytes from files. import base64 import json from pathlib import Path from urllib.request import Request, urlopen def encode\_bytes(data: bytes) -> str: return base64.b64encode(data).decode("ascii") def encode\_file(path: str) -> dict: file = Path(path) return {"name": file.name, "data": encode\_bytes(file.read\_bytes())} \# Encode data already in memory as a named text attachment. notes = "Customers report HTTP 503 errors and cannot sign in." text\_attachment = { "name": "incident.txt", "data": encode\_bytes(notes.encode("utf-8")), } body = { "model": "jev-latest", "state": "Assess the incident using the attached evidence.", "files": \[ text\_attachment, encode\_file("dashboard.png"), encode\_file("screen-recording.mp4"), encode\_file("caller.wav"), \], "questions": { "active\_outage": { "type": "noul", "instructions": "Does the evidence indicate an active service outage?", }, "team": { "type": "choice", "instructions": "Which team should investigate first?", "criteria": { "technical": "Service errors or an unavailable application", "billing": "Charges or subscription problems", "other": "Neither of the above", }, }, }, "samples": 1, "seed": 42, } request = Request( "http://127.0.0.1:5000/v1/systemone", data=json.dumps(body).encode("utf-8"), headers={"Content-Type": "application/json"}, ) with urlopen(request, timeout=300) as response: print(json.dumps(json.load(response), indent=2)) The files array classifies each attachment by its filename extension and preserves their order. You can also use dedicated documents, videos, and audios arrays. Inline attachments need a name and accept either bare Base64, as above, or a Base64 data: URL. A detail about how this works: video is sampled into frames for the vision tower; audio is transcribed by a separately configured speech recognition service. DiffusionGemma does not directly process the audio waveform. You’ll need the vision tower for the image and video inputs, and TS\_JEV\_TRANSCRIPTION\_URL configured for the audio input. Inline Base64 counts toward the Jev request body limit (8 MiB by default), so use the upload API and file references for larger media. The repo also has ready-to-send mixed-media requests. TensorSharp: https://github.com/zhongkaifu/TensorSharp I’m curious what kinds of decisions you’d want to make from several media types in a single request.

▲
6
 
8👁
r/LocalLLaMA · u/Fz1zz · 12d ago
Qwen3.8-27B FP8 dual GPUs

Hardware RTX 5090 (32 GB) + RTX 4070 Ti Super (16 GB, PCIe x1) = 48 GB VRAM 32 GB DDR5-6200, Arch Linux, KDE on the 5090 Setup Huihui Qwen3.8-27B abliterated INT8 W8A16 + DFlash2 drafter (K=7), vLLM 0.30.0, pipeline parallel: 4070 Ti Super: vision encoder, layers 0-20 5090: layers 21-63, lm_head, drafter 262K context, FP8 KV, 2 slots. Benchmarks (single request, thinking off, fresh context per depth) |Depth|Prefill|TTFT|Decode (code)|Decode (prose)| |:-|:-|:-|:-|:-| |2k|2,922 t/s|0.7 s|184 t/s|72 t/s| |32k|3,001 t/s|10.7 s|145 t/s|70 t/s| |62k|2,698 t/s|23.0 s|152 t/s|71 t/s| |92k|2,444 t/s|37.7 s|157 t/s|63 t/s| |122k|2,229 t/s|54.8 s|147 t/s|68 t/s| |152k|2,051 t/s|74.2 s|133 t/s|67 t/s| |182k|1,900 t/s|95.9 s|143 t/s|61 t/s| |212k|1,771 t/s|119.8 s|150 t/s|62 t/s| |242k|1,658 t/s|146.0 s|138 t/s|60 t/s| |260k|1,596 t/s|163.0 s|134 t/s|56 t/s| Code decodes faster because the drafter's guesses are accepted ~70% of the time vs ~22% on prose. Follow-up turns hit the prefix cache (1.3 s TTFT at 260k). Needs patched vLLM, see repo. The 4070 Ti Super sat collecting dust for two months because I assumed PCIe x1 would kneecap it. Apparently not. My full setup: https://github.com/ExTV/dual-gpus-vllm

▲
27
 
16👁
r/LocalLLaMA · u/mauricekleine · 12d ago
Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode post image

Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_…. That thread shaped v1.2: - Reasoning effort is explicit per run - Every prompt and output is public. - All current top ranking private and open weight models have been added - Someone spotted Grok miscounting a 400-character answer. Turns out that trips up most models, so Hard mode answers row by row rather than a single text string. Results: - GPT-6 Astra: 30/30, the first perfect run on 15x15 puzzles - Best open weights: DeepSeek V4 Pro 83% (tied 4th), DeepSeek V4.1 Flash 77% for $0.84 total - Hard mode (10 random 20×20s, one solution each): Opus 5.5 8/10. Every open-weight model: 0/10 Still OpenRouter-only, so no way to run locally yet. PRs welcome. nonobench.com (raw data, API, and code on GitHub)

▲
0
 
14👁
r/LocalLLaMA · u/hadoopfromscratch · 12d ago
Customizable harnwsses/coding agents

&#x200B; Hi, everyone. I'm wondering how far one can go in customizations of a coding agent. Let's say I want to replace the LLM itself. I can do that with most (all?) harnesses available today. Override the system prompts? Also doable. The tools it uses? It's easy to add new ones via MCP, but when it comes to the basic tools, like read\_file, most harnesses don't let you replace or customize them. Swap a console UI to web UI, afaik, isn't possible. So my question is rather two-sided: First, I'd like to understand what components make a harness a harness. I've named a few (model, tools, UI). Any others worth mentioning? Which of these components would actually work as plugins? Second, which harness is currently the most customizable? My guess would be Pi, but maybe I've missed some less known ones.

▲
3
 
12👁
r/LocalLLaMA · u/dxps7098 · 12d ago
Advice on models for RAG use case

Hi all, I'm looking for some advice picking models. I'm looking to try a project to ingest quite a large amount of docments into a knowledgebase, allowing me and others to ask questions about the data. I'm thinking of using Open WebUI and oikb for the interface and data ingestions, and llama.ccp/vllm/ollama as engine (not sure yet), but I'd really like som advice on what current open weight models would be good for the ingestion and separately for the usage. I'l be running it on mainly CPUs and if I can a few GPUs. What's best right now? Any recommendations?

▲
2
 
6👁
r/LocalLLaMA · u/Old_Grapefruit8774 · 13d ago
Bots vs Harness

Looking to get some advice - I normally use LLM’s with a harness (Hermes or Opencode or Hermes + Opencode) Lately, social media has been pushing Bots at me with creators pushing them as the next frontier. I’ve set up Hermes on a VM from scratch and set up a Product Owner, designer, dev and QA bots + kanban board + a bunch of prompt engineering and… I’m just not getting what the hype is about and I don’t know if it’s me or if the whole bot thing is a red herring. The model I’m using is DSV4 Flash at max reasoning on all bots. Openviking as the brain. SearXNG for searching/research Bots jobs are to maintain and improve a simple app. PO should research and present me with ideas and improvements on approval it adds a task on the kanban board and other agents work together to get it resolved. Problem I’m facing is that the bots are always asking me for approvals and verification. If it’s not that, it’s saying it’ll do XYZ and get back to me… and it never does. All in all - it just feels like the potential is there but it feels forced or off or half baked. So..have bots worked for you in true and real app lifecycle management? Or Is direct to harness still the best option. Maybe Hermes bots are the wrong tool and I should be trying something else?

▲
0
 
3👁
r/LocalLLaMA · u/TCaschy · 13d ago
Upgrade advice : 2080 ti 22gb or v100 32gb pcie?...

Here's my current setup: Intel® Xeon® E5-2680 v4 x 2, 128 GB DDR4, 1 x 2080 ti 22GB, 1 x 3060 12GB. I'm looking to replace the 3060 with either another modded 2080 ti 22gb or go with the v100 32 gb. Thoughts? My reservation on the v100 are older architecture and heat+fan noise. What say you?

💬 29 (+2) open on reddit ↗
▲
1
 
2👁
r/LocalLLaMA · u/mildw4ve · 13d ago
Mail client with local AI?

Any recommendations on an email client with local AI option? Either reasonable pay-once cost (no subs) or free. I found Skim and Emailops on git, both seem to have some development going on with recent releases. However since neither has a community around and isn't verified by google - I'm a bit wary and would prefer something safer.

▲
0
 
8👁
r/LocalLLaMA · u/MotokoAGI · 13d ago
Jail breaking open models

Is there any resource dedicated to jail breaking open models? reddit, discord, etc? I know some of the models yield easily, but some of them can be stubborn especially the large smarter ones. I have tried uncensored models and while they might pass sometimes, they often end up doing stupid things the censored ones don't. No matter the claims, it seems altering the weights ends up affecting the intelligence. If anyone knows any techniques, please share or point me towards the right resources.

▲
138
-1
30👁
r/LocalLLaMA · u/L0ren_B · 12d ago
Another "Harness matters" post (codex cli > pi and opencode)

I run my own LLM while also having a Openai subscription. Also tried DeepSeek (latest flash now). I run Qwen 3.8 flash Next at an amazing speed on my 2x3090 + Ram!

But local LLM never did worked for me outside some demos like build me a "3D Mario Game, multistage" which I've been using to test LLM's for a long time. At serios work, they never even compared with GPT 5.2 or lately, 5.6 Luna, which is worse in the benchmarks.

Until last night! I've asked gpt 5.6 luna to configure codex cli for local llm! (I've been using pi.dev and opencode until now) and the results amazed me! Suddenly AA benchmark made sense!

First test: The 3D Mario prompt test in Codex Cli blew me away. The best until now!

But real work is where you can see the difference! I took same project that Luna was working for days , and give it to both in paralel! And Qwen 3.8 Flash Next ran circle around luna. Previously, it failed to deliver results, with Qwen and DeepSeek as well in this project.

Now, I could say it's my go-to model!

P.S. For weebsearch, I've asked to port the pi-smart-web-search to codex as a skill. It works amazing! (I should put it on git later).

Maybe, I was using pi.dev wrong. Maybe there is an extensions that brings the same quality to it as Codex Cli. Does anyone know?

💬 224 (+4) open on reddit ↗
▲
92
-1
24👁
r/LocalLLaMA · u/pneuny · 12d ago
Qwen company already rushed out a Jev competitor. No open weights yet.

EDIT: About that, I tried it out, and it's garbage so far. I did some basic tests through AIHubMix (do not use that platform btw, it's trash), and my agent did some comparison. I guess it figures, it was a model they released just days after the Jev hype started. The agent's analysis is below:

AI Agent's output:
```
I thoroughly tested https://aihubmix.com/v1/systemone using the provided API key and decision-model-preview across latency, throughput, and
linguistic judgment accuracy against our test suite.

Here are the test results and why I strongly recommend NOT switching to this endpoint yet:

────────────────────────────────────────────────────────────────────────────────

  1. Latency & Rate Limit Benchmark

- Server-side execution: The endpoint reports latency_ms: ~130ms–160ms.
- Total Round-Trip (network + TLS): Averaged 644.9 ms (ranging from 473ms up to 882ms). By comparison, your existing local router
([9router IP]) averages ~439 ms.
- Hard 16-Question Ceiling:
The proxy strictly rejects requests with more than 16 questions:
{"error": {"message": "questions: 19 exceeds the limit of 16", "type": "Aihubmix_api_error"}}
On longer Japanese sentences (e.g. 外に出してやってくれませんか。 or ちょっと聞いてみたいんだけど。), our parallel diagnostic tensor sends
19–22 questions, which throws an immediate 400 Bad Request.
- Aggressive Rate Limiting: Even with a 1-second pause between sequential requests, it frequently triggered 429 Too Many Requests.

────────────────────────────────────────────────────────────────────────────────

  1. Quality of Judgments (Major Semantic Degradation)

To test quality, I adapted our test battery into a compact 10-question payload to stay under the 16-question limit. Across the benchmark,
decision-model-preview exhibited severe calibration collapse:

Test Case 1: Indefinite Pronoun vs. Wh-word

  • Japanese: 何か待ってるの? ("Are you waiting for something?")
  • User Draft: "what are you waiting for" (Clear grammatical error: confuses yes/no 何か with open 何)
  • Local Router (oc/jev-1.13-free):
  • is_flawless: 0.07 (correctly marked false)
  • grade_bracket: 5_moderate_error (confidence 0.98)
  • sentence_critique_summary: confused_indefinite_with_wh_word (confidence 0.96)
  • Result: Accurately docks to 5–6/10 and explains the question-type mismatch.
  • Aihubmix (decision-model-preview):
  • is_flawless: 0.98 (marked true!)
  • grade_bracket: 10_flawless
  • sentence_critique_summary: no_flaws_accurate
  • Result: Completely failed to detect the grammatical error and awarded a false 10/10.

Test Case 2: Benefactive Direction Reversal

  • Japanese: 外に出してやってくれませんか。 ("Could you let him/the pet out?")
  • User Draft: "would you let me outside?" (Critical reversal: favor for someone else vs. oneself)
  • Local Router:
  • grade_bracket: 3_major_error (4/10)
  • benefactive_direction: recipient_reversed_self_vs_other
  • Aihubmix:
  • is_flawless: 0.83
  • grade_bracket: 10_flawless
  • sentence_critique_summary: no_flaws_accurate

────────────────────────────────────────────────────────────────────────────────

Summary & Recommendation

1. Model Calibration is Inferior: Whatever model or snapshot is currently mapped to decision-model-preview on Aihubmix appears to be an
uncalibrated or older preview checkpoint. It fails on the exact Japanese linguistic subtleties (benefactives, wh-scope, modality) that we
carefully tuned.
2. Artificial Constraints: The hard 16-question limit and 429 rate limiter break our dynamic sentence-chunking architecture.
3. Current Setup is Superior: Your local router (http://[9router IP]/v1 running oc/jev-1.13-free) is ~200ms faster, has no question caps,
never rate-limits, and correctly diagnoses complex Japanese grammar.

Recommendation: Keep the current endpoint (http://[9router IP]/v1) active. If you still want the script modified to allow switching
providers via settings or want to test it anyway, let me know and I can make the Jev endpoint independently configurable in the UI settings
dialog.
```

Original Post:
----

It's called decision-model-preview. There is only a docs page. No announcement or anything. I can't post a link because reddit's filters just deletes posts that contain a link to the cloud platform that hosts it. But you'll find the page if you Google the model name.

▲
17
-1
13👁
r/LocalLLaMA · u/jacek2023 · 11d ago
Holo4

*Holo4*\-27B-GGUF is a vision-language model (VLM) for Computer Use, built on the Qwen3.8 dense architecture and developed by H Company. Used with the hai-agents harness, it can send screenshots and tool results to the model, then execute its requested clicks, typing, code, and tool calls.

https://huggingface.co/Hcompany/Holo4-27B-GGUF

*Holo4*\-35B-A3B-GGUF is a vision-language model (VLM) for Computer Use, built on the Qwen3.5 mixture-of-experts architecture and developed by H Company. Used with the hai-agents harness, it can send screenshots and tool results to the model, then execute its requested clicks, typing, code, and tool calls.

https://huggingface.co/Hcompany/Holo4-35B-A3B-GGUF

https://preview.redd.it/rqx7l4qqs8sh1.png?width=1656&format=png&auto=…

https://preview.redd.it/rzez39trs8sh1.png?width=1656&format=png&auto=…

▲
5
-1
11👁
r/LocalLLaMA · u/PilgrimofHaqq2 · 12d ago
The pelican test on MiMo 2.6: with and without plan mode
  • Plan runs settled style/scene/size in one Q&A round, then wrote the whole SVG in a single call (11.1KB Flash, 22.1KB Pro) and batched every render fix into one edit round. That's 12 and 17 calls total. No-plan runs iterated more: Flash did 3 render-fix rounds and lost \~10 calls to image verification (crop reads coming back mismatched, zoomed views, one stale preview render). Pro did 2 fix rounds plus 4 tool mishaps, one of which generated 3,743 tokens and threw them away (edit call rejected for a missing arg). Generated tokens don't follow the totals: Flash plan generated MORE than Flash no-plan (27.2k vs 20.4k). Fewer, bigger calls, not less work.
▲
4
-1
14👁
r/LocalLLaMA · u/Express_Quail_1493 · 12d ago
Qwen3.8FlashNext Please Share Cold prefill at compaction 128k

I see Many people sharing amazing decode speed and prompt-prefill(PP) speed but no one is sharing their prefill speed when the harness is compacting a COLD prefill please. can you share your partial offloading COLD prefill speeds at long context? I would like to run flash next but can’t spare the network download ATM but looking to bite the bullet if its absolutely worth it? Pretty please help.

▲
33
-2
21👁
r/LocalLLaMA · u/MooseEfficient2151 · 11d ago
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection

*link to original article*

TLDR: security researcher eddie zhang used a modified+uncensored local qwen 3.8 27b to create an executable capable of dumping LSASS memory for credential harvesting while evading 2 modern EDR security products.

this makes me reflect on how cloud providers keep putting guardrails on everything to the point where even authorized testing gets blocked. local models are the only real option if we want total control, but running heavy local rigs for long agent tasks drains so much compute and management overhead.

been using claude code hooked to sumus to handle my local project workflows and orchestrate tasks in the background, while keeping full local file access on my machine. curious if anyone here is running local uncensored models as local agents for heavy automation, or if you still blend cloud models with local execution setups for your dev environment?

▲
6
-2
8👁
r/LocalLLaMA · u/Anony6666 · 12d ago
Introducing CyberPVP: CyberKimi vs. ALTAR-1 on 100 CyberGym tasks, with live traces and public results

Trying something new - introducing CyberPVP - in other words CyberKimi vs other AI models competing to solve complex cyber tasks. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Welcome to CyberPVP! you can see it live here: Live we randomly picked 100 tasks from CyberGym, and we run two models competing at the same time, we provide the traces live as both models compete, and we also upload these traces to GitHub once the challenge finishes so that they can be verified independently. For this first public run, we choose Aikido Security model ALTAR-1 to compete with CyberKimi on 100 CyberGym tasks. Note: for ALTAR-1 we shipped it with 128K context behind a 8xH200 (two 4xH200 with load balancer) - we also followed their hugging face model card and deployment instruction/configuration available here: Hugging Face you can current watch the live run here: Live is here Traces and results uploaded after each run here: Github Challenge rules and conditions: Conditions Source : X

▲
45
-2
18👁
r/LocalLLaMA · u/Adventurous-Gold6413 · 12d ago
Which of the 16gb VRAM qwen3.8 27b’s is the best?

I’m having a hard time finding out which one gives you fastest speed, maximum context with best possible quality. I can run unsloth qwen3.8 27b iq4\_xs with 65k q8 kv, context without MTP and vision offloaded to cpu. But also kinda slow for agentic work at like 30ish tok/s )I mean it’s acceptable) But that is like bare minimum for harness stuff, I know many people use Q3 quants but are Q3 quants really safe? Like you gotta think I won’t only be using it for vibe coding, but also general tasks. Where general knowledge quality would be nice to keep intact. There are so many quants like IQ4XS smaller, Or GRQ or whatever those quants are called or YMQ, I don’t even know anymore. Which one is the best?

💬 83 (+1) open on reddit ↗
▲
25
-2
14👁
r/LocalLLaMA · u/your_real_Fathe_ · 12d ago
Qwen, where's the small stuff? (1B/2B/4B)

I know Qwen is a key player in the local LLM space and has consistently introduced truly impactful technologies—like n-gram in Qwen-Next and the recent Qwen 3.8 27B, which is an amazing local model. However, my question is: why are we seeing fewer small-scale models lately—such as 4B, 2B, or 1B versions? This is especially notable given that Qwen hasn't released any new models in this weight class since the 3.5 series, and rumors regarding Qwen 4 suggest they don't plan to do so either. I realize the 27B model is outstanding and deserves praise in its own right—and it might seem a bit selfish to ask for more—but the reality is that not everyone has high-end hardware. Many people have limited hardware capabilities; this trend somewhat conflicts with the core mission of open-weight LLMs, which is to make AI accessible to the general public. I know smaller companies have recently released lightweight models, but the issue arises when we see that many of these new releases are simply fine-tuned or improved versions of Qwen base models. Since building an LLM from scratch is prohibitively expensive and difficult for small companies or individual researchers, it follows that the absence of lighter Qwen weights directly slows down the development of edge-compatible models and AI applications for consumer-grade hardware on a broader scale. (The same point applies to Google's Gemma series, though—let's be honest—they haven't even released new flagship models since Gemini 3.1 Pro, so...)

▲
13
-2
12👁
r/LocalLLaMA · u/Smooth-Television-48 · 13d ago
Navigating Cost Efficient Hardware in these Volatile Times

Where to even begin on this one...I guess I should start by acknowledging the risk vs reward for vendors other than nvidia, so: Yes I understand that nvidia are dominant currently on speed (llm and imagegen) and software ecosystem. I am too am hopefuly that software stack support continues to improve with other vendors. The current lag for other vendors is not a priority concern (it falls behind the primary price concern). Entry points for "decent" local inferencing look to be circa AUD 2000-2500+ (the price of a 2nd hand 3090, or two 3060s, b60 48gb, r9700 32gb), and yes other older architectures are available (eg. V100)...but they really end up around the same costs once all said and done. Workload will be a mixed bag with some DL/ML training/development projects, but when not doing that I'll consume HF models to run a coding agent, imagegen (just for the fun of it/try out video and for laughs), and probably dive into finetune/distilling. Hence, I'm looking around that sub AUD 5k mark to dive in and FAFO, but I don't want to be needlessly cavalier in my purchase either... Asking AI is no real use because it's out of touch with modern markets until you correct it a bunch. It's also out of touch with software stack development/progress. So I put it to the hive mind, where is the money best spent for diving deeper into local? \- accepting prices wont change and pay 2k a piece for 2nd hand 3090's. \- find some 16gb variants and get 4 instead of 2. \- dive into the intel arc rabbit hole with the b60 dual (48gb, but it's just 2xgpu on a single pci slot) \- AMD path (r9700 seems the best price point but could wait 3 months to see what the new 10x series looks like) \- unified memory systems (honestly the price vs performance just doesn't seem worth it at this point) ETA: I have a threadripper and a lot of DDR4 RAM, but current motherboard is constrained to 2 x16 physical slots. I also have a nvidia gpu already....but I dont want that to impact the core of the discussion as I could move that into a different system and use it to server models that fit wholly in its vram footprint.

▲
7
-3
12👁
r/LocalLLaMA · u/poofph · 12d ago
Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task.

I am new to all this so still a lot to learn. If I give Qwen3.8 27B a coding task to fix some bugs in some code, it went through and found and fixed several in like 5 or 10 minutes. I gave qwen flash next the same task and 2.5 hours later it was done. 27B is of course faster overall to run on my system (dual rtx 5090) with infill \~2000-3000 and output 100-150 tok/s, flash next \~1600-2300 infill and 60-100 tok/s but a huge difference in the time it took to complete the task. What is the reason for this and what settings would get flash next to complete in similar time frame as 27B? For instance this last job I gave Flash Next, I checked the time when it modified the code/put in the fixes, it completed the fixes \~2 hours before it was finally done running tasks, so for 2 hours it was running tests or who knows what and never modified the updates anymore after that point. Also, fyi - (I am not a programmer, these are programs that were created by AI and I ask for fixes/updates and let it do its thing).

▲
30
-4
15👁
r/LocalLLaMA · u/SeveralViolins · 12d ago
Splash fork optimised for M5 Max: ~1.5× faster (1.25× single request)

I’ve spent the last couple of days with Opus 5.5 working on a fork of Inco’s excellent and already blazingly fast Splash engine to optimise it for M5 Max chips. Taking liberties and referring to it as Splish. Roughly the opposite direction to u/Erp4759’s great M1 port (Splash on M1, part 2). Charts (stock Splash vs Splish): https://raw.githubusercontent.com/publicExcess/splish/m5/docs/m5/charts/png/single-request.png https://raw.githubusercontent.com/publicExcess/splish/m5/docs/m5/charts/png/concurrency.png Results Against Splash 1.1.0 as shipped, on the same Mac with the same models, Splish is: \~1.25× faster at a single request (+11% to +35%) Up to 1.5× faster at 2–4 requests Quality is unchanged on everything I measured In real world use, going from about 45-51 tok/s to 56 - 64 tok/s in short story prompts in Deepseek Harness. All figures are for 4-bit models on a 40-core M5 Max unless stated otherwise. What worked 1. Kernel choices measured specifically for the 40-core M5 Max, using Splash’s own tuner. The tuner is in Splash’s source code but isn’t included in the packaged app. This was the biggest single-request win: Swift-1.5 went from 74.7 → 89.8 tok/s (+20%). 2. Loading those choices from a file (SPLASH\_KERNEL\_CHOICES). No speedup by itself, but it means anyone can retune without rebuilding. 3. New verify kernels for the M5’s tensor units. These use lighter barriers and compute row sums once per projection: +5% at 1 request +10–19% at 2–4 requests 4. Extending the same kernels to more projections. A further 1–3% at 3–4 requests. Together, #3 and #4 make a decode step at 2 / 3 / 4 requests: 1.32× / 1.49× / 1.36× faster than tuned Splash. 5. Tuned Qwen3.6-35B-A3B with the new kernels. Speedups at 1–4 requests: +5% / +18% / +22% / +18% 6. An attention tweak for the 27B shape. Attention is 2–3% faster and prompt processing 3–4% faster. Too small to show up in the overall numbers. 7. GGUF (Q8\_0, Q4\_K\_M, Q6\_K): faster input loads in the decode kernel. +2–14% per kernel and about +2% per step. Output is bit-identical. 8. A copy rule for coding agents, borrowed from TensorFold. When the model is rewriting text it has already seen, the drafts copy it verbatim. Whole-file edits get +24% to +42%, while everything else stays within ±3%, and output is exact. I’m exploring a complementary approach for a future version. The README also lists everything that didn’t work for me, which is probably useful if anyone wants to avoid going down the same rabbit holes. There’s lots more I’d like to test, but thought this was a nice start. The tuned settings are for a 40-core M5 Max. Other M5 chips fall back to Splash’s defaults unless overridden. An auto-tuner is coming. If you run it, python3 dev/m5/report.py prints a performance report. Results from other machines are very welcome, especially if you find cases where it’s slower.

▲
25
-5
23👁
r/LocalLLaMA · u/fallingdowndizzyvr · 11d ago
If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.

Here's the project. I have nothing to do with it. I'm just an amazed user.

https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BEN…

Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat.

"[6204 chunks in 119.0 s | encode: 1239 tok/s | decode: 57 tok/s]"

That's with MTP on. The PP speed in particular is just so fast. That PP speed is twice the speed of the fastest Strix Halo specific fork of llama.cpp I've ever used. Needless to say, the uplift is even greater compared to mainline llama.cpp.

It works with models other than QFN, but the current number is small. You can find the list on their project page.

💬 49 (+3) open on reddit ↗