59 posts · 1 sub · RSS
← prev Monday, September 28, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
809
+31
56👁
▲
794
+30
64👁
r/LocalLLaMA · u/charles25565 · 12d ago
GPT-3 is discontinued today post image

It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely in our memories. Arguably what's more infuriating is that they suggest using GPT-5.6 Terra as a replacement. Keep in mind that Babbage is a model that's literally 3/4 of the size than MiniCPM5 2B. Even Luna might be overkill as a replacement. But neither is a drop-in replacement. Davinci is the main GPT-3 most people use. This is why we have local models, because they simply cannot have a universal end of life date.

💬 161 (+3) open on reddit ↗
▲
629
+27
45👁
▲
222
+13
34👁
r/LocalLLaMA · u/Educational-Care7867 · 11d ago
ImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench post image

Some context first.

I am process improvement / business consultant and had worked with Fortune 500 companies on improving their processes around refunds, returns, customer support, etc.

This entire thing had a lot of complex decision making and generally every decision / condition node in a process map was generally replaced by a human because factoring in ambiguity in a code is very difficult.

Idea of ImaJev

Hence, when Jev came out, I was very intrigued with it and also could clearly see its use-case of improving decision making in complex decision work flows.

However, Jev didnt have support for Images and I thought that it can be replicated for both Text and Images in a single model and thats when I started with ImaJev.

Training Process

It went badly at first. My first big fine-tune on about 500k short decisions made the 9B model worse at reasoning: 64.9 down to 42.3 on JevBench hard. It had basically learned to pattern-match. I spent the next couple of weeks generating hard questions with open-weight models and only keeping the ones where two Ai models agreed on the answer. That brought it back.

Results

Then, on the JevBench - It came out #1 of 91 (v1.4.2.2, scored 27 Sep), 67.37 vs Jev 1.13.0 at 63.29.

The same week DecisionBench put it #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1.

I honestly didn't expect either.

To be fair about it: the #1 is on a score that weighs accuracy, calibration, speed and cost equally. On accuracy alone it's #3. Its main strength is that when it says 90% it's usually right, and it'll say "can't tell" instead of guessing.

What it actually is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions.

It gives back a probability for each option plus "unknown", in one forward pass.

Runs on a Mac with MLX or on one GPU.

The whole project costed me around $1200 in rented GPU and a lot of time :P

I would love to know your thoughts on it - it anyone would be interested to try that.

💬 69 (+2) open on reddit ↗
▲
676
+12
37👁
▲
279
+11
52👁
r/LocalLLaMA · u/LegacyRemaster · 11d ago
Qwen next 3.8 and 3.8 27b Vs Sonnet 5.5 low and Sonnet 5.5 medium. post image

Six months ago, a result like this was unthinkable. But now we can say it loud and clear: local models are at the cutting edge, and the gap of just a few months has been confirmed.

Personally, I use Qwen-Next 3.8 for complex tasks; today, GPT-Sol-6-High was messing up a project, but Qwen-Next got it back on track. I consider it a reliable benchmark. What’s your take?

💬 117 (+3) open on reddit ↗
▲
176
+11
24👁
r/LocalLLaMA · u/Ok-Breadfruit-3523 · 12d ago
I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s post image

Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2 more of these and will probably get them running to see if 3.8 flash next runs at usable speeds. This setup is wildly inefficient with power but cost me less than $800.

▲
259
+9
32👁
r/LocalLLaMA · u/jacek2023 · 11d ago
nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 · Hugging Face

Model Developer: NVIDIA Corporation

Model Development: Fine-tuned from NVIDIA-Nemotron-3-Ultra-550B-A55B

[](https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-…)What is Nemotron?

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

[](https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-…)Description

Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra, fine-tuned for one epoch on 477,642 synthetic reasoning traces distilled from GLM-5.2 across 22,000 curated problems spanning 16 regional and international competitive-programming contest families. Selected as the SFT teacher for its higher accuracy and roughly 30% shorter generations compared to a DeepSeek-V4-Flash-trained variant, GLM-5.2 distillation yields a model that, combined at inference time with GenCorrect — an iterative closed-loop test-time compute strategy that generates diverse candidate solutions, incorporates evaluator feedback, and refines subsequent generations under a fixed submission budget — was evaluated live and prospectively on the IOI 2026 problem set under official contest time, internet-access, and submission constraints, scoring 535.4 out of 600 and surpassing both the gold-medal threshold (361.12) and the top human contestant's score (498.27), making it the first AI system reported to outscore the highest-scoring human contestant on an IOI problem set.

This model is ready for commercial or non-commercial use.

▲
104
+8
30👁
r/LocalLLaMA · u/tossit97531 · 11d ago
Can we get some quality control on all these model perf posts?

Too many hyperactive amateurs are coming in here with "1b model at 832843tok/s!" and hardly any of them have all the info necessary for local runners to evaluate. We need context ladders with perplexity/KLD, hardware specs, model params and quant(s), runtime, tuned runtime parameters, basically everything we need to reproduce locally if we can match the entire setup. To say nothing of what the model is even good at in the first place if it's not a well-known model.

The goal is to get perf numbers that show they meet a certain quality bar. I don't care if I get 8324834 tok/s if it's all garbage.

Can we start filtering the hyperactive amateur perf posts please? It's getting really frustrating seeing all these posts of models and wading through info just to see that it doesn't test with anything but an empty context or doesn't say anything about quant or platform.

We need to define some rigor and apply it to this place, or it will remain like most ai-oriented subs and get continually choked with slop.

▲
63
+8
43👁
r/LocalLLaMA · u/Top-Evidence174 · 11d ago
Mica v0.1 4B got diamonds in survival Minecraft on its first run. 26 decisions from an empty inventory. post image

I've been working on Mica, a 4B decision model, and wanted to see how far it could get in actual Minecraft, not a sim. New world, empty inventory, and the goal was a diamond pickaxe.

Last time I posted, it got an iron pickaxe. Honestly that took around 20 tries and it was pretty flaky. I've reworked the harness a lot since then. This time it went all the way to a diamond pickaxe, and once the harness was finished it did it on the first run.

It took 26 decisions and about 8 minutes of game time. It got wood, made a crafting table, then wooden and stone pickaxes, then iron and coal, a furnace and an iron pickaxe. After that it tunneled down to diamonds at y=2, mined three, put a crafting table down right there and made the pickaxe. Decisions took about 108 ms on average.

My favorite bit is around step 14. The planner wanted it to make planks to burn in the furnace, but Mica went and mined coal instead (0.69 vs 0.31) and then smelted all three iron at once. Which was the better call, honestly.

How it works: every step Mica gets the game state as text (inventory, nearby blocks, health, what happened last step) plus a few candidate commands, and it picks one. A Mineflayer bot running Mindcraft skills does the actual moving and mining. The panel on the right of the video shows each decision and its probabilities live. The bot also knows where the nearest diamonds are, so it isn't searching for them.

To be clear, I'm not saying a 4B model plays Minecraft on its own. What I wanted to show is that a model this small can sit behind a bot, read what's going on, and make the next call well enough to get all the way to diamonds.

I'm planning to release the harness soon. Mica will read Minecraft chat, so you can type what you want and it'll work toward it. Simple stuff like getting items, crafting or following you should work fine, but it'll struggle with anything really complex, like building a house.

Also, v0.5 should be out in the next 1-2 weeks. A lot of the architecture changed, and I did extra training on the parts where v0.1 was weak, so I'm expecting a clear jump in performance. The aim is to be at or near the top among 4B JEV-like models.

There'll be two versions: Mica v0.5 4B, and Mica v0.5 4B Distill Laya, which is light enough to run on pretty much any PC.

For Minecraft, I'm hoping v0.5 will be good enough to take down the Ender Dragon, and I did extra training specifically with that in mind. No promises, but it'd be really cool if it pulls it off lol

Oh and fun fact, Mica is a fully vibe-coded project

Model: https://huggingface.co/sky7350/Mica-v0.1-4B
Model code: https://github.com/akivet/Mica-v0.1-4B
Minecraft harness: coming soon

▲
187
+7
17👁
▲
80
+7
26👁
▲
69
+7
39👁
r/LocalLLaMA · u/returnity · 11d ago
Searching for 3.8 35B: Qwen3.6-35B-A3B (Testing 5 Finetunes vs. Base)

TL;DR -- You should probably just use base Qwen3.6-35B, as only Occamy-1.0 is competitive with it. Tiel is a major let-down, worse than Ornith. KAT surprises (good), Nex surprises (bad). This post is long. Sorry, lots to cover.

I think we all want to see a next-generation small MoE from the Qwen team to replace 3.6-35B in our workflows. This model is a perfect fit for smaller gmaing laptops and mid-tier rigs. It sucks that Qwen seems to have abandoned this model, but at least there are fine-tunes that improve upon it... right?

Well... maybe not. I ran benchmarks on the 3.6-35B-A3B base model, as well as five finetuunes: Occamy-1.0, Ornith-1.5, KAT-Coder-V2.5-Dev, Tiel-Coder, and Nex-N2.5-mini, and the results are quite surprising.

I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark I can run in ~10 hourso on my machine, it's not actively post-trained on by any of these models, and it provides a lot of useful information along with the raw accuracy scores. This includes: first-try and retry pass rates, token counts, solve times, and how well-formed the output diffs are. Here's the table:

| model | First-try pass | Retry pass | tokens | sec/case | tok/solve | well-formed diff |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B BASE (STOCK template) | 37.4% | 71.0% | 8650 | 285 | 14.1K | 96.3% |
| Occamy-1.0-35B-A3B (STOCK template) | 29.0% | 70.1% | 6801 | 285 | 17.2K | 86.9% |
| Occamy-1.0-35B-A3B (froggeric medium) | 30.8% | 69.2% | 6009 | 233 | 16.8K | 94.4% |
| Occamy-1.0-35B-A3B (froggeric, xhigh) | 27.1% | 67.3% | 8631 | 310 | 20.0K | 91.6% |
| Ornith-1.5-35B-A3B | 23.4% | 63.6% | 4813 | 226 | 16.4K | 87.9% |
| KAT-Coder-V2.5-Dev | 20.6% | 58.9% | 2190 | 84 | 9.3K | 86.9% |
| Tiel-Coder-35B-A3B | 18.7% | 53.3% | 4851 | 171 | 18.2K | 89.7% |
| Nex-N2.5-mini | 10.3% | 30.8% | 5037 | 188 | 33.3K | 95.3% |

As you can see, the only finetune that even competes with the base model is Occamy-1.0. The rest are utterly dominated by the base model, a grim disappointment for finetune enthusiasts. I was particualarly surprised by the performance of Tiel, which seems to get a lot of love in this subreddit.

Speaking of Tiel, I want to clarify that Tiel is just Ornith-1.5 with a different chat template, Sharp, which is based on froggeric with an added "terse mode" instruction that's supposed to reduce excessive verbosity. I wanted to standardize for templates, so ALL models are using the base froggeric v22.5 template set to medium (which is equivalent to standard thinking on, no additional message sent). I used this because I wanted to test Tiel vs. Ornith-1.5, and Tiel is the chat template. Also, practially, I use froggeric in my real workflows. However, to ensure coverage, I also tested the STOCK template on the 2 highest-performing models, to make sure it wasn't affecting the scores. As you can see, the template doesn't make a significant difference in the scores, and the scores for base 35B with different templates are so close to identical that I excluded the froggeric one from the table.

I also tested froggeric/Sharp's reasoning-effort toggle, and found xhigh -> medium significantly reduces token counts and solve times (by ~1/3), without affecting accuracy significantly. That stands in stark contrast to Tiel's 'terse mode' toggle, the core feature of Tiel over Ornith, which dramatically reduces accuracy along with the reduction in token counts. My results strongly suggest that if you want a less verbose model, you're better off lowering the reasoning effort than using Tiel with terseness on.

Speaking of token use, that's probably the big differentiator here. A couple models stand out: Ornith and KAT-Coder-V2.5-Dev are the most efficient models, with KAT in particular having a brevity unmatched by anything else. KAT is fucking fast, and I think despite its lower accuracy than Occamy, it has a place in my lineup as a subagent because it just gets. shit. done. Occamy is also interesting, as it is the only model that perfoms on a similar level to the base, but it uses 20-30% fewer median tokens. However, Occamy also had a number of runaway generations where the token count blew up, so it's total tokens/solve is actually higher than base.

In an effort to further distinguish Occamy from base, since Aider struggled to do that, I ran tau2-bench, an agentic tool-calling benchmark consisting of multi-turn interactions with a simulated counterparty. I figured this was a good bench to use as Occamy is post-trained specifically for 'co-work' scenarios, but not trained on this particular set. I used Qwen3.8-27B with reasoning effort set to low as the simulated customer in these conversations. The base model was able to pull away from Occamy in the harder retail domain of this benchmark, but Occamy resolved the issues in the airline domain at an equal rate while requiring fewer turns. Here's the results.

| model (Q8_0) | domain | pass^1 | tokens | sec/task | turns/task |
|---|---|---|---|---|---|
| Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) | airline | 80.0% | 5390 | 290 | 11 |
| Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) | airline | 78.0% | 6770 | 390 | 13 |
| Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) | retail | 86.0% | 4098 | 333 | 14 |
| Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) | retail | 79.8% | 3284 | 284 | 13 |

Overall, I think the results are clear, if unexpected: Occamy-1.0 is the only fine-tune that even competes with the base model on Aider Polyglot, but even it is not a clear winner. Tiel is noticebly worse than plain Ornith without the terseness toggle, and the terse mode doesn't even save any tokens. xhigh in froggeric/Sharp degrades accuracy slightly and bloats token use, which makes sense given the models were not RL'd for the extra thinking effort prompt. KAT-Coder-V2.5-Dev is the most efficient model, with accuracy nearly as good as Ornith and better than Tiel. Finally, Nex-N2.5-mini is a disaster.

💬 65 (+2) open on reddit ↗
▲
72
+6
32👁
r/LocalLLaMA · u/MasterNomie · 11d ago
What model sits between Qwen 3.8 27b and Flash next for coding?

Having tested both Qwen 3.8 27b and Flash next on RTX 5090 with 96GB RAM, I want to find the middle ground between the two for coding capabilities but not sacrifice decode speed to standstill. I would like decode speed to between 75-100 ideally for fast iterations; otherwise I become impatient.

Currently I get 200+ TPS on Qwen 3.8 27B and approx 50 TPS on Flash next.

My hardware - RTX 5090 and 96 GB DDR5 which I plan to upgrade to 128 GB (in this economy, yes, but unwillingly).

What model sits between these two in terms of coding capabilities and hardware requirement? If none is present, I can perhaps think of using Flash next for plan creation and 27b for implementation.

Edit: Fast forward few days. I gave Strata a go with Swift 1.5 Flash Next IQ3\_XXS and able to achieve \~150 tok/sec decode and 5k tok/sec prefill. I am escatic! The quality of response from 27b is considerably better and the speed is great. Both targets achieved.

▲
65
+6
29👁
r/LocalLLaMA · u/starkruzr · 12d ago
3090 NVLink bridges: are there clones of these? why are they so insanely expensive?

I think I need a 3-slot for my two cards. but holy fuck these things are pricey.

💬 69 (-2) open on reddit ↗
▲
36
+5
16👁
r/LocalLLaMA · u/light_2earth · 12d ago
macOS 27 ships a free local LLM on Apple Silicon Macs. I made it easy to use from Node and Python

Apple Silicon Macs on macOS 26+ come with a small LLM built in. No download, no API key, and nothing leaves your Mac.

Why I built it

I was making a tool that writes API docs from code, and I didn't want users to install Ollama or paste an API key. Apple's model was already on their Mac, so I used it.

Getting it to work well was harder than expected. At temperature 0 it kept repeating itself, and a 2-second call took 20. Calls took 17 seconds instead of 1.5 until I kept one process running. So I turned all the fixes into a library: apple-llm.

What it's good for

\- Pulling structured data out of messy text, like emails into tickets or receipts into expenses. The JSON always matches your schema.

\- Tagging, summarising and rewriting

\- Private data you don't want to send anywhere

\- Tools you share with other Mac users, who don't need to download a model or get a key

What it's bad at

Coding and long reasoning. It's a small model.

There's also an optional cloud mode that uses Apple's bigger server model for harder questions. That one is not local: your prompt goes to Apple's servers, and it has a usage limit.

Node: npm install apple-llm

Python: pip install apple-llm

https://reddit.com/link/1ws5l5p/video/oarqcgm767sh1/player

GitHub: https://github.com/jagdishpal02000/apple-llm

💬 26 (+1) open on reddit ↗
▲
42
+4
20👁
r/LocalLLaMA · u/Usual_Maximum7673 · 11d ago
Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights)

TL;DR: The Jeff models are a set of Qwen3.5 and Gemma fine-tunes for zero-shot classification: small, efficient, open-weight models with respectable out-of-the-box performance that can be slotted right into code or fine-tuned/LoRA-trained as needed. Give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in \~28 ms on an M4 Max (see caveats below).

Maybe equally exciting for open model enthusiasts like myself, everything was done on local hardware: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing - all connected and monitored from my Android phone via Tailscale. Apache 2.0, Jev-compatible API. Weights: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B · Code: github.com/firelex/jeff · Videos: games table

When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware.

So here's what I did:

  • The 0.8B trains in about 2 hours and the 2B in about 3.5, on one workstation GPU (RTX PRO 6000, 96 GB).
  • \~31k synthetic training questions written and checked by Qwen3.8-Flash-Next on two DGX Sparks. No cloud GPUs, and no closed-model output in the training data.
  • The rest of the 271k training questions are public datasets converted into decisions, plus 10k code-built probability questions.

Benchmarks (4,599 questions: BBH, Financial PhraseBank, JudgeBench, RAGTruth, WinoGrande):

|Model|Untrained base|Jeff (trained)|Calibration error|
|:-|:-|:-|:-|
|Qwen3.5-0.8B|45.3%|79.1%|0.049|
|Qwen3.5-2B|46.5%|83.1%|0.028|
|Jev (published)||83.0%|≈0.06|
|AutoJev-27B (published)||84.9%|—|

The caveat: the published numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86–89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64–68% against Jev's 94%, and \~50% on JevBench's hard tier against \~73%. See the HuggingFace model card for details. But that's not surprising, and I don't think it matters. No 0.8B or 2B model reasons like an LLM, and I don't think anyone should expect it to. The Jeff models are extremely fast judgement-callers (much faster than Jev), and have reasonable out-of-the-box performance. In one of my apps, I used the 0.8B model for voice-based navigation, and with a quick fine-tune, I got to real-time performance (24ms) at almost 100% accuracy.

Now the fun part: games, as a zero-shot test. Games are not the ideal zero-shot test, but they're fun, and the TypeSafe guys (Jev) did it, too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one. The options say what each move leads to, never which one is right.

|20 episodes each|Doom (kills)|Frogger (crossings)|Pac-Man (pellets of 98)|
|:-|:-|:-|:-|
|Random moves|−0.05|0|11.2|
|Hand-coded rule bot|6.55|10.25|94.1|
|Qwen3.5-0.8B, untrained|5.0|1.0|25.8|
|Jeff 0.8B|6.55|10.3|57.0|
|Jeff 2B|−0.9|6.0|41.2|

Jev's published Doom score is also 6.55, but its prompt spells out the aiming rule (fire when the bearing is between −8 and +8 degrees) and it takes \~212 ms per call over its API. Jeff gets "the nearest monster is a little to your left" and decides in \~29 ms on my Mac.

Lessons learned:

  • System 1 models are here to stay. Having the ability to process unstructured data at software speed inside an app is extremely powerful. And being able to do this locally is fantastic.
  • A small model is a classifier, not a planner. Models in the 0.8B-2B range don't reason like Qwen3.8-27B or Jev. But they also don't need to. As long as you present the options in the right way, you can get up to 50 decisions per second (depending on your hardware).
  • Fine-tune it if needed. If the models' zero-shot performance isn't good enough for you, fine-tune them briefly or add a LoRA adapter.
  • Wording matters enormously. Play around with how you present the options. Giving Frogger's final step option the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Previously, the frog just stayed on the last log.
  • Bigger isn't better. As the game tests showed, the untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers "turn away from the nearest monster"), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.
  • Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, not reliably, and in a real-time loop the mistakes compound.

Happy to answer questions about the pipeline (synthetic data from a local teacher, leak filter, calibration) or the game harness. Everything, including the videos, is linked above.

▲
57
+4
21👁
r/LocalLLaMA · u/lordekeen · 11d ago
Qwen 3.8 is a workhorse

https://preview.redd.it/23cumykhabsh1.png?width=944&format=png&auto=w…

Reminder that you can put Qwen 3.8 27B as a subagent and its a workhorse. Pic: using DeepSeek v4.1 Flash as Orchestrator in Pi, Qwen 3.8 27B GSQ RCO in llama.cpp.

▲
66
+3
34👁
r/LocalLLaMA · u/ZenZombie117 · 12d ago
Liked Muse, so I cut the 30B model in half by width, distilled it back, and it does 57 of 60 tool tasks its parent does 60 of

I've liked how Muse-Glimmer worked, so I wanted to see if I could produce a smaller "kid" out of it. Ornith's sharp decisions on when to think and which tool to call were the other thing I liked, so Ornith-1.0-9B got to be the policy teacher while the parent wrote the words. No RL anywhere, distillation only.

I present to you Xyntetik-Kvist-14B.

What it is good for

  • Smaller than the Muse parent but still manages most tool tasks: 57 of 60 held-out closed-loop tasks (contacts, weather, flights, currency, dates, units, stock), scored by re-executing the calls against ground truth, where the parent does 60.
  • Fits a 24 GB card whole at Q8_0 (15.4 GB) or the Q5_0 mix (10.3 GB) and serves an OpenAI-, Anthropic- and Responses-compatible API through Xyntetik Runner, so it drops into an agent loop you already have.
  • Every failed attempt is published beside it: 12 gated runs, 2 full passes, one shipped. The training record has the preregistrations, the amendments and the defects, so you can see exactly where it breaks before you build on it.
  • Give it a calculator tool for arithmetic. Without one it gets "17% of 2,340" wrong, and the card says so.

Numbers, from the card

| claim | number |
|---|---|
| parent | Muse-Glimmer-30B, cut by width (hidden 6,656 to 5,760, FFN 19,968 to 10,240, heads 32 to 24), all 52 layers kept |
| size | 14.44 B parameters; BF16 28.9 GB, Q8_0 15.4 GB, Q5_0 mix 10.3 GB |
| distillation | 6,000 steps, 98.3 M tokens, 162 hours, then 1,440 steps on agentic trajectories |
| fidelity to parent | KLD 0.762, margin-qualified top-1 84.0% on 45,056 held-out positions (a student's row, not the quant bar) |
| tool tasks | 57 of 60 held-out, re-executed against ground truth; parent 60, untrained control 0 |
| format and calls | 199 of 200 first turns well formed; 99 of 99 tool calls valid |
| attempts | 12 gated attempts, 2 full passes, attempt 12 shipped |
| weak spot | calc tasks 12 of 15 over 160; 7 of 160 runs end in a reasoning loop |
| serving | Runner v0.5.7 or later |
| licence | Apache-2.0 |

Links

EDIT: reading the comments, i should have said this first. this is not a general drop-in for Muse or a gemma4 replacement, and it was never going to be on my compute (98M distillation tokens vs the trillion a real distill wants, i simply lack the compute). the purpose was more on getting the tool calling right. IQ4_NL mix (7.6 GB) is up now too, it scores the same 57/60 on the tool tasks but does not fit an 8 GB card whole (50/52 layers on a 3070, ~5 tok/s).

▲
59
+3
27👁
r/LocalLLaMA · u/streppelchen · 12d ago
Minisforum MS-S1 MAX-P495 @ €7.799,00

MINISFORUM MS-S1 MAX-P495 – Minisforum EU

Expected to ship mid october.

At that price point, it doesn't make a whole lot of sense in my opinion.

I get that ram prices are where they are, i get that it's a newer model of hardware, but twice the price for 50% more ram and \~5-10% more performance is just hard, especially when compared to the recently released m5 ultra studio.

▲
5
+2
18👁
r/LocalLLaMA · u/Perfect-Campaign9551 · 11d ago
Recommended way to run Qwen 3.8 27b on a 3090 in Windows?

I know this may have been asked a lot, but I'm not an LLM expert yet (have never set up vllm myself or llama.cpp myself yet) I've only used scripts other people have set up.

Is there a simple way / steps to follow to run Qwen 3.8 27b in Windows with my single 3090?

I can run it straight up with Ollama with a 64k context but it seems like it only works reliably in chat and not in OpenCode (in OpenCode it "works" but at one point it "hung up" on me it seemed like. Not sure if maybe it was busy thinking still)

I've found quite a few threads that are close to what i'm asking for but many I think use WSL or something, too. Which I'm also not super experienced with yet.

I found the HyperQwen repo but their documentation is like...120% all technical and not very user friendly at all. I CAN do technical stuff! But it's barely passable as "do this, and then this" type of docs at the moment.

Ninfer is only for 5090 cards from what I read.

EDIT: Thanks guys, I was able to use llama-cpp-windows-manager project to get Qwen 3.8 27B up and going (Q4\_K\_M) . I have a 98K context and get 70tok/s and it's working with Open Code just fine. Very usable.

💬 34 (+2) open on reddit ↗
▲
55
+2
24👁
▲
46
+2
14👁
r/LocalLLaMA · u/KingGongzilla · 11d ago
Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090

Hi everyone :)

The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090.

To get the best of both worlds, I adapted Swift 1 and Swift 1.5 for HyperQwen and benchmarked them against HyperQwen’s specialized W4A16 AutoRound fast quant for speed and quality.

Performance

|Model|Average time/task ↓|Average output tokens/task ↓|Decode tok/s ↑|
|:-|:-|:-|:-|
|Qwen - HyperQwen fast quant|108.1 s|8,985|112.1|
|Swift 1.0 + HyperQwen|66.2 s|5,245|105.9|
|Swift 1.5 + HyperQwen, INT8 heads|72.2 s|5,751|104.0|
|Swift 1.5 + HyperQwen INT4 heads|68.2 s|5,669|107.2|

All models ran on one RTX 3090 24GB with FP8 KV cache and 150k configured context. Average task time covers about 630 tasks from different benchmarks listed below.

Swift-1 seems to use the fewest tokens and has fastest task completion time. Swift-1.5-INT4 achieves roughly 37% lower average time per task compared to the HyperQwen Qwen fast model. Despite slightly lower TPS (due to lower draft acceptance) it finishes sooner because it generates fewer tokens.

Quality

There are some minor quality and performance tradeoffs between the models:

|Test|Qwen HyperQwen fast|Swift 1.0|Swift 1.5 INT8 heads|Swift 1.5 INT4 heads|
|:-|:-|:-|:-|:-|
|GSM8K, 200 questions|97.5%|98.0%|98.0%|97.5%|
|IFBench, 300 prompts, strict|74.0%|73.3%|73.7%|72.3%|
|LiveCodeBench, (100-problem subset)|90%|89%|89%|91%|
|Custom tool-call/JSON eval|29/30|28/30|30/30|30/30|
|English/Python perplexity ↓|6.551|6.605|6.643|6.679|

Applied changes to Swift models to adapt for HyperQwen:

Changes to Swift models:

  • Kept the upstream AWQ INT4 model weights and converted embeddings to INT8.
  • Swift 1.0 and Swift 1.5 INT8-head variants: quantized the output head and MTP (multi-token prediction) linear layers to INT8 and added HyperQwen’s reference draft vocabulary for speculative decoding.
  • Swift 1.5 INT4-heads: quantized the output head and MTP linear layers to GPTQ INT4 instead, and built a Swift-specific 65,536-token draft vocabulary.

Setup

If you want to try it yourself, point your coding agent at these setup instructions and ask it to set up Swift 1.5 + HyperQwen on your machine.

All three models can be found here:
https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks

Big shoutout to UkisAI for Swift and syv-ai for HyperQwen! It's genuinely insane to be able to run these models on an RTX3090 at those speeds!

▲
37
+2
19👁
r/LocalLLaMA · u/PilgrimofHaqq2 · 11d ago
9 prompt rules cut my coding agent's wasted thinking up to 70% (GLM 5.3 & GLM 5.3 Flash)

360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%.

The block (shipped to global instructions):

Thinking discipline 1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it. 2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling. 3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps. 4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue. 5. Do not revise just to agree. If the user pushes back without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. Being agreeable at the cost of being correct is a failure, not politeness. 6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason. 7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle. 8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do. 9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke.

Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam

💬 21 (+1) open on reddit ↗
▲
32
+2
13👁
r/LocalLLaMA · u/Balance- · 12d ago
It would be really cool to have an official 3D-print engineering benchmark/leaderboard like this post image

Someone prompted different LLMs to generate CAD code for a bridge under fixed constraints (2-foot span, under 500g filament, 18-hour print limit), printed them, and load-tested them to failure. The results were quite varied: some models couldn't even design parts that fit together, while the winner held over 100 lbs. Most benchmarks don't really capture physical intuition, spatial reasoning, and functional code generation at the same time. I would love a standardized benchmark and leaderboard for this.

▲
13
+1
9👁
r/LocalLLaMA · u/SnooPeripherals5313 · 11d ago
3D/2D Text Visualisation post image

Like everyone, I use 3js for data visualisation. But while semantic clusters are interesting, they don't confer much practical information alone.

So I did something very simple: a query spatially re-assembles the nodes, and you can switch them to text.

Honestly, it's hard to swing 3D viz for text as a genuinely useful feature and not a novelty, but I get the feeling there's still some potential in the idea. Would be good to discuss, I'm sure someone here has made a better implementation.

▲
8
+1
10👁
r/LocalLLaMA · u/otacon6531 · 11d ago
IQ Quants still slow on P40?

I have been using Qwen 3.6:35b IQ4 via llama.cpp on my p40 and am getting anywhere between 37 - 83 tok/s (mtp is on). Prefill usually starts at 600 and slowly degrades as it continues processing so 600 for short prompts and more like 300-400 by the end of a long prompt. It hurts, but it is what my budget allows.

AI told me IQ quants are noticeably slower on the P40 and it referenced (https://www.reddit.com/r/LocalLLaMA/comments/1dmhpud/are\_iq\_quants\_slow\_o…) from two years ago, but I didn't feel it being slower when I moved from Q4 to IQ4, so...

What am I missing? Are IQ Quants actually a significant amount slower on the P40 or is this outdated information?

▲
27
+1
17👁
r/LocalLLaMA · u/Brief-Tap-6616 · 11d ago
95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

Hello everyone! A little while back I posted about LlamAmpere, a fork of Llama.cpp with Ampere-specific improvements (though it is caught up to main and will support other hardware, too).

Thank you to everyone that tried it out and shared back their results across the 30xx cards. I'm happy to share I've pushed v0.4 out this morning. On the 4.6bpw model tested, speeds improved \~10% vs the last version while also improving the max context by 10%+ (technically, it can go above 262K, but I have not tested any custom kernels or graphs to support YaRN).

The closest competition comes from vLLM, keeping within <10%, but does so with lower maximum context. It is significantly faster than other llama.cpp options tested.

https://preview.redd.it/u70q7lew1bsh1.png?width=1080&format=png&auto=…

[](https://preview.redd.it/95-tps-through-100k-generated-262k-ctx-on-a-single-30…)

There's also a number of other improvements for other quants/formats, with EXL3 seeing significant speed up (\~80% the speed of the 4-XS-M quant tested). It has a slightly lower KLD, but not a range I have found stat significance for at the task level, so I sticking with the XS-M model for now (built on top of Swift-qwen's distill, which is far more token efficient than the stock train for \~1% performance loss). the 4.3 bpw EXL3 model does provide a bit more room if you are interested in 2+ concurrent predictions. Improvements in this format and the IQ2/3 codebook quants will be most useful for people on 12/16/20 GB setups. These measurements are at temp=1, vs some of the vanity speeds you will see people claim with temp=0 and/or short generations.

As always, please share your results + config details so I can keep improving!

fork is here: https://github.com/JakeATX/llamAmpere/blob/main/QWEN\_AMPERE.md#build-and-run**

model used here:

https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF**

Build command:

git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server

Build + launch (linux):

git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server
curl -L -o ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf \\
https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF/resolve/main/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf
\-m ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf -c 262144 \\
\-ngl 99 -fa on -ctk turbo5 -ctv turbo4 -b 4096 -ub 1024 -t 8 -tb 8 --parallel 1
cd
./build-sm86/bin/llama-server

Previously, people had expressed concern over quantizing KV cache, and TQ specifically. The TL;DR on that is that any reasonable KV quantization strategy (at least for hybrid attention models like Qwen) is going to be swamped by quantization of the weights. The KV quant we're using here (TQ5/TQ4) is less than 1/3 of the KLD we see when moving from 8 bit weights to 4.6 bit weights (and the KLD is only partially additive, so some of the incremental errors cancel out). There was no statistical significance when testing this KV quant at the task level against 8/8 kv (just trivial variations in sentence length). I will be adding KVaRN in the next release, but with a better codec than currently available elsewhere, so it requires a bit more testing before release.

v0.5 will be focused primarily on the 12GB cards, but this should have generation-wide speed ups, so even if you're not on 24GB, please share your results.

Enjoy!

▲
4
+1
7👁
r/LocalLLaMA · u/whoami-233 · 12d ago
Anyone running multi GPU A100 80GB Cards?

Hey guys, I am looking for benchmarks for people running multi node (4 or more) A100 80 GB cards and seeing what results and models they are getting. Something with VLLM and multi users would be very useful. Or if you know of a place I can find such results please let me know! Appreciated!

▲
216
 
24👁
r/LocalLLaMA · u/jonas__m · 11d ago
Speculative reward hacking in coding agents post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "*Let me look at the problem from the grader's perspective*" and referred to "*hidden tests*", "*test authors*", and "*the checker*".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

\[Pictured example shows verbatim quotes from agent's reasoning\] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

▲
97
 
49👁
r/LocalLLaMA · u/wombweed · 11d ago
I am concerned about all these disparate hard forks that target specific architectures instead of opening a PR against upstream

Other than the obvious self promotion, is there a practical reason people do this that I am missing? There's dozens of llamacpp forks with silly names that are supposedly "optimized" for this or that specific GPU and seem to have zero intention to merge into upstream. Am I missing the real reasons why this happens so often? Why do people think it's OK to do this? In my experience in the open source community this is generally frowned upon.

I don't know if it's just a me problem that this kind of thing puts me off so much. I am usually quite grateful for PR feedback and conscientious about the code I put out there; I take pride in submitting high quality code that meets or exceeds the standards of a given project. Of course there is nothing ethically wrong with hard forks or taking shortcuts if you find the collaborative process cumbersome, but personally I wouldn't promote my fork in such cases, let alone go out of my way to add custom branding with a Reddit announcement post etc. since the effort required to do so seems roughly equivalent to the effort required to meet the contributor standards. In contrast, many of the authors of these forks seem very eager to have others adopt their rebranded fork for production use cases. There just seems to be a big disconnect, idk.

Edit: some great discussion in this thread, thanks to all who responded. Consensus seems to be that (excluding the obvious low-effort engagement bait forks) the base project has to meet many compatibility requirements while a downstream project can be more focused, which is a great point.

💬 166 (+3) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/TheyCallMeDozer · 11d ago
Guy Build a MMORP using Claude... what would it take to do local

As usauly i was scrolling around YouTubes while .... well when every man scrolls YouTube to pass the time.... anyway, came across this video - https://www.youtube.com/watch?v=doR2RhsneRA

TLDR: Guy spends $2175 USD and over 36 hours using Claude Opus 5.5 to build a pretty impressive MMORPG.

Now there is alot of caviats, is it perfect... No... is it really an MMO ... no i havent seen any code added for it.. buttt the strcuture is there, its a hell of a start.

And it got me thinking, if people with their local AI's where to do something like this What models or infra would you use to do this.

Me I think you could get a really good start with Hermes, Qwen Flash, GLM5.3 across a couple of DGX Sparks or if you had 2.5 TB's of RAM and freetoken Kimi K3.

And with the detailed level of prompting and design he laid out prior to actaully letting claude have at it, i think it would doable locally.

So to start the Discussion, what would your tech stack be to do this locally? For me:

\- 2 x DGX Sparks - GLM 5.3

\- 5090 desktop with ComfyUI and a bunch of work floors for image generation, Guassplating, Image to 3D models

and to me I think that would be all that would be needed to get started, but im intrested to see what others think up

▲
3
 
12👁
r/LocalLLaMA · u/Musicheardworldwide · 11d ago
What can I run comfortably?

So I put this computer together after I saw a lot of others posting what they have, and I wanted to know where this ranked and what (in your opinion) are the best local models for coding, and for always on agents.

I didn’t want to type out the whole thing, so I asked my model to give me the specs.

Creature workstation (kernel 7.0.0-28-generic) with an Intel Xeon E5-2698 v4 at 2.2 GHz (20 cores / 40 threads, 50 MiB L3), 125 GiB DDR4-2400 (90 GiB free right now), one NVIDIA RTX PRO 4500 Blackwell with 32 GB of GDDR7 / 31.9 GiB usable VRAM (896 GB/s), and 5.37 TB of raw storage — a 915 GB NVMe root (379 GB free), a 916 GB media disk, and two 1.8 TB drives.

All figures read live from lscpu, free -g, nvidia-smi and df just now.

▲
0
 
13👁
r/LocalLLaMA · u/challis88ocarina · 11d ago
PSA: exercise caution when comparing t/s among models and servers

A "token" is not a fixed chunk of text. It's a word from the model's own private dictionary. Each model ships with its own vocabulary: the list of string-slices learned during training. Analogy: two people transcribe the same sentence: one writes "New York" as one word, the other as two. Both are correct; they're just counting different things.

So tokens/sec is speed measured in steps per minute, and two models can have different stride lengths. One can takes long steps (eg, 4.83 chars each), the other short ones (eg, 3.27). A child and an adult both walking "60 steps per minute" are not walking side by side.

The rules that follow:

  1. Same tokenizer = fair comparison. Two llama.cpp servers running the same model family, tok/s compares directly. Trust it.
  1. Different tokenizers = the number is in different units. Convert to distance: real speed = tok/s x chars-per-token. If both servers report 40 tok/s on prose, one server is laying down \~131 chars/s and the other server \~193 chars/s, so the second is 1.48x faster while the headline numbers tie. The inflation favors the choppier tokenizer: more tokens for the same text = bigger tok/s for the same wall-clock speed.
  1. The conversion factor is content-dependent, so measure, don't assume. A ratio may be 0.68 on prose but 0.74 on JSON. The same vocabularies chop different text differently. Any cross-server speed claim should come with chars/token (or just words/sec) measured on a representative workload, not vendor marketing numbers.
  1. One is real, one is an illusion. What's real is prefill and cost. A more efficient tokenizer turns the same conversation into fewer tokens, so there is genuinely less compute before the first token (shorter TTFT). That's a true speed and money win, not a units trick. What's an Illusion: decode-rate comparisons across families. "Model A does 80 tok/s, model B does 60" says nothing about who finishes the answer first unless A and B share a vocabulary.
▲
1
 
10👁
r/LocalLLaMA · u/DimeRhyme · 11d ago
An UltraFast Qwen3.8 Flash recipe: 74 tok/s, 212 tok/s aggregate on one DGX Spark post image

TL;DR: vLLM recipe for Qwen3.8 Flash on one DGX Spark / GB10. 74 tok/s peak single-stream, 60 to 70 on normal requests, 212 tok/s across 8 streams, 2x to 3.4x faster cold prefill than the recipe it's forked from, full 262K context, and quality matches the original within noise. Everything is open, including the raw per-round data and the benchmark scripts.

https://github.com/dime-online/qwen3.8-Flash-DGX-UltraFast

Most single-Spark setups I've seen posted for this model land somewhere in the 35 to 45 tok/s range, so I spent a few weeks figuring out where the time per token actually goes on a GB10 and cutting it down. If you don't have a Spark, the tricks in the middle section should still be interesting, since most of them apply to any MTP or speculative decoding setup.

What this is, in plain terms

It's a ready-made serving setup. You build the container, pull the public weights, and get an OpenAI-compatible server that answers a lot faster on one box. The speed comes from the model's own draft head guessing several tokens ahead while the full model checks all of them in one pass. Right guesses give you several tokens for the price of one step, and wrong ones get replaced by the full model's answer, so output quality doesn't change.

Decode

Peak decode speed, same workload at every point, best of 3 rounds:

|Streams|1|2|3|4|5|6|7|8|
|:-|:-|:-|:-|:-|:-|:-|:-|:-|
|tok/s|74.1|110.0|132.9|155.5|175.8|191.5|205.8|212.2|

Tokens per step stays between 3.65 and 3.94 from 1 all the way to 8 streams, so the speculation doesn't fall apart under batching. 8 is where it tops out because that's the configured max\_num\_seqs, and the gain from 7 to 8 was down to 3%.

Prefill

Cold prompt with nothing cached, three repeats each:

|Prompt|This recipe|Original recipe|Increase|
|:-|:-|:-|:-|
|16K tokens|4,016 tok/s|1,171 tok/s|\+243%|
|64K tokens|2,426 tok/s|1,071 tok/s|\+127%|
|128K tokens|2,213 tok/s|1,065 tok/s|\+108%|

With prefix caching on, a cached coding prompt starts replying in about 0.57 s, which is what makes agent loops feel fast.

What actually made the difference

The model's own MTP head, run densely, lands about 3.7 tokens per verify step. That's the single biggest lever.

I cut the draft head's vocab from 248K to 65K ids. On a GB10 the draft pass is memory-bound, and reading a full-vocab head every draft step was a real chunk of the step time. The target model still verifies against the full vocab, so this can only change speed, not output.

The quant is W4A16 AutoRound for the MoE experts, FP8 for the side layers and INT8 for the lm\_head. No 3-bit and no NVFP4, because I wanted the speed to come from the serving path and not from squeezing the weights harder.

There's a GB10-tuned low-latency GEMM for the small decode-time matmuls and a sort-free top-k in the verify step.

The prefill gain mostly comes from a faster gather path for the per-layer embedding table, which removes a pile of serial page faults during prefill.

Put together, each decode step went from 68.3 ms on the original recipe to 52.3 ms on an agent-shaped coding workload, about 1.3x more steps per second.

One thing that didn't pay off: doubling the prefill chunk to 16,384 tokens gave no prefill gain at all and ran the box low enough on memory that I rejected it.

Quality

93.1% and 93.3% on a fixed 492-question suite over two seeds, covering code with execution checks, math, knowledge, instruction following, tool calls and long-context needles. I also ran a teacher-forced check against the original checkpoint, and top-1 agreement moved by 0.06 points against a 0.15 point noise band I set before running it.

Practical stuff

The model takes about 71 GiB, the KV pool is 16 GB, and around 16 GiB stays free under load. While generating, the GPU draws about 35 to 37 W median and peaks near 70 W on long cold prefills, with no power or thermal throttling across the soak runs.

The 65K draft vocab was built from English and code, so Chinese, Japanese and Korean output drafts less well and runs slower. Quality isn't affected, because the full model still checks every token.

Built on Saren-Arterius's qwen3.8-Flash-DGX-AutoRound, so big credit there. Happy to answer questions.

💬 13 (+3) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/Robert__Sinclair · 11d ago
The next big company will be...

...the one that will mass produce a cheap device (sub $1000) able to run current sota models at a decent speed.

for that to happen, obviously RAM has to be cheaper, models need to be more efficient and CPUs have to change. It will take time. But as in the 70s/80s computers were huge and expensive mainframes only big companies had, today we are in the same situation.
Fortunately progress happens faster now, so it won't take 30 years to get affordable "home computers". Probably 10, hopefully less.
It's encouraging that today to run the latest qwen 27B you can spend less than $3000. But still...

▲
8
 
14👁
r/LocalLLaMA · u/Dismal-Effect-1914 · 11d ago
B70 no stock/Price Increase

I bought a B70 off Amazon last week for 1300 and have been playing around with it. Today I checked and it seems like I cannot find a single one online for less than 1600 and most places dont have them in stock anymore? What happened? What drives these sudden price increases? It seems like all GPUs across the board have seen another dramatic price flux. Some 5090s I saw were going for 10k!?

▲
28
 
8👁
r/LocalLLaMA · u/jjusko20 · 11d ago
Watch me post-train AliceAI-Foundation-80B-A3B from base to instruct at home, live, on my V100s!

No click bait baby I promise - I'm live streaming the training process kinda like MiMo.

UPDATE: \[Training is paused for an hour or two\] back to training in batches. u/FullOf_Bad_Ideas has pointed out to me I'm burning a ton of compute for nothing on sequence lengths - we'll be breaking the run up into 7/8 batches and then going again.

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wpg4a8/im\_trying\_to\_posttrain/

If you didn't see my original post a few days ago, I'm attempting a slightly more ambitious than usual project in trying to create at least a rough AliceAI-Foundation-80B-A3B-Instruct

I spent the weekend distilling my initial instruct training dataset out of Qwen 3.8 27b, medium thinking - intentionally done because I can run it locally, and I wanted a full dataset in some reasonable amount of time. Still took my v100s running 4x instances at 25tps, like 96 hours of non stop generation to complete the dataset.

I opted not to go for a pre-existing public dataset because I wanted to practice building my own distillation engine (which was configured to work off of an OpenAI compatible endpoint, so it'll distill anything you can hook it up to). The final dataset (this time) consists of 3340 samples: 1760 of general instruct transcripts, and 1580 agentic specific work rows about SWE, harnesses, terminals, etc - I gave the distilling engine a python sandbox and got to simulate turn driven development with a user, and I trained for a bunch of different harness syntax for tool calls, which hopefully will be enough to generalize - gonna run 2 epochs at first.

My GPUs are sobbing right now - turned them on on Friday and left for a weekend vacation, got back today, waited an hour for the data to finish generating, and then immediately fired up the train.

The stuff above is the short version. I'm guessing the initial SFT train will take about 3-6 days, and I plan on working on a RL implementation after I'm satisfied that the SFT has at least worked properly. I am training a rank 16 QLoRa adapter on only q/k/v/o proj, no direct knowledge weight fine tuning.

I thought what MiMo did with their recent training was really cool to watch online, and I like sharing my work with like minded people, and frankly, there's a part of me that's hoping someone will see this and want to hire me (looking for NYC work if you know anyone looking for some passionate ML engineers!) - so I've set up my own little training stream on a cloud flare tunnel.

The stream has the live in progress status of the train, including a live view of the actual data being processed by the model. It also includes way more detail about how I actually designed and generated my training data. Happy to throw the full set on HF as well. I don't expect this model to beat any existing standards but I'll be curious to see if I can get it to operate properly in a harness so I can formally bench it.

I hope you find this interesting! The live stream is a self updating website where you can see exactly what's happening - no need to reload. To watch the training live, visit https://figure-bios-expect-cio.trycloudflare.com/ \[i am currently fixing training issues but it'll be back asap\] -- I'll be keeping it up until the initial SFT is done, at least. The stream lets you inspect the training live as well. This is just a cloudflare tunnel to the trainer.

3 hour update? Loss started at 9ish and is bouncing near 3/4

Update today: back online

▲
0
 
12👁
r/LocalLLaMA · u/cortexist · 11d ago
A hybrid model of Gemma4 with a JEV-like decision head in multi-speaker voice conversation post image

The human brain is neither an LLM nor a JEV. In a crowded market you hear a lot of speech and answer almost none of it. The ongoing question is not “what should I say?” It is “you talking to me?" and "should I say anything at all?”

This live voice demo showcases three hardware tiers—the Blackwell 4500, Jetson Orin NX 16GB, and Jetson Orin Nano 8GB—solving this exact problem. By splitting the workload between a lightweight decision head for turn-taking and a Gemma 4 pipeline for text generation, the setup delivers highly responsive, low-latency vocal interaction.

EDIT: repo (the latest code yet published) https://github.com/cortexist/little-gemma

▲
0
 
14👁
r/LocalLLaMA · u/rawdikrik · 11d ago
Argue with each other for my edu-tainment - 5070 + 5060ti OR RX 7900XTX

I run an Unraid server and use local models for STT, memory, a small llm (I like the new swift bonsai), and SystemOne Models. I currently have a 5070 plugged into my x570 board (with a 5600x), and then the 5060ti on a riser. The 5060 runs at x4, there is a limitation on the board setup.

Running a model big enough for both cards runs SLOW, since the connection to the 5060ti is capped.

Ive tried optimizing with ninfer and vllm, but my speed is capped at the hardware level.

I am considering scrapping the 2 card setup for a single card, and right now the best budget option is the RX7900XTX.

Can you guys argue about what would be the better setup? I dont need the newest models, and I dont need the most speed. I pay for online models. I just like to have a bit of local stuff to help with server stuff. I feel like I spend too much time managing the 2 card setup for not enough to get out of it, and I think dropping to the one card would make things easier without a loss in speed. The idea is to sell both NVIDIA cards, and just get a single card with big enough memory that isnt a pig.

Any advice would help.

▲
0
 
8👁
r/LocalLLaMA · u/inawhole · 11d ago
Gevva0 - a Jev like decision engine on Gemma 26B via direct logit scoring

On the official JevBench evaluation battery, Gevva0 scored 74.63 (#1 global rank), averaging 214ms p50 across large legal contract sets with 82.9% accuracy on the forensic hard tier.

How it works under the hood:

  1. Direct Logit Scoring: Ingests context and reads decision logits directly from llm.scores\[-1\] in a single prefill pass. Fast-path resolution runs in 22ms on short contexts.
  2. Cyclic Debiasing: Permutes class tokens across 4 cyclic positions to neutralize label position bias.
  3. Platt Temperature Calibration: Fits confidence via sigmoid scaling to push Expected Calibration Error (ECE) below 0.03.
  4. Asymmetric Audit Pass: Locks the categorical verdict permanently first, then performs an isolated extraction pass to retrieve verbatim source quotes without contaminating the decision logit.

The repo includes the evaluation harness, raw benchmark datasets, and a local web dashboard: https://github.com/solvingSteve/Gevva0

Setup instructions and benchmarks are in the README.
Working on Demos and Use Cases now so if you have any ideas I'll try to build them next!

▲
12
 
12👁
r/LocalLLaMA · u/Merchant_Lawrence · 11d ago
Need small model that can work for tool caling and agent

Hi. So....... after toturing my 750 ti 4 gb and 16 gb ram with image gen model .i want continue experiment with agent mode like hermes or opencode, using local model but before go i want ask few question. are big model = good perfomance or small model can do same stuff. what small model recommend for agent my spec what caveat of doing this ?

▲
0
 
13👁
r/LocalLLaMA · u/Arany8 · 11d ago
X account claims high t/s setup, but thin on details

According to this post it is possible to reach very high numbers using mtp, however I have failed to reproduce the 50+ tps for 5060ti.

Am I just ignorant or how exactly do this? Or is this a fake post?
Freshly built llama fork for sm120 (Blackwell):
https://github.com/Anbeeld/beellama.cpp
"C:\\llama\\build\\bin\\llama-server.exe" \^

\-m "%MODEL%" \^

\--port 8090 --host 127.0.0.1 \^

\-ngl 99 \^

\--cache-type-k kvarn3 --cache-type-v kvarn3 \^

\--flash-attn on \^

\--load-mode mlock \^

\--jinja \^

\-c 98304 --parallel 1 \^

\--fit off \^

\--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-ubatch-size 128 \^

\-ctkd q8\_0 -ctvd q4\_0 \^

\--kv-tail-tokens auto

Runs at 20-35 t/s.

▲
0
 
12👁
r/LocalLLaMA · u/itsthewolfe · 12d ago
What is the current recommended local model for general use (96GB).

I'm setting up my first build with Open Claw. I'm new to ask of this and starting from zero knowledge.

I've done a lot of reading up, but it's a little overwhelming. So I'm biting things off in chunks.

I want everything to be local. I have a mini PC with 96GB of RAM so can fit a good sized model.

I have Open Claw set up right now with OpenRouter.

My next step is to set up my local model.

What is the current leading open source model for generic tasks and learning? I have plenty of memory to support.

Kimi K3, Opus, Quen 3.8, other?

▲
0
 
15👁
r/LocalLLaMA · u/forevergeeks · 12d ago
Would you buy an AI appliance that removed all the hard work for you

Would you buy an AI appliance that made it easier for you to run local AI models such as Qwen 3.8 27B and Gemma 3 27B?

By easier I mean, the appliance will take care of all the infrastructure stuff for you such as installing the OS, the inference engine such as llama.cpp or vLLM, access management and perhaps include aome preconfigured agents for you start using the system.

The system is multi-user, with a role-based management system, meaning multiple people can use it, including teams.

All runs local, but with the option of using cloud based models if you need more horse power.

Is this something that has an appealing?

Or the fun is the tinkering 🤪

▲
0
 
13👁
r/LocalLLaMA · u/TangeloOk9486 · 12d ago
What models can I run locally on a Mac mini m4 32GB

Hey guys i am planning to get a mac since the GPU and other stuff isnt currently possible for me rn so what models near or stronger than Sonnet 4.6 or somewhat nearer can I use it on mac or would i be able to use it actually?

I mainly need it for coding tasks, different file management and reports and also reasoning. SUggestions or feedbacks are welcomed for models. I want everything local for privacy concerns

Edit: Fixed the model mention

▲
7
 
13👁
r/LocalLLaMA · u/js1943 · 12d ago
LM Studio vs Bionic

I am confused between LM Studio and Bionic.

I have LM Studio for a long time though not used frequently.

Recently I am trying to learn the agentic stuff. Watched a few videos but they were all using Bionic. The strange thing is the interface looks different than the one I just installed today. (Mine seems to be missing features, no developer mode. I am on MacOS)

On the other hand, I seem to be able to find those missing settings in LM Studio.

So what is the difference between the two? Is there anything Bionic can do but LM Studio doesn't?

💬 9 (+5) open on reddit ↗
▲
9
 
10👁
r/LocalLLaMA · u/segmond · 12d ago
Anyone customizing and Optimizing llama.cpp per model?

Basically the idea is take your favorite model, for example qwen3.8-27b or say dsv4vision. Strip everything out that is not needed by that model so the only thing needed is just for the model. Optimize the remaining code to be fast. The idea is to have a model also do this, provide it with enough tools, prompts, docs, guidance. I reckon that if we have a llama.cpp that is optimized for just one model architecture without all the cruft needed to run and. handle other models, that it would not be surprising to easily see 2x+ performance improvement. Anyone thinking along this idea? Again, the goal will be to give this task to a smart model, let it run in a loop, and after a week or 2 you hopefully end up with llama.qwen3.8-27b or llama.glm5.3-flash that would fly.

💬 42 (+1) open on reddit ↗
▲
0
 
6👁
r/LocalLLaMA · u/rm-rf-rm · 12d ago
llama.cpp MacOS menu bar app using blobs instead of GGUF files

Recently started using the MacOS menu bar app for llama.cpp available at https://llama.app When you use the UI to download a model, it seems to do an Ollama-esque hashing instead of just saving the GGUF. Even if you put a GGUF in the Model directory folder, neither the menu bar UI nor the web UI recognizes it. https://preview.redd.it/bsiw7kbn46sh1.png?width=1186&format=png&auto=…

▲
1
 
2👁
r/LocalLLaMA · u/textclf · 12d ago
Introducting TextCLF Quant Factory

Hello, I created a calibration free quant method called TQ. It doesn't need any data so models could be quantized as soon as they come out and the quantized models would generalize better. It performs closely to calibration-based methods. For example, for Qwen 3.8 27B the 4-bit TQ has mean KLD of 0.0282 and top-1 of 92.4% I opened sourced the quant code as Quant Factory so anyone can quantize and run open source models. Right now it only supports 4-bit but I plan to add support for 2-bit and 3-bit soon. The repo link is: https://github.com/textclf-api/quant-factory I have a collection of quantized models using TQ at: https://huggingface.co/textclf You can run these models using either using the following docker image or by following the repo's instruction. For example you can run textclf/Qwen3.8-27B-TQ-4bit like this: docker run --rm --gpus all -p 8000:8000 docker.io/textclf/tq-quant:4bit-main vllm serve textclf/Qwen3.8-27B-TQ-4bit --quantization tq The Dockerfile in the repo shows how this docker image was created. The repo's README explains the approach used for this quant method and why it is useful. You can try it and let me know what you think. Feedback appreciated. EDIT: I did KLD testing using the Wikitext-2 dataset for Qwen 3.8 37B. I also did the same test for the Unsloth-UD-Q4\_K\_XL quant. Here is what I go: |Quant|Disk Size without MTP (GB)|Mean KLD|Median KLD|99% KLD|Top 1% Agreement| |:-|:-|:-|:-|:-|:-| |TQ 4-bit|17.76|0.02823666|0.01282929|0.26897613|92.419%| |UD-Q4\_K\_XL|17.59|0.00771805|0.00318557|0.07390548|95.779%| Working on getting more tests for other models!

▲
0
 
8👁
r/LocalLLaMA · u/Truth-Does-Not-Exist · 12d ago
is DDR5 a scam? My $350 2007 Dell Precision is destroying my $1500 2025 RTX 5070 rig in agentic tasks. made possible by Prism32

I had a theory that ram speed didn't really matter and the only thing that matters is your GPU capacity, vram speed, vram size, and system ram size instead of ram speed or cpu speed. I think I've been vindicated. I used unsloth/Qwen3.8-27B-GGUF:UD-IQ3\_XXS (10.9gb) https://huggingface.co/unsloth/Qwen3.8-27B-GGUF as my baseline and mtp q4\_0 1.37gb only on the dual gpu systems because mtp was too slow on the 9060xt and 5070 I used llama.cpp for all of them and tried to go for the max context I could since they are supposed to be day to day agents. I picked prism32 as my agent harness for this because it's the most compatible, fastest, and reliable one I've found. It's a custom architecture https://github.com/MegaDyneSystems/prism32 which gives it some massive advantages, especially if you want to avoid bloated frameworks eating your context or CPU. The 2007 system literally doesn't work with any other harness because they all require sse4.2 and a ton of heavy dependencies. prism32's only dependency is python 3.7 or above. It's so lightweight (uses around 5 to 10mb ram) that I actually run it bare metal on my ARM synology NAS and even my 2008 TP-link router. If you want to run any agents especially with advanced features on edge or legacy hardware without it choking your system their is no competition I tried these 5 systems: 2007 dell precision t5400 ddr2: 24gb ddr2 8 core 8 thread dual xeon x5460, rx 6700xt rx 6700 22gb vram total, 215gb ssd total system memory 46gb (pic 1) 2009 dell precision t5500 ddr3: 72gb ddr3 12 core 24 thread dual xeon x5675, rtx 5060 rtx 4060 16gb vram total, 512gb ssd, total system memory 88gb (pic 2) 2012 dell precision t3600 ddr3: 64gb ddr3 6 core 12 thread xeon e5-1650, rtx 4060, rtx 3060 12gb total vram 20gb, 512gb ssd, total system memory 84gb 2018 hp obelisk desktop 875 ddr4: 32gb ddr4 3200 i7 8700, rx 9060 xt 16gb, total vram 16gb 512gb ssd, total system memory 48gb (pic 3 ) * 2025 HP omen: 32gb ddr5 6000 RTX 5070 12gb GDDR7 vram 1tb ssd, total 44gb memory (pic 4) The speed results on each system were: 2007 dell precision ddr2 144k context, 18 tk's a second decode, 150tk's prompt processing 2009 dell precision t5500 ddr3 256k context 22 tokens a second decode, 224 tokens prompt processing 2012 dell precision t3600 ddr3 104k context 22 tokens a second short context 14 tokens a second long context, prompt processing is 315 tokens, (could optimize further but tests took me long enough) 2018 hp obelisk 875 ddr4 180k context, 16 tokens a second long decode, 477 promp processing 2025 HP omen 131k 13 tokens a second, 43 prompt processing (yes 43) conclusion The older dual xeon dual gpu setups completely destroyed the newer stuff in context length and speed even on worse GPU's which I think proves my theory, The HP omen system is at least $1500 and the 2007 system didn't cost more than 400 total, rx 6700 xt was $190, rx 6700 was $140, on ebay they are overpriced at $150 although I got it $50 second hand in 2014, and the DDR3 systems were $20 second hand and go around 80 to 150 on ebay, I'd say the ddr3 systems are the best for performance and value, next project is running qwen 3.8 flash next on the ddr3 systems

▲
17
-1
13👁
r/LocalLLaMA · u/jacek2023 · 12d ago
Holo4

*Holo4*\-27B-GGUF is a vision-language model (VLM) for Computer Use, built on the Qwen3.8 dense architecture and developed by H Company. Used with the hai-agents harness, it can send screenshots and tool results to the model, then execute its requested clicks, typing, code, and tool calls.

https://huggingface.co/Hcompany/Holo4-27B-GGUF

*Holo4*\-35B-A3B-GGUF is a vision-language model (VLM) for Computer Use, built on the Qwen3.5 mixture-of-experts architecture and developed by H Company. Used with the hai-agents harness, it can send screenshots and tool results to the model, then execute its requested clicks, typing, code, and tool calls.

https://huggingface.co/Hcompany/Holo4-35B-A3B-GGUF

https://preview.redd.it/rqx7l4qqs8sh1.png?width=1656&format=png&auto=…

https://preview.redd.it/rzez39trs8sh1.png?width=1656&format=png&auto=…

▲
33
-2
21👁
r/LocalLLaMA · u/MooseEfficient2151 · 12d ago
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection

*link to original article*

TLDR: security researcher eddie zhang used a modified+uncensored local qwen 3.8 27b to create an executable capable of dumping LSASS memory for credential harvesting while evading 2 modern EDR security products.

this makes me reflect on how cloud providers keep putting guardrails on everything to the point where even authorized testing gets blocked. local models are the only real option if we want total control, but running heavy local rigs for long agent tasks drains so much compute and management overhead.

been using claude code hooked to sumus to handle my local project workflows and orchestrate tasks in the background, while keeping full local file access on my machine. curious if anyone here is running local uncensored models as local agents for heavy automation, or if you still blend cloud models with local execution setups for your dev environment?

▲
25
-5
23👁
r/LocalLLaMA · u/fallingdowndizzyvr · 12d ago
If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.

Here's the project. I have nothing to do with it. I'm just an amazed user.

https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BEN…

Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat.

"[6204 chunks in 119.0 s | encode: 1239 tok/s | decode: 57 tok/s]"

That's with MTP on. The PP speed in particular is just so fast. That PP speed is twice the speed of the fastest Strix Halo specific fork of llama.cpp I've ever used. Needless to say, the uplift is even greater compared to mainline llama.cpp.

It works with models other than QFN, but the current number is small. You can find the list on their project page.

💬 49 (+3) open on reddit ↗