349 posts · 1 sub · RSS
← prev Sep 27, 2026 → Oct 2, 2026 next →
2026-09-27 → 2026-10-02 hourdayweekmonthyearall
allr/LocalLLaMA
▲
2455
+162
102👁
r/LocalLLaMA · u/BannedGoNext · 10d ago
Anthropic just dropped the greatest advertisement for GLM ever.

Like.. yea bro, I knew GLM was cool. Now everyone does.

💬 531 (+26) open on reddit ↗
▲
2325
+1171
114👁
r/LocalLLaMA · u/StayLameBro · 7d ago
I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. post image

\*\*DISCLAIMER\*\* THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE.

Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context is all that fits next to Qwen 3.8 27B (IQ4\_XS), even with the wired limit raised to 20480. An iPhone 17 Pro Max was sitting in my pocket, so I figured what can I do to make use of this extra silicon.

Turns out a 10 Gb/s USB-C cable & some software is all you need. The Mac runs layers 1–40 of each 256-token batch and streams the activations to the phone. The phone runs layers 41–64 on its GPU while the Mac starts the next batch. The A19 Pro's GPU has matrix units (Metal 4 tensor ops), and they make the phone's half 2.4x faster than the same phone without them.

Same build, phone off vs. on, prefilling a 2,000-token file into a saved agent session:

  • 8k context: Mac alone 132 tok/s → Mac + iPhone 177 tok/s (+35%) (measured two days earlier, same bench)
  • 16k context: Mac alone 109 tok/s → Mac + iPhone 157 tok/s (+44%)
  • 32k context: Mac alone 101 tok/s → Mac + iPhone 130 tok/s (+29%)
  • 48k context: Mac alone 87 tok/s → Mac + iPhone 113 tok/s (+30%)

A fresh 27k-token agent session, cold: 245 s on stock llama.cpp, 228 s on my fork with the Mac alone, and 168 s with the phone.

Past 64k the phone switches jobs. The oldest KV pages move to the phone and the Mac runs all 64 layers. For every attention layer, the phone computes attention over the old keys on its GPU, and the Mac merges that with its own part. While writing, the phone's Neural Engine takes part of that work too: each 16k-key page of old context is compiled into a Neural Engine model with the keys as its weights. At 140k that took writing from 279 to 176 ms per token compared with the phone's GPU alone.

The server allocates 196k–229k of 8-bit context based on the phone's free memory; that's up to \~5.7 GB of KV cache living on the phone instead of the Mac, so the Mac's memory use stops growing at 64k. I've tested a growing session to 128k at 8-bit, with 3/3 planted facts recalled. Separately, at 140k in 4-bit, the run passed the gate with greedy output matching the Mac-only run for 32 generated tokens.

What it doesn't do: speed up writing below 64k. That's the Mac's job. My fork's kernels (SME2 on the M4 CPU and Metal fusions) plus DFlash2 speculative decoding take it from 11.3 tok/s on stock llama.cpp to 25 tok/s at about 30k context with medium thinking, phone or not. SME2 also adds up to 29% to prefill on the Mac alone. Past 64k the phone does share the writing (attention over the old keys), and without it the Mac would have to drop to 4-bit context to reach 128k. In real use I have seen upwards of 30 TPS at lower context.

The phone joins prefills over about 512 tokens. In one real session, that was 7 of 36 requests, but about 83% of the tokens read. Past 64k it holds the context and does the old-key attention, but it stops running layers 41–64 there for now; doing both is next. One request at a time.

I'm curious what this setup could do with newer model architectures. DeepSeek V4.1-Flash reports 890 bytes per token for its global KV cache and adds n-gram embedding tables (Engram). Qwen3.8-Flash-Next, the Qwen 4 architecture preview, has an n-gram lookup table too. Those aren't features of the 27B model I tested, and I haven't benchmarked either architecture here. The real gold is within the newer phones and models working together. With the A20 Pro in the iPhone 18 Pro Max, I bet there is a lot more for me to push.

Code, setup and bench scripts: https://github.com/StayLameBro/backburner

Still a lot of work to do but I built this with Opus 5.5. Happy to answer anything.

💬 331 (+105) open on reddit ↗
▲
1326
+76
72👁
▲
1220
+66
75👁
r/LocalLLaMA · u/Dany0 · 10d ago
AMD's new 256 core EPYC has 16-channel DDR5-12800, 91% memory bandwidth of an RTX 5090

Here are some inspirational quotes you can put into the comments:

  • God is dead and we killed him
  • I am become death
  • All this for 1.5 tok/s?
  • Sir this is LocalLLaMA not RichPeopleofLocalLLaMA
  • Sweet! A 2TB DDR5-12800 RDIMM kit is going to cost only 2 kidneys and a small micronation's GDP
💬 246 (+1) open on reddit ↗
▲
943
+297
94👁
r/LocalLLaMA · u/kvyb · 7d ago
Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following

Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since.

I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly:

incapable of producing more than a few words at a time.
single default personality which no amount of prompting can overcome
will not use tools, at all, whatsoever.
There needs to be a middle ground

They were right. The tool calls didn't actually work, and when people asked it to do something it would sometimes just say it's busy or going to bed. Very human. In a bad way.

So I spent the last three weeks on 2.0. The goal was simple: keep the voice people liked and lose the drawbacks.

What 2.0 does now

  • With no system prompt, it's a normal person texting. Not an assistant, not a catgirl.
  • Give it a character card and it becomes that person, and still texts like one.
  • Ask for a formal email, numbered steps or a proper explanation, and you get exactly that. Then it goes back to texting.
  • Don't want the lowercase texting? Tell it "from now on write in full sentences" (or put it in the system prompt) and it sticks to that until you say otherwise. v1 ignored this completely.
  • It calls tools, and it asks when something is missing instead of making it up. This is the part I'm happiest about. Ask the base model to book a flight without saying where from and it picks JFK. 2.0 asks where you're flying from.
  • It writes code and does math at roughly base-model level.

It's a colleague and a humanlike companion, not an assistant. Use it for chat, roleplay, agents or actual work.

How I trained it

v1 was plain SFT on real and synthetic conversations (139,845 messages from 1,396 conversations). That copies habits, including the bad ones.

For 2.0 I used on-policy distillation. The model writes its own replies and a teacher grades every token. There are two teachers:

  • v1 plus a hidden "text like a person" instruction, for chat and characters;
  • the plain base model, for instructions, tools and code.

The student never sees the hidden instruction, so it learns the behaviour without needing a prompt. Same 27B, a second LoRA on top, merged.

Numbers (vs the model I trained on, huihui-ai's abliterated Qwen3.8-27B; same prompts, same run, thinking off)

|Benchmark|Base (abliterated)|2.0|
|:-|:-|:-|
|IFBench (instruction types I never trained on)|37.3|43.7|
|When2Call (call, ask or refuse correctly)|48|58|
|BFCL irrelevance (don't call a tool when none fits)|60|78|
|IFEval, GSM8K, BFCL simple|81.9 / 89.1 / 97|83.5 / 89.1 / 98 (ties)|

Full chart in the images.

Where it's still worse: knowledge (MMLU-Pro 72.5 vs 78.5) and competitive code (LiveCodeBench 51 vs 56).

Is it actually more human? I built a benchmark for this, "ishuman":

  • It takes 150 fragments from unseen chats.
  • Has each model write the next message.
  • Shows a judge the real message and the model's without labels, and asks which one a person wrote.

|Model|Judge thought it was the real person (50% = can't tell)|
|:-|:-|
|Qwen3.8-27B abliterated (huihui-ai, the model I trained on)|0.3%|
|Same abliterated model + a "text like a human" system prompt|6.8%|
|Qwen3.8-27B official (unmodified, via OpenRouter)|15.1%|
|Qwen3.8-27B-Humanlike-Chat 2.0|23.5%|

So no, you can't just prompt your way there. In a separate test of 16 live multi-turn chats with invented people, 2.0 was picked over the base model 16 out of 16 times.

Links

Big thanks to everyone who left feedback last time, especially the ones who were critical. Tell me where it still sounds like an assistant.

Edit: safetensors are up for vLLM and SGLang:
GPTQ-Int4 (24 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-GPTQ-Int4
FP8 (48 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
BF16 (80 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0

💬 191 (+29) open on reddit ↗
▲
809
+31
56👁
▲
803
+39
63👁
r/LocalLLaMA · u/charles25565 · 11d ago
GPT-3 is discontinued today post image

It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely in our memories. Arguably what's more infuriating is that they suggest using GPT-5.6 Terra as a replacement. Keep in mind that Babbage is a model that's literally 3/4 of the size than MiniCPM5 2B. Even Luna might be overkill as a replacement. But neither is a drop-in replacement. Davinci is the main GPT-3 most people use. This is why we have local models, because they simply cannot have a universal end of life date.

💬 161 (+3) open on reddit ↗
▲
731
+57
70👁
r/LocalLLaMA · u/xenovatech · 9d ago
We just open-sourced the world's fastest WebGPU kernels for local AI on Hugging Face post image

The collection includes kernels for more than 200 common ML operations, all of which can run entirely locally in your browser on WebGPU. We're also working to upstream these optimizations to Transformers.js, ONNX Runtime Web, LiteRT.js, and more!

Kernels: https://huggingface.co/kernels?platform=webgpu
Blog: https://huggingface.co/blog/webgpu-kernels

💬 45 (+2) open on reddit ↗
▲
725
+223
109👁
r/LocalLLaMA · u/LegacyRemaster · 7d ago
Does anyone know if any new releases from Mistral are planned? post image

It’s been a long time since the last models came out. I notice they are selling GLM on the site, and I wonder if they are developing something, given the long silence.

💬 333 (+50) open on reddit ↗
▲
686
+22
36👁
▲
645
+5
26👁
r/LocalLLaMA · u/Available_Pressure47 · 13d ago
42x Faster Prompt Lookup Drafting in llama.cpp
💬 181 (+1) open on reddit ↗
▲
629
+27
45👁
▲
622
+17
52👁
r/LocalLLaMA · u/am17an · 12d ago
Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy

Meta came out with a banger paper https://arxiv.org/pdf/2606.00206, but it did not look at various quantizations supported in llama.cpp. So I did a run on 50 random MATH-500 questions (https://huggingface.co/datasets/HuggingFaceH4/MATH-500) and ran it on various quantizations of https://huggingface.co/bartowski/Qwen\_Qwen3.5-4B-GGUF and tried

--logit-bias 466-2 --logit-bias 694-2 --logit-bias 1362-2 \
--logit-bias 1412-2 --logit-bias 1921-2 --logit-bias 1990-2 \
--logit-bias 2086-2 --logit-bias 2361-2 --logit-bias 2441-2 \
--logit-bias 2493-2 --logit-bias 2892-2 --logit-bias 3222-2 \
--logit-bias 3315-2 --logit-bias 3384-2 --logit-bias 3404-2 \
--logit-bias 3482-2 --logit-bias 3655-2 --logit-bias 4213-2 \
--logit-bias 4370-2 --logit-bias 4598-2 --logit-bias 4611-2 \
--logit-bias 4808-2 --logit-bias 5752-2 --logit-bias 6970-2 \
--logit-bias 7014-2 --logit-bias 7643-2 --logit-bias 8106-2 \
--logit-bias 10179-2 --logit-bias 10451-2 --logit-bias 11746-2 \
--logit-bias 13264-2 --logit-bias 13428-2 --logit-bias 14673-2 \
--logit-bias 15029-2 --logit-bias 16036-2 --logit-bias 21143-2 \
--logit-bias 21979-2 --logit-bias 33955-2 --logit-bias 35999-2 \
--logit-bias 36563-2 --logit-bias 37201-2 --logit-bias 37781-2 \
--logit-bias 41484-2 --logit-bias 62586-2 --logit-bias 66073-2 \
--logit-bias 73071-2 --logit-bias 84485-2 --logit-bias 85152-2 \
--logit-bias 95500-2

these correspond to the paper's overthinking markers:
\[

" perhaps", " maybe", " wait", " Wait", " actually",

" hold", " Hmm", " hmm", " Alternatively", " alternatively",

" However", " however", " instead", " Instead", " But",

" but", " though", " although", " yet", " rather",

" unless", " otherwise", " nonetheless", " nevertheless", " regardless",

" still", " anyway", " Or", " or", " either",

" whether", " uncertain", " unsure", " possibly", " might",

" could", " another", " different", " reconsider", " rethink",

" backtrack", " retry", " revisit", " doubt", " confused",

" wrong", " mistake", " error", " incorrect"

\]

Here are the results, surprisingly even BF16 leads to better accuracy. Caveats being this is one test on one model. Try it out and see it helps!

|Format|Accuracy: baseline → penalty|Reasoning tokens|
|:-|:-|:-|
|BF16|74% → 84%|−19.4%|
|Q8\_0|76% → 80%|−11.0%|
|Q4\_K\_M|60% → 66%|−14.8%|
|Q3\_K\_M|52% → 66%|−17.5%|
|Q2\_K|12% → 24%|−11.5%|

💬 135 (+7) open on reddit ↗
▲
581
+186
63👁
▲
506
+118
79👁
r/LocalLLaMA · u/paf1138 · 7d ago
New in llama.cpp: Decision Models
💬 123 (+14) open on reddit ↗
▲
458
+24
68👁
r/LocalLLaMA · u/Rombodawg · 9d ago
Least to most expensive (Somewhat modern) GPU's with 32gb of vram (Under $1600) Based on ebay listings post image

I was researching prices on ebay and fed claude a bunch of images of listings. I had it make a chart and thought it would be useful to share.

💬 260 (+24) open on reddit ↗
▲
458
+70
56👁
▲
446
+7
48👁
r/LocalLLaMA · u/speedb0at · 12d ago
The Opus 5.5 posts about motion graphics are cool but Qwen 27B made this on a 4090

Saw the hundreds of tweets where people just keep asking Opus 5.5 for motion graphic videos. Decided to ask qwen to look at them and make its own. Quite amazing what local can achieve.

\*\*EDIT\*\* It looks laggy because of reddits .gif limit btw

the full high res version (with sound) is here: https://x.com/mkultraware/status/2104192428664127555

Promted and built in: https://github.com/mkultraware/accuretta

https://i.redd.it/5gmotgxx82sh1.gif

💬 123 (+2) open on reddit ↗
▲
426
+6
47👁
▲
410
+28
45👁
▲
404
+11
52👁
r/LocalLLaMA · u/professormunchies · 12d ago
Qwen plays World of Warcraft post image

Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without installing the game and it has mobile controls. Afterwards, created a custom mcp to drive the client and have finer game control than a generic browser agent. The agent harness can plug into your local or cloud LLMs and be used to drive the game. For best results have a model that can output >50token/sec. No visual input is used in the making (might be beneficial in the future but incur more latency). The mcp and agent are only running on my dev server but if folks are interested in trying the game go to https://jankcraft.xyz/

Still vibing but I’ll make some more content of it … I think those Pokémon benchmarks have become a little too easy and they need a new challenge like speed running to 80 in wraith of the lich king.

💬 128 (+7) open on reddit ↗
▲
343
+47
59👁
r/LocalLLaMA · u/-p-e-w- · 8d ago
Heretic is on PewDiePie!

So I haven’t played a computer game in 20 years, and I know nothing about Minecraft, and I definitely prefer classical literature over YouTube culture, but even I have heard about the individual called PewDiePie, for two reasons:

  1. His monicker starts with the initials of my own name
  1. I remember a recurring Internet meme a few years ago where he was competing for the most subscribers with an Indian film music channel

I had never watched a single one of his videos, however.

Well, until today, when people started spamming me with messages informing me that Mr. Kjellberg aka PewDiePie has tried out Heretic and made a video where he talks about it:

https://m.youtube.com/watch?v=ODDJXGY_1kQ

(Heretic mentioned around 9:00)

Obviously I’m thrilled that a less technical audience is being exposed to my work, and the more people understand what is possible the better. I expect to be receiving a couple hundred more mails in the coming days asking how to run Heretic on ChatGPT (you can’t), or accusing me of working for the CIA (I don’t), but other than that, the more the merrier I guess 😏

Heretic 2.0 coming soon…

💬 90 (+3) open on reddit ↗
▲
338
+14
58👁
r/LocalLLaMA · u/politefella0 · 10d ago
Deepseek Harness app is out now!!!

Downloading now.

💬 91 (+5) open on reddit ↗
▲
312
+12
58👁
r/LocalLLaMA · u/writesfw · 10d ago
Are you worried about a potential ban of Chinese open weight models?

Anthropic released the GLM article today. Trump is getting very involved.

Do you foresee Chinese open weight models getting banned soon?

💬 506 (+12) open on reddit ↗
▲
283
+15
51👁
r/LocalLLaMA · u/LegacyRemaster · 11d ago
Qwen next 3.8 and 3.8 27b Vs Sonnet 5.5 low and Sonnet 5.5 medium. post image

Six months ago, a result like this was unthinkable. But now we can say it loud and clear: local models are at the cutting edge, and the gap of just a few months has been confirmed.

Personally, I use Qwen-Next 3.8 for complex tasks; today, GPT-Sol-6-High was messing up a project, but Qwen-Next got it back on track. I consider it a reliable benchmark. What’s your take?

💬 117 (+3) open on reddit ↗
▲
271
+75
43👁
r/LocalLLaMA · u/Recoil42 · 7d ago
New Architecture from Percepta: Spotlight — separating intelligence from memory, allowing knowledge and skills to grow without changing the model's weights. post image

https://www.percepta.ai/blog/can-llms-grow-their-own-capabilities https://www.percepta.ai/blog/spotlight-memory "Our new architecture, Spotlight, replaces attention with a memory that escapes this trade-off: it is the first architecture to achieve infinitely growing memory without increasing the access cost. Every token reads from and writes to an unbounded memory, but because the model learns to index individual memory cells, each token only touches a small number at a time. While other sparse architectures fix the fraction of capacity used at each step—a mixture-of-experts model, for instance, always activates the same number of experts out of a fixed set—Spotlight is arbitrarily sparse, touching the same number of cells regardless of how the memory grows. The fraction of memory it uses can shrink as far as we want. Spotlight separates an intelligence module, which performs computation, from memory, which holds knowledge, procedures, and working state. The intelligence module stays the same size, and the weights don't change as memory grows. The memory is writable, and the model itself decides what to load and when to overwrite it, token by token. Because memory can hold skills as well as facts, the model can gain new capabilities without retraining: what it can do is not limited by the size of its intelligence module."

💬 32 (+4) open on reddit ↗
▲
270
+20
26👁
r/LocalLLaMA · u/Spiritual_Impress_30 · 9d ago
Thank You, Mradermacher.

best iq quants in the biz, got me gemma 4 26b to run 75tok/s tg and 1500 pp on 2x 4060 8gb using lmstudio serving to hermes, much work has been done.

▲
262
+21
45👁
r/LocalLLaMA · u/jacek2023 · 10d ago
Reflection 70B was released two years ago (September 2024)

You may think that jev, OpenClaw or TurboQuant are super cool, but actually the coolest LLM invention happened two years ago

As we all know, the best source of reliable information about LLMs is YouTube:

https://preview.redd.it/4tv0ctbwyfsh1.png?width=2544&format=png&auto=…

Back in September 2024, Reflection 70B appeared out of nowhere and was announced as an open-source model that supposedly destroyed GPT-4o

There was only one small problem. People downloaded it. And tested it :(

https://preview.redd.it/kamhrp9izfsh1.png?width=1514&format=png&auto=…

It turned out that Reflection 70B was basically a Llama 3.1

https://preview.redd.it/ghmof4l50gsh1.png?width=1524&format=png&auto=…

but at the end the mystery was solved

https://preview.redd.it/44eb3jw90gsh1.png?width=1530&format=png&auto=…

Let this be a moment of reflection on the current hypes in LocalLLaMA.

great summary by Maziyar PANAHI https://x.com/MaziyarPanahi/status/1838559480658710982

▲
255
+5
31👁
r/LocalLLaMA · u/jacek2023 · 11d ago
nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 · Hugging Face

Model Developer: NVIDIA Corporation

Model Development: Fine-tuned from NVIDIA-Nemotron-3-Ultra-550B-A55B

[](https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-…)What is Nemotron?

NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.

[](https://huggingface.co/nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-…)Description

Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra, fine-tuned for one epoch on 477,642 synthetic reasoning traces distilled from GLM-5.2 across 22,000 curated problems spanning 16 regional and international competitive-programming contest families. Selected as the SFT teacher for its higher accuracy and roughly 30% shorter generations compared to a DeepSeek-V4-Flash-trained variant, GLM-5.2 distillation yields a model that, combined at inference time with GenCorrect — an iterative closed-loop test-time compute strategy that generates diverse candidate solutions, incorporates evaluator feedback, and refines subsequent generations under a fixed submission budget — was evaluated live and prospectively on the IOI 2026 problem set under official contest time, internet-access, and submission constraints, scoring 535.4 out of 600 and surpassing both the gold-medal threshold (361.12) and the top human contestant's score (498.27), making it the first AI system reported to outscore the highest-scoring human contestant on an IOI problem set.

This model is ready for commercial or non-commercial use.

▲
255
+17
52👁
r/LocalLLaMA · u/jacobpederson · 8d ago
Why am I like this? (Full Chat and Image generation on a 286 Tandy 1000 TL/3) post image

40 year tech gap? No problem! The Tandy runs DeskMind, a native DOS program. It talks over WiFi (a PicoMEM 2 card with mTCP) to a small Python server on my PC. That server drives Qwen3.8-27B (NInfer on a 5090) and Krea 2 (ComfyUI on a 4090). The 286 never sees JSON, base64 or a PNG. It gets plain text lines and pictures that are ready to copy into video memory.

Drawing from chat without tool calling. The system prompt tells Qwen to wrap a picture request in \<draw>...</draw>\. The server catches the tag mid-stream, runs Krea 2, dithers the result, and streams a \picture ready\ line. "Draw me a 286 AI logo" takes about 9 s from Enter to a thumbnail in the chat.

Qwen Vision sees what the Tandy sees. When you ask about a picture, Qwen gets the original and the 16-colour dithered version, so "why does the sky look striped?" has context. The latest picture stays in context for follow-ups.

The model knows where it lives. The system prompt knows it's talking IN a 286 with 80 columns and 16 colors. It keeps answers short and plain ASCII, and when asked about games it suggests Wolfenstein 3D or Commander Keen.

Streaming cleanup for a 1990 screen. Reasoning is stripped, Markdown is removed on the fly, Unicode becomes code page 437 (bullets turn into the CP437 block character), and tiny tokens are merged into \~48-character lines so the 286 isn't redrawing for every token.

Per-request reasoning effort: low for chat (replies start in \~2 s), medium for rewriting image prompts.

Prompt "enhancement" tuned for dithering: bold shapes, strong contrast, simple backgrounds. The rewrite shows up in an edit box on the Tandy before drawing, and the rules themselves can be edited from the Tandy.

\- \*\*A Dither Lab\*\* in the server GUI: Floyd-Steinberg, Atkinson, Bayer, Yliluoma and more, previewed at the Tandy's real aspect ratio.

Numbers: Krea 2 at 1024x768 in \~10 s 8 steps, the target is 640x200). About 1 s to send a 64,000-byte picture over the PicoMEM WiFi (56-79 KB/s).

Code (GPLv3): https://github.com/RowanUnderwood/DeskMind

added image gallery https://imgur.com/a/wysqoM1

💬 129 (+12) open on reddit ↗
▲
245
+12
29👁
r/LocalLLaMA · u/LambdaHominem · 9d ago
AI CEO Interviews (2026) post image
💬 34 (+1) open on reddit ↗
▲
229
+25
45👁
r/LocalLLaMA · u/Significant-Price695 · 9d ago
Oído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)

I'm part of the Lokutor team that built this.

Model: NVIDIA Conformer-CTC Small (13M params, int8). It runs on an ESP32-S3 with 8 MB PSRAM, no GPU or NPU. LibriSpeech WER is 3.7 / 8.2, versus 6.3 / 15.9 for Whisper tiny.en on a laptop. Under real noise (DEMAND: car, kitchen, cafeteria) plus babble and reverb, mean WER is 8.4 vs 12.1 for Whisper tiny.en. You can try the exact chip arithmetic on your laptop mic with live_demo.py. https://github.com/lokutor-ai/oido

💬 40 (+1) open on reddit ↗
▲
222
+13
34👁
r/LocalLLaMA · u/Educational-Care7867 · 11d ago
ImaJev-4b: I spent 15 days fine-tuning a 4B model to make business decisions from text and photos, and it just ranked #1 of 91 on JevBench & ahead of GPT-5.6 Luna on DecisionBench post image

Some context first.

I am process improvement / business consultant and had worked with Fortune 500 companies on improving their processes around refunds, returns, customer support, etc.

This entire thing had a lot of complex decision making and generally every decision / condition node in a process map was generally replaced by a human because factoring in ambiguity in a code is very difficult.

Idea of ImaJev

Hence, when Jev came out, I was very intrigued with it and also could clearly see its use-case of improving decision making in complex decision work flows.

However, Jev didnt have support for Images and I thought that it can be replicated for both Text and Images in a single model and thats when I started with ImaJev.

Training Process

It went badly at first. My first big fine-tune on about 500k short decisions made the 9B model worse at reasoning: 64.9 down to 42.3 on JevBench hard. It had basically learned to pattern-match. I spent the next couple of weeks generating hard questions with open-weight models and only keeping the ones where two Ai models agreed on the answer. That brought it back.

Results

Then, on the JevBench - It came out #1 of 91 (v1.4.2.2, scored 27 Sep), 67.37 vs Jev 1.13.0 at 63.29.

The same week DecisionBench put it #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1.

I honestly didn't expect either.

To be fair about it: the #1 is on a score that weighs accuracy, calibration, speed and cost equally. On accuracy alone it's #3. Its main strength is that when it says 90% it's usually right, and it'll say "can't tell" instead of guessing.

What it actually is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions.

It gives back a probability for each option plus "unknown", in one forward pass.

Runs on a Mac with MLX or on one GPU.

The whole project costed me around $1200 in rented GPU and a lot of time :P

I would love to know your thoughts on it - it anyone would be interested to try that.

💬 69 (+2) open on reddit ↗
▲
220
+56
55👁
r/LocalLLaMA · u/jacek2023 · 7d ago
microsoft/FrogNano-4B-2609 · Hugging Face

An agentic model from Microsoft for the GPU poor

https://huggingface.co/bartowski/FrogNano-4B-2609-GGUF

FrogNano is derived from Qwen/Qwen3.5-4B, a general-purpose post-trained model designed for language, reasoning, coding, agentic, and multimodal tasks. FrogNano inherits Qwen3.5-4B's dense 32-layer hybrid Gated DeltaNet and gated-attention architecture, but its additional post-training is text-only and focused on repository-level software engineering. The model is further trained using reinforcement learning on approximately 1,500 synthetic SWE task environments generated and calibrated against the evolving policy using TaskPilot. Training uses the five-tool Leaf harness and executable test-based rewards over complete multi-turn coding trajectories.

The additional post-training is intended to improve long-horizon repository navigation, debugging, code editing, test execution, and patch generation in a compact 4B model. Unlike approaches based on behavioral distillation, FrogNano does not train on stronger-model solution trajectories, actions, reasoning traces, or patch targets. This specialization also introduces limitations and risks: performance is sensitive to the Leaf harness and test quality, training data are Python-heavy and primarily English, and generated patches may be incorrect or insecure despite passing available tests. When integrated with the Leaf harness, FrogNano generates structured tool calls that can propose repository changes. Leaf executes authorized tool calls within an isolated repository environment to produce a candidate patch; FrogNano does not itself deploy the changes. Any resulting patches require human review, regression testing, and security validation before use or deployment.

💬 69 (+14) open on reddit ↗
▲
218
+4
32👁
r/LocalLLaMA · u/Acceptable-Cycle4645 · 10d ago
Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models post image

I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.

Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.

And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.

The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.

All figures here: https://github.com/0xShug0/audio.cpp/tree/main/assets/figure/

💬 24 (+1) open on reddit ↗
▲
218
+8
47👁
r/LocalLLaMA · u/jacek2023 · 9d ago
add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp

now you can use GLM-5.3-Flash on your home computer

💬 89 (+10) open on reddit ↗
▲
216
 
24👁
r/LocalLLaMA · u/jonas__m · 11d ago
Speculative reward hacking in coding agents post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "*Let me look at the problem from the grader's perspective*" and referred to "*hidden tests*", "*test authors*", and "*the checker*".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

\[Pictured example shows verbatim quotes from agent's reasoning\] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

▲
207
+3
26👁
r/LocalLLaMA · u/jacek2023 · 8d ago
Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp

now you can use MTP with Qwen Flash Next, time to switch from Qwen 3.8 27B?

(merged after 17h of development)

quants: https://huggingface.co/ggml-org/Qwen3.8-Flash-Next-GGUF

link to the previous discussion (I deleted the old post to avoid duplicates): https://www.reddit.com/r/LocalLLaMA/comments/1wur4lt/qwen\_flash\_next\_mtp\_work\_restarted/

▲
198
+36
66👁
▲
186
+6
16👁
▲
184
+7
31👁
r/LocalLLaMA · u/Educational_Sun_8813 · 9d ago
Preorder for new AMD Ryzen™ AI Max 400 Series 192GB from framework just started

Framework Desktop
Framework Desktop DIY Edition (AMD Ryzen™ AI Max 400 Series) 192GB

💬 129 (-1) open on reddit ↗
▲
184
+8
34👁
▲
181
+8
46👁
r/LocalLLaMA · u/Creative-Type9411 · 8d ago
Finally got my 4th card in (64gb total) post image

I was waiting on the blowers for the T4s and posted this before it was finished, other than some braided cable sleeves for the fan wires its pretty much good, I was going to upgrade the CPU, but I'm getting great speeds comparatively to a CPU in my old box that had way more cores, so I don't think it's going to make a difference.

Fractal Design Torrent Mid-Tower Case w/Tinted Glass
SuperMicro X11SPA-T Motherboard
Xeon W3225
768GB DDR4 ECC 2666
4xTesla T4 16GB GPU
4x1tb Samsung 870 EVO SATA SSD Raid

Ubuntu 26.04/llama.cpp/openwebui+custom powershell harness

now i want more cards 👀

💬 108 (+5) open on reddit ↗
▲
179
+16
54👁
r/LocalLLaMA · u/ea_man · 8d ago
Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now post image

pi-llama-skip-reasoning is an extension for the Pi.dev harness that forces a local llama.cpp model to stop reasoning and answer / act immediately.

When you are deep into the ctx session and ask 27B a simple question about a fact or need a direct action, the model may still feel the urge to indulge in copious deliberation in the reasoning trace. This extension allows the user to force the model to snap out of the reasoning stage and provide the answer immediately.

Disclaimer: don't skip the reasoning for important problem-solving, that would hurt quality.

This uses the same mechanism the llama.cpp web interface uses to skip reasoning, so it's native to llama.cpp, this extension is meant for Pi.dev yet the same mechanism could work for other harnesses.

Usage: /skip-reasoning command or shortcut Alt+T ,
Install: pi install npm:pi-llama-skip-reasoning

\- https://pi.dev/packages/pi-llama-skip-reasoning

💬 82 (+4) open on reddit ↗
▲
176
+11
24👁
r/LocalLLaMA · u/Ok-Breadfruit-3523 · 11d ago
I’m calling this the Monstrosity. 5 ex mining BC-250 boards Qwen3-Coder-Next Q4 at 40 tok/s post image

Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2 more of these and will probably get them running to see if 3.8 flash next runs at usable speeds. This setup is wildly inefficient with power but cost me less than $800.

▲
163
+6
31👁
r/LocalLLaMA · u/WebAssemblyMan · 9d ago
DeepSeek now trained on Ascend 950 post image

26 months ago Liang Wenfeng said:
"Someone must step onto the frontier."

Now they are training their models on Ascend 950

▲
161
+3
50👁
r/LocalLLaMA · u/MLDataScientist · 9d ago
Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine

I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling and the developer fixed them.

Inference engine (only runs on Nvidia for now; AMD support is experimental): https://github.com/Niko1221/Strata

Here are some metrics with screenshots. My laptop has 5070ti 12GB VRAM, 64GB ddr5 RAM, Intel 275HX CPU, gen4 SSD.

Aquarium test \(unsloth studio connected via local API\)

The model I used was https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/tree/main/IQ3\_XXS which has a good quality for its size. Above, the model generated the aquarium test. At 43k context depth, it was running at 51 t/s. Stock llama.cpp reached only 23t/s with the same quant.

32k context read at 1500t\/s \(unsloth studio via local API\)

This quant could only reach 100t/s PP with stock llama.cpp using the same quant. Strata was reading 32k context text at 1500t/s. This is way above my expectation. This quant can load with up to 200k context at 8bit. However, I was only using 131k context.

Memory utilization

As you can see it is utilizing 11GB VRAM and 56GB RAM (includes system/OS programs).

This engine is specifically built for one model only and only select ggufs (ISTA-DASLab) work with it. You can use IQ3\_S from ISTA-DASLab which they claim recovers full model's performance on coding benchmarks. I tested IQ3\_XXS for some time and I would say it is an excellent model.

I never thought 12GB VRAM would be enough to run frontier models from 6 months ago locally on a laptop. What a time to be alive!

💬 143 (+3) open on reddit ↗
▲
157
+1
32👁
r/LocalLLaMA · u/soyalemujica · 9d ago
If one hour of AI is costing me 0.12€ is paying for frontier a cheaper option?

Running Qwen flash next of even Qwen 27b dense, I can do any,burning sticking to flash due to its speed, and the kwh cost is at 0.25€ where I live in, ranging from 0.11€ to 0.35€, so I used chatgpt to help me calculate the total kwh consumption on my 7900xtx plus 9800x3D, and well that is the result.

Judging by this, if deepseek flash is indeed then faster to use per 1m token, does it mean that frontier is cheaper for me or am I calculating something wrong ?

💬 175 (-1) open on reddit ↗
▲
147
+9
32👁
r/LocalLLaMA · u/nullmove · 12d ago
Naive-N0.5-Flash - 309B-A15.5B

https://naive.ai/en/research/

  • Built for coding and AI R&D
  • 1M context
  • Hybrid SWA/DSA
💬 38 (+2) open on reddit ↗
▲
138
-1
30👁
r/LocalLLaMA · u/L0ren_B · 12d ago
Another "Harness matters" post (codex cli > pi and opencode)

I run my own LLM while also having a Openai subscription. Also tried DeepSeek (latest flash now). I run Qwen 3.8 flash Next at an amazing speed on my 2x3090 + Ram!

But local LLM never did worked for me outside some demos like build me a "3D Mario Game, multistage" which I've been using to test LLM's for a long time. At serios work, they never even compared with GPT 5.2 or lately, 5.6 Luna, which is worse in the benchmarks.

Until last night! I've asked gpt 5.6 luna to configure codex cli for local llm! (I've been using pi.dev and opencode until now) and the results amazed me! Suddenly AA benchmark made sense!

First test: The 3D Mario prompt test in Codex Cli blew me away. The best until now!

But real work is where you can see the difference! I took same project that Luna was working for days , and give it to both in paralel! And Qwen 3.8 Flash Next ran circle around luna. Previously, it failed to deliver results, with Qwen and DeepSeek as well in this project.

Now, I could say it's my go-to model!

P.S. For weebsearch, I've asked to port the pi-smart-web-search to codex as a skill. It works amazing! (I should put it on git later).

Maybe, I was using pi.dev wrong. Maybe there is an extensions that brings the same quality to it as Codex Cli. Does anyone know?

💬 224 (+4) open on reddit ↗
▲
138
+26
47👁
r/LocalLLaMA · u/matteiuspi · 7d ago
Two 96 GB Ascend cards crun Qwen3.8-flash-next hardware notes, vLLM work, benchmarks, and what is next

I have been building a somewhat unusual local inference machine around two Huawei Atlas 300I Duo cards. They are relatively inexpensive, passive, dual-accelerator PCIe cards with 96 GB of device memory apiece. They are also absolutely not drop-in CUDA replacements.

When I first brought up Qwen3.8 Flash-Next these past two weeks, it was often incoherent and lived around 1 generated token per second. Some runs were below that. Today the same two-card machine is producing coherent output at roughly 30 tok/s for one request and about 61 tok/s aggregate at four-way concurrency on my short decode benchmark. It also completed the full 198-question GPQA Diamond set.

This post is the start of a guide for these cards: what the cards physically are, how I cool them, what “96 GB” really means, what I changed in vLLM and vLLM Ascend, which optimizations actually mattered, and which problems are still open.

The short version is the hardware is capable. My work has been mostly on the software stack, ubuntu-26.04 driver support, model architecture support, memory layout, custom operators, and getting every asynchronous state transition exactly right.

The hardware: one card is really two devices

My current machine has two Atlas 300I Duo cards, which enumerate as four Ascend 310P3 devices.

Current system / Planned system

Physical cards: 2 / 3

Ascend devices/npus/AIcpus: 4 / 6

Nameplate device memory: 192 GB / 288 GB

Approx. runtime-visible memory with this configuration: 172 GiB / 258 GiB

Combined maximum accelerator-board power: 300 W / 450 W

Each card has two accelerator SoCs and 96 GB of LPDDR4X in total, or 48 GB local to each chip. It is not one unified 96 GB allocation. A model that does not fit on one 48 GB device still needs tensor, expert, pipeline, or another form of model parallelism. The card is PCIe Gen4 x16, full-height/full-length, and a surprisingly thin single-slot design. Huawei rates it at 408 GB/s aggregate memory bandwidth and 150 W maximum board power. The official specifications are here (https://support.huawei.com/enterprise/en/doc/EDOC1100285916?section=j00e).

Also, despite the generic “HBM” terminology used by a lot of accelerator software, the memory on these cards is LPDDR4X.

This two-chip-per-card layout matters. Communication within a model still goes through the distributed runtime, and memory remains local to a rank. I use HCCL collectives and explicitly map tensor and expert ownership across all four chips. Thinking of the machine as four 48 GB ranks is much more useful than thinking of it as two 96 GB GPUs.

https://reddit.com/link/1wvt1m4/video/7rju7qbrw1th1/player

Passive cooling is not a deal-breaker

The cards have large heatsinks and no onboard fans. They were designed for server airflow, so putting them in an ordinary workstation and hoping a rear case fan will sort it out is a bad plan. There is a useful teardown here (https://videocardz.com/newz/huawei-atlas-300i-dual-ai-gpu-with-96gb-memory-worth-1400-has-been-taken-apart) if you want to see the heatsink and heat-pipe arrangement.

I give them direct, high-volume airflow and run the room on AC/heat-pump cooling. Under real multi-hour model loads, the cards can crank continuously without drama. Across my recorded Qwen runs, peak device temperatures were generally 72–78 °C. My watchdog limit is 96 °C, and the cards have not approached it.

So I would not bat an eye at adding another passive card. The actual checklist is mundane:

• Keep unobstructed airflow through the heatsink fins.

• Make sure the chassis fans have enough static pressure.

• Budget another 150 W of board power per card, plus the rest of the host.

• Exhaust the heat from the room instead of recirculating it through the rack.

• Log temperature during long prefill, decode, and concurrency tests rather than trusting an idle reading.

Passive does not mean low-power or self-cooling. It means the chassis and room are the cooling system. Once that is handled, these have behaved like ordinary 150 W server cards for me.

Note: Nothing heats up these cards more than loading/moving things around in their ram -- the npus at full utilization run cooler than large block memory assignments. We keep this in mind when optimizing the model serving code paths.

ECC, nameplate memory, and what is actually usable

My cards arrived with ECC enabled by default. I disabled it to reclaim device memory. This is an inference and development box, and I consciously prefer capacity over ECC protection here. That is a reliability tradeoff, not a universal recommendation.

Even with ECC disabled, firmware, the runtime, communication buffers, graph captures, workspaces, and allocator reservations consume memory. In practice, the software sees roughly 43 GiB per 48 GB chip. Four chips therefore provide about 172 GiB of useful aggregate capacity, but it is still four separate local pools. The exact free number also changes with the CANN build and launch configuration.

That distinction has shaped nearly every model decision. The question is not only, “Does the checkpoint total fit in 192 GB?” It is, “Does each rank's weight shard, recurrent state, KV/cache allocation, graph capture, collective workspace, and worst-case temporary allocation fit in its own 43 GiB?”

Why I forked vLLM as well as vLLM Ascend

The public work lives in the OpenSensor vLLM Ascend fork (https://github.com/opensensor/vllm-ascend) and paired vLLM fork (https://github.com/opensensor/vllm). I needed both sides because this was not just a missing device kernel.

I am currently the only person developing these forks. The software bus factor today is one. I have made a lot of progress, but a fast-moving one-person fork should not be confused with the maturity, test coverage, or support depth of mainline vLLM on NVIDIA.

I also ran into a bizarre tooling problem: in my sessions, Claude repeatedly refused to engage with prompts about this architecture because the cards are Huawei hardware from China. These were ordinary engineering discussions about serving, sharding, cooling, and performance—not requests to build a restricted application. I am describing my direct experience rather than claiming that every Claude version or account will behave identically, but it made Claude unreliable as a development assistant for this project. Whatever anyone thinks about the politics, that is a real practical constraint when choosing tools around this hardware.

Qwen3.8 Flash-Next combines MoE routing, Gated DeltaNet recurrent layers, sparse quadratic-attention layers, packed low-bit experts, long context, and an MTP draft model. Supporting that cleanly touched model integration, the v1 runner, cache accounting, scheduling, graph capture, distributed state, model loading, and Ascend-specific operators.

The current development sprint has been roughly two and a half weeks of nearly continuous bring-up and optimization, with hundreds of fork commits, repeated full checkpoint loads, profiler captures, operator microbenchmarks, and multi-hour quality runs. This was not one magic kernel patch.

What took Qwen from incoherent \~1 tok/s to where it is now

These are the architectural changes that moved the needle.

  1. Make the hybrid model correct before making it fast

The early model could generate tokens, but generation was not a correctness test. I found failures that only appeared at production geometry: incorrect Gated DeltaNet gate-vector handling, recurrent-state precision and lifecycle problems, incomplete sparse-attention score width, and mismatches between the host operator API and the installed kernel package.

One particularly nasty GDN issue looked fine in small-head tests but accumulated state error across the real 36 recurrent layers and produced incoherent text. Keeping recurrent state in FP32 and fixing the production-shaped data movement was foundational. So was treating the custom OPP package and Python host code as one ABI-versioned unit. A stale kernel can look like a model problem for a long time.

  1. Shard the model at load time instead of loading everything everywhere

The Qwen checkpoint is about 169 GiB and contains 1,610 safetensor files. It was larger than available host RAM in one of my bring-up configurations, never mind the memory on an individual NPU.

I built an expert-aware loader that reads only the experts owned by each rank, keeps dense/shared tensors where required, and avoids materializing the whole expert bank before throwing most of it away. Host-side expert data is mapped and moved lazily. This changed model loading from an accidental memory stress test into a deterministic TP4/EP4 layout.

  1. Keep low-bit weights packed and do the work on the NPU

My first usable W4 reference dequantized packed weights in the eager path and then called a regular matmul. It was useful for correctness and managed only 0.229 tok/s on one recorded smoke case.

The production path keeps the weights packed, routes tokens to experts on the device, and uses custom AscendC Cube kernels for grouped expert projections. I added FRACTAL\_NZ layouts, fused gate/up handling, tiled reductions, route and tile reuse, and dedicated W4A8 execution instead of repeatedly expanding W4 weights into a larger temporary representation.

This is both a speed win and a capacity win. Avoiding transient expanded expert banks leaves memory available for state, cache, graphs, and concurrency.

  1. Build caches for the model I actually have

Flash-Next is not a conventional all-attention transformer. My configuration has 36 recurrent GDN layers and 12 sparse QSA layers. Treating all of that as a normal dense KV cache wastes memory and misses the state semantics.

I implemented separate recurrent-state management, compact physical cache layouts, prefix-state tiers, sparse page selection, direct NZ gathers, and 310P-specific QSA paths. The service is configured for a 262,144-token context limit, and the cache planner retains capacity for about 4.08 such windows. That is a memory-planning result, not a claim that every possible four-by-262K workload has completed an end-to-end soak.

  1. Remove synchronization and launch overhead from the token loop

On these devices, a stray device-to-host scalar read can serialize the whole pipeline. I removed hot-path .item() calls, reused per-step tensors, deferred collectives behind useful work, tightened CPU affinity, and moved routing and sparse selection away from Python.

Once the eager path was correct, I added decode-only ACL graphs and MTP2 speculative decoding. I capture the actual concurrency shapes I serve rather than pretending one graph is universal. Fused operators and graph replay matter enormously when a decode step otherwise consists of many small kernel launches.

  1. Optimize the service, not just an isolated kernel

Several kernels won a microbenchmark and lost end to end. I kept the ones that reduced real request time and rejected or quarantined the others. Multi-request QSA, grouped MTP experts, expert-route caching, cache accounting, cold-prefill chunking, and collective overlap were all measured at the API boundary.

That last part is why four concurrent requests reach roughly 61 aggregate tok/s even though one request is around 30 tok/s. The extra work can occupy parts of the machine that a single token stream leaves idle.

Qwen performance today

These are milestones from different stages and workloads, not one controlled single-variable benchmark:

Qwen3.8 Flash-Next milestone / Measured result

Earliest uncontrolled service: Roughly 0.2–1.9 tok/s, often incoherent

Correct eager W4 dequant reference: 0.229 tok/s

show remaining 8,027 characters

Stable W8 service baseline: About 18–19 tok/s

Native W4A8, MTP2, graphs, one request: 29.74 tok/s median; 34.34 peak

Same optimized service, four requests: 60.88 tok/s aggregate median

Two concurrent 40K warm-prefix requests: 30.76 tok/s aggregate

40K cold prompt: About 120 seconds TTFT; 30–32 tok/s afterward

The 30/61 tok/s figures are short, fixed-output decode tests. Long reasoning requests tell a less flattering and more useful story.

https://reddit.com/link/1wvt1m4/video/rpc5c2c1s1th1/player

For quality, I ran all 198 GPQA Diamond questions on the four-chip Ascend W4 service and, as a reference, an RTX 6000 Pro running a different IQ4\_XS GGUF in llama.cpp. Both scored 140/198 (70.71%) with the same AISBench-style answer extractor. This is evidence that the Ascend path is coherent; it is not a pure hardware or quantization comparison because the runtimes, quantizations, chat templates, and concurrency differ.

On the final uninterrupted 106-case Ascend phase, four workers emitted 507,256 tokens in 2 h 52 m 46 s: 48.93 aggregate tok/s. Median per-request client rate was 12.50 tok/s on these long reasoning generations. The server completed every request in that phase with no eager fallback or zero-acceptance interval. The RTX reference was much faster per request, so I am not presenting this as an NVIDIA killer. I am presenting it as a large model working correctly and usefully on hardware that initially produced slow nonsense.

GLM is my next hard(er) model

I am also bringing up the roughly 304B-parameter GLM-5.3-Flash architecture. It combines 34 KDA linear-attention layers, 11 DSA sparse-attention layers, 288 experts, latent MLA history, mHC mixing, and a mixed W2/W4 expert checkpoint. The checkpoint is about 151.6 GiB, with approximately 35.6 GB of loaded weights per rank in one four-rank profile.

It now loads and generates on the same machine. The speed progression so far has been:

GLM four-chip milestone / One request / Four-request aggregate

Initial full-service profile / 0.764 tok/s / 1.474 tok/s

Batched NZ workspace write / 0.912 tok/s / 1.822 tok/s

NZ-packed code layout / 1.386 tok/s / 2.946 tok/s

Fused mHC/MLA work / 1.743 tok/s / 3.269 tok/s

Latest measured integrated runner / 2.236 tok/s / 4.476 tok/s

An 8,232-token GLM prompt prefills at about 39.1 prompt tok/s and then decodes at about 2.16 tok/s. Those numbers are far from my target, and the 32-token test completion was too short to establish answer quality. GLM is currently a bring-up and optimization result, not a service recommendation.

I have already found an important quantization lesson there. An early W2 checkpoint showed residual growth all the way to an RMS around 525 in the last layer. Moving the affected experts to a no-clip W4 treatment kept the network bounded. Separately, my custom blocked-dequant Cube kernel was about 17 times faster than the eager reference in its isolated test. As Qwen taught me, both numeric behavior and end-to-end integration have to pass before either result means “done.”

The third card: 50% more memory and cores, not just a spare

I plan to add a third Atlas 300I Duo to this system. That takes the machine from four to six 310P devices and from 192 GB to 288 GB of nameplate device memory. With the same ECC and runtime reservations, I expect roughly another 86 GiB of runtime-visible capacity, for approximately 258 GiB across the six ranks.

The n-card architecture is designed to use it. Expert ownership is distributed across ranks, so the two new chips add local expert capacity and accelerator cores; they are not merely passive storage. Tokens route to the ranks that own their experts, and the additional ranks participate in the model's compute and collectives. I expect useful scale from EP6, although the exact speedup will be measured rather than advertised in advance.

The extra capacity gives me several options:

• Keep larger expert sets or higher-precision layers resident.

• Fit models that are just over the four-chip limit without host offload.

• Spend more memory on long-context state and cache.

• Reduce aggressive quantization where the quality trade is not worthwhile.

• Run a large distributed model while retaining room for another smaller service or evaluation workload.

The cost is one more card, 150 W of maximum board power, two more device ranks, and another passive heatsink that needs real airflow. In the cooled room, none of those are architectural concerns. The interesting cost is communication: six ranks change route balance, collective sizes, and PCIe/HCCL traffic. I will tune and benchmark that topology, but the software is already organized around n-card expert distribution rather than hard-coded four-way ownership.

Known issues and active work

This is what is still on my bench:

• I just fixed one real cross-stream race: a prefix-Mamba state slot could be spilled or reused before its pending NPU writer completed, allowing an older checkpoint to be restored. That fix has an NPU regression test.. A later 106-case run survived 12 state spills without degrading, which is encouraging but not a root-cause proof. I am continuing long mixed-load and eviction/reuse soaks.

• Qwen cold prefill. A 40K cold prompt still takes roughly two minutes even though subsequent decode is fast. Profiling points primarily at the 12 QSA layers, especially sparse selection and tiled attention. This is now a more important target than another tiny decode micro-optimization.

• Qwen long-context qualification. The planner has the capacity, but I am separating configured context, allocated capacity, and completed end-to-end long-context tests. I want retrieval and concurrent-fill evidence, not a screenshot of a launch flag.

• GLM coherence and performance. I am requalifying the full model after KDA, MLA, QSA, and runner integration changes, then moving the grouped mixed W2/W4 expert path, cache layout, graph replay, and eventually MTP through the same correctness-first gates used for Qwen.

• GLM loading and memory. The filtered loader can skip large amounts of peer-owned or superseded checkpoint payload before tensor materialization. I saw one roughly 20% load-time improvement, but it needs controlled reruns and byte-accounting before I call it a result.

• Six-device expert parallelism. When the third card arrives, I will measure rank balance, per-card temperatures, collective time, model capacity, and c1/cN throughput on the exact six-rank topology.

• Other model adapters. The same loader, packed-expert, cache, and operator infrastructure is feeding ongoing DeepSeek and other hybrid/MoE work. I am avoiding model-name conditionals where the underlying contract can be made generic.

Should you buy one?

I plan to list my first additional card in the OpenSensor storefront (https://www.opensensor.io/) next week at $2,900. The listing is not live yet. I believe it is a good price for what the hardware can already do and what the software should unlock. It is not yet the same kind of turnkey purchase as a supported NVIDIA card running mainline vLLM.

I think the right buyer is a developer, lab, or systems-minded end user who is comfortable with both of the following:

  1. You own the airflow solution. These are passively cooled server cards. Depending on the chassis and motherboard, that may mean high-static-pressure case fans, a duct, or a 3D-printed shroud. I use Fusion 360 and am happy to help with additional shroud designs. There are too many motherboard layouts, card spacings, fan sizes, and case geometries to pretend that one printable design will fit everything.
  2. You are adopting an active development fork. I am the sole developer on the vLLM work today as upstream is focused more on their server grade accelerator modules not yet available to the US. The progress is real and the benchmarks in this post are from actual hardware, but more bugs and better approaches will be discovered. Buyers shoul
💬 118 (+6) open on reddit ↗
▲
136
+6
20👁
r/LocalLLaMA · u/yahbluez · 10d ago
Is AI Profitable Yet?
▲
134
-2
18👁
r/LocalLLaMA · u/Automatic-Arm8153 · 12d ago
Mimo v2.6 flash MOPD
▲
129
+19
69👁
r/LocalLLaMA · u/aya-ifm · 8d ago
AMA about K2 Horizon, Meet our team from IFM

Hi r/LocalLLaMA

We’re researchers at the Institute of Foundation Models (IFM), an AI research lab dedicated to open and independent development of frontier-class foundation models.

We recently released K2 Horizon a connected fleet of six fully open models with size ranging from 0.9B to 375B. In addition to weights, we also open-sourced training data and recipes, training code, intermediate checkpoints, fine-grained training logs and evals.

Ask us anything about pre-training and data mixes, post-training, small models on-device, MoVA and sparse attention, deployment, what’s out now, and what’s coming next.

Participating in the AMA:

  • Hector Liu u/hunterhector
  • Alexander Moreno u/IFMAlex
  • Mikhail Yurochkin u/my-moonfolk
  • Rupesh Srivastava u/k2pt
  • Junlin Chen u/Junlin_Chen110
  • Haonan Li u/East-Career9147

We'll be live Mon, Oct 5, 8–10 PM PT. Questions are open now, so drop yours anytime!

Join IFM on Discord: https://ifm.ai/discord

https://preview.redd.it/awu0r6fbowsh1.png?width=3240&format=png&auto=…

💬 135 (+108) open on reddit ↗
▲
126
+11
83👁
r/LocalLLaMA · u/FutureStriking283 · 8d ago
Anyone wonder why americans lag so far behind in the open LLM market?

I mean , DeepSeek, Kimi, GLM, MiniMax -- the list of Chinese LLM's is such a long freaking list. As American's -- why don't we feel .. a little funny .. about being so far behind? Chinese entrepreneurs are peneuring like crazy and American's .. just obsess on .. what ?

update -- already I'm starting to see some clear answers. American's are putting money ahead of technology. Completely understandable.

second update -- I asked "why can't we have american AI as good as or better than the chinese" and BY FAR the number one most supported comment? "Have you even thought of shareholder value!?" . We .. America , are so fucked.

💬 425 (+28) open on reddit ↗
▲
123
+7
24👁
r/LocalLLaMA · u/autoencoder · 10d ago
Another case of censorship post image

I was optimizing my diet using a cloud AI provider, and I found my Pi agent stuck like this. Looks like I had too many mushrooms lol.

💬 24 (+1) open on reddit ↗
▲
119
+1
33👁
r/LocalLLaMA · u/Usual_Maximum7673 · 8d ago
Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory

A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO 6000, but as a general zero-shot classifier it trailed the big models.

Then it occurred to me that most decisions an agent makes in front of a local model aren't open-ended. They fall into a handful of recurring kinds: is this a prompt injection, which tool to call, how urgent is this ticket, is this answer grounded in the sources. So I trained 9 LoRA adapters, one per job, and you pick the ones you need. The server loads the base once plus whichever adapters you choose (about 40 MB each), and every request either names an adapter or goes to plain Jeff.

That means you keep both: the base model stays untouched, so you still get Jeff's general zero-shot ability for anything new, and the adapters give you near-perfect accuracy in the domains you care about. Each adapter was also trained with 10% of the base model's own training data mixed in, to help it keep its general skills.

Everything is on jeffhub.ai: the adapters, the results, the docs. Code on GitHub, models on Hugging Face, and you can try all nine adapters in your browser.

The headline: I let Jeff + adapters answer first and pass only the queries it's unsure about to Qwen3.8-27B. Same test rows both ways, on an M4 Max:

|Measure|Qwen3.8-27B alone|Jeff + adapters, 27B only when unsure|
|:-|:-|:-|
|Accuracy (mean of 8 adapters\*)|86.6%|95.3%|
|Time per decision (mean)|8.1 s|0.25 s (38× faster)|
|Wrong answers|13.4%|4.7%|
|Memory|28.6 GB|under 2 GB for Jeff, even with all 9 adapters loaded (+6.9%)|

On the five decisions an inbox agent makes for every message (guard, triage, support intent, tool choice, grounding) alone: 87.7% → 95.7%, 39× faster. Jeff wins outright on 8 of the nine adapters and ties on grounding (96.3% vs 96.7%, at 20× the speed). On their full held-out test sets, six of the nine adapters score 97–98%. On a GPU, a decision takes about 30 ms, whether you load one adapter or all nine.

\*Emotion is left out of the averages: picking the single strongest of 27 emotions (or neutral) in short Reddit comments is hard even for people, and the human labels often disagree. Jeff + adapter scores 60.6% there against the 27B's 35.6%, at 42× the speed. Including it, the average across all nine adapters is 91.4% for Jeff + adapters against 80.9% for the 27B, so leaving it out makes the gain shown above smaller, not larger.

Caveats, up front:

  • the 27B ran in 8-bit with step-by-step reasoning off (with reasoning on, the speedup would be even more dramatic);
  • each task used a fixed sample of 300 held-out rows (500 for emotion and legal-clauses);
  • each adapter's "pass it on" threshold was chosen on separate calibration rows, before the test rows were scored.

Data: 4 adapters are trained on public data sets. 5 are mostly synthetic. Every generated row records which model wrote it, and the cards give the counts. Every data set went through a shortcut check and an independent review before training, and a lot of first drafts failed: things like the answer being given away by length.

What's open: weights (Apache 2.0), code (MIT), and each adapter's test and calibration sets, so you can check every number. The training data isn't published.

This is a community preview: I'd love feedback.

Next: over the next \~36 hours I'll train v1.3, a long-term-support base. The fixed parts of a prompt come first, so servers can prepare them once and reuse them, which means faster decisions. I'll then retrain all nine adapters on it and keep the request format stable, so others can build and submit their own adapters. The adapter kit, with the data checks I used, is in the repo.

I've got access to more hardware now, so if there's a decision you'd like an adapter for, tell me and I'll train it.

The goal: when the next generation of local models lands (like everyone, I'm watching for Qwen 4), anyone running one locally should also have a tiny, fast, well-calibrated decision layer in front of it.

💬 33 (+1) open on reddit ↗
▲
110
 
40👁
r/LocalLLaMA · u/rmonsurate · 9d ago
Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)

We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next.

Victoria (coding and agents)

  • Qwen3.8-Flash-Next cut down by 44% using a paper / technique called REAP: 512 down to 288 per layer.
  • Retrained at 4-bit (NVFP4) afterwards, so it's trained for the format it ships in rather than just quantized after the fact.
  • Terminal-Bench 2.1: 70.0%, averaged over 3 runs with an 8h per-task timeout. Our previous NVFP4 build scored 62.5%.
  • HumanEval: 159/164.
  • 48.0 GiB of weights, including the draft head. The 95.4 GiB n-gram table is separate and not counted in that number.
  • 280 tok/s single stream on one B300 with the draft head, versus 135 without it.
  • GGUF Q4_K_M is 49.17 GiB. It scored 75.3% on Terminal-Bench (a single run, so treat it as noisy) and 93.2% on HumanEval (averaged over 5 runs).
  • Uses 35% fewer output tokens than our previous build.

Maple (Canadian questions)

Most models answer questions about taxes, benefits and regulations as if you live in the US. Maple is fine-tuned to default to Canada. On 600 held-out questions, with search:

  • Cites an official Canadian source: 6.0% before fine-tuning, 62.9% after.
  • Fully correct answers: 6.6% before, 21.8% after.
  • "No answer" responses: 47.2% before, 23.7% after.
  • It pushes Canada onto people who said they live somewhere else less often: 2.9% before, 1.0% after.

Coding holds up: 157/164 on HumanEval. Grading was done by an AI judge panel; human review hasn't happened yet.

Links:
https://huggingface.co/rmonsurate/Victoria
https://huggingface.co/rmonsurate/Maple

Happy to answer questions about running them.

Edit: llama.cpp users. The GGUF carries our draft head, and mainline llama.cpp doesn't know about it yet, so it fails with "expected 1256, got 1224". Your download is fine. For now, build from our fork: github.com/rmonsurate/llama.cpp, branch qwen4exp-mtp. Prebuilt binaries are on the way. Thanks to the reader who caught this.

💬 42 (+3) open on reddit ↗
▲
105
+4
32👁
r/LocalLLaMA · u/Loginhe · 10d ago
[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw post image

We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed.

Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16.

What's inside

  • Four quantized GGUFs, 2.40 to 3.50 bpw (66.4 to 83.6 GB), and the BF16 vision projector
  • Expert-pruned Coder GGUF, 58.4 GB in total, of which 29.6 GB must remain resident
  • GSQ (Gumbel-Softmax Quantization): post-training scalar quantization that jointly learns grid assignments and group scales, closing most of the gap between scalar and vector quantization at low bit-widths while remaining deployable in standard GGUF types
  • RCO (Riemannian Constrained Optimization): enforces exact budgets by gradient descent on the task loss, without per-constraint tuning. It serves two roles in this release: assigning a quantization type to every tensor, and selecting which experts to retain in the Coder build, where it enforces several exact budgets simultaneously, one per layer

Results:

At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.

  • IQ3\_S (3.50 bpw, 83.6 GB): AIME25 100.00, GPQA-Diamond 92.93 against 91.92 for BF16, LiveCodeBench v6 86.86 against 87.43. Task average 93.26 against 93.12.
  • IQ3\_XXS (3.00 bpw, 75.8 GB): AIME25 100.00, GPQA-Diamond 91.41, LiveCodeBench v6 86.29
  • Q2\_0 (2.40 bpw, 66.4 GB): zero-shot average 78.00, above the BF16 value of 76.94, at approximately one fifth of the size

Coder (capability pruned model):

Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits.

The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.

  • SWE-bench Verified: 75.60 against 82.80 for BF16, retaining 91.3%
  • LiveCodeBench v6: 86.28 against 87.43, retaining 98.7%

Both measured at xhigh reasoning effort.

Links

Both repositories ship the complete per-tensor RCO allocation.

The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome.

From the ISTA Deep Algorithms and Systems Lab.

💬 77 (+5) open on reddit ↗
▲
105
+52
74👁
r/LocalLLaMA · u/basnijholt · 7d ago
Self-hosting AI does not save money, and I do it anyway

Hi folks, I'm a long-time lurker and big fan of this subreddit and a massive self-hosting fan (also outside of AI).

I doubt many people will disagree with me here because I see the same arguments being made in many posts. However, I thought it might be interesting to share anyway. I wrote down why self-hosting AI does not save money: https://www.nijho.lt/post/self-hosting-ai-is-not-cheaper/

EDIT: didn't think this would be so controversial 😅 I do say explicitly in my blog post "I would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not".

EDIT 2: Comparing $200 sub with Opus 5.5 or Astra with Qwen 3.8 27B is not apples to apples.

💬 194 (+42) open on reddit ↗
▲
104
+8
30👁
r/LocalLLaMA · u/tossit97531 · 11d ago
Can we get some quality control on all these model perf posts?

Too many hyperactive amateurs are coming in here with "1b model at 832843tok/s!" and hardly any of them have all the info necessary for local runners to evaluate. We need context ladders with perplexity/KLD, hardware specs, model params and quant(s), runtime, tuned runtime parameters, basically everything we need to reproduce locally if we can match the entire setup. To say nothing of what the model is even good at in the first place if it's not a well-known model.

The goal is to get perf numbers that show they meet a certain quality bar. I don't care if I get 8324834 tok/s if it's all garbage.

Can we start filtering the hyperactive amateur perf posts please? It's getting really frustrating seeing all these posts of models and wading through info just to see that it doesn't test with anything but an empty context or doesn't say anything about quant or platform.

We need to define some rigor and apply it to this place, or it will remain like most ai-oriented subs and get continually choked with slop.

▲
97
 
49👁
r/LocalLLaMA · u/wombweed · 11d ago
I am concerned about all these disparate hard forks that target specific architectures instead of opening a PR against upstream

Other than the obvious self promotion, is there a practical reason people do this that I am missing? There's dozens of llamacpp forks with silly names that are supposedly "optimized" for this or that specific GPU and seem to have zero intention to merge into upstream. Am I missing the real reasons why this happens so often? Why do people think it's OK to do this? In my experience in the open source community this is generally frowned upon.

I don't know if it's just a me problem that this kind of thing puts me off so much. I am usually quite grateful for PR feedback and conscientious about the code I put out there; I take pride in submitting high quality code that meets or exceeds the standards of a given project. Of course there is nothing ethically wrong with hard forks or taking shortcuts if you find the collaborative process cumbersome, but personally I wouldn't promote my fork in such cases, let alone go out of my way to add custom branding with a Reddit announcement post etc. since the effort required to do so seems roughly equivalent to the effort required to meet the contributor standards. In contrast, many of the authors of these forks seem very eager to have others adopt their rebranded fork for production use cases. There just seems to be a big disconnect, idk.

Edit: some great discussion in this thread, thanks to all who responded. Consensus seems to be that (excluding the obvious low-effort engagement bait forks) the base project has to meet many compatibility requirements while a downstream project can be more focused, which is a great point.

💬 166 (+3) open on reddit ↗
▲
96
+4
28👁
r/LocalLLaMA · u/JLeonsarmiento · 12d ago
... so, yeah. post image

Finally got 3.8-Flash-Next running on my M4Pro 48GB Mac with https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

Dense 3.8-27B is just faster... and maybe better due to quantization level...

EDIT:

Hold a second, Flash-Next is actually performing faster than 27B after some key flags on llama.cpp. it's Holding up to 131K without OOM-ing..... maybe...

0.36.940.283 I srv          load:   --top-k

0.36.940.283 I srv          load:   20

0.36.940.284 I srv          load:   --ctx-size

0.36.940.284 I srv          load:   131072

0.36.940.284 I srv          load:   --cache-type-k

0.36.940.285 I srv          load:   q4\_0

0.36.940.285 I srv          load:   --cache-type-v

0.36.940.285 I srv          load:   q4\_0

0.36.940.285 I srv          load:   --flash-attn

0.36.940.285 I srv          load:   on

0.36.940.285 I srv          load:   --load-mode

0.36.940.286 I srv          load:   mmap

0.36.940.286 I srv          load:   --lazy-mode

0.36.940.286 I srv          load:   on

EDIT 2

Yes, this model is brutal. This quant at Q2\_0 in llama.cpp is out performing 27B at oQ4e in prompt processing, speed generation, but most importantly, the only thing that matters, sheer intelligence.

What a time to have 48 GB of ram !!!

▲
92
-1
24👁
r/LocalLLaMA · u/pneuny · 12d ago
Qwen company already rushed out a Jev competitor. No open weights yet.

EDIT: About that, I tried it out, and it's garbage so far. I did some basic tests through AIHubMix (do not use that platform btw, it's trash), and my agent did some comparison. I guess it figures, it was a model they released just days after the Jev hype started. The agent's analysis is below:

AI Agent's output:
```
I thoroughly tested https://aihubmix.com/v1/systemone using the provided API key and decision-model-preview across latency, throughput, and
linguistic judgment accuracy against our test suite.

Here are the test results and why I strongly recommend NOT switching to this endpoint yet:

────────────────────────────────────────────────────────────────────────────────

  1. Latency & Rate Limit Benchmark

- Server-side execution: The endpoint reports latency_ms: ~130ms–160ms.
- Total Round-Trip (network + TLS): Averaged 644.9 ms (ranging from 473ms up to 882ms). By comparison, your existing local router
([9router IP]) averages ~439 ms.
- Hard 16-Question Ceiling:
The proxy strictly rejects requests with more than 16 questions:
{"error": {"message": "questions: 19 exceeds the limit of 16", "type": "Aihubmix_api_error"}}
On longer Japanese sentences (e.g. 外に出してやってくれませんか。 or ちょっと聞いてみたいんだけど。), our parallel diagnostic tensor sends
19–22 questions, which throws an immediate 400 Bad Request.
- Aggressive Rate Limiting: Even with a 1-second pause between sequential requests, it frequently triggered 429 Too Many Requests.

────────────────────────────────────────────────────────────────────────────────

  1. Quality of Judgments (Major Semantic Degradation)

To test quality, I adapted our test battery into a compact 10-question payload to stay under the 16-question limit. Across the benchmark,
decision-model-preview exhibited severe calibration collapse:

Test Case 1: Indefinite Pronoun vs. Wh-word

  • Japanese: 何か待ってるの? ("Are you waiting for something?")
  • User Draft: "what are you waiting for" (Clear grammatical error: confuses yes/no 何か with open 何)
  • Local Router (oc/jev-1.13-free):
  • is_flawless: 0.07 (correctly marked false)
  • grade_bracket: 5_moderate_error (confidence 0.98)
  • sentence_critique_summary: confused_indefinite_with_wh_word (confidence 0.96)
  • Result: Accurately docks to 5–6/10 and explains the question-type mismatch.
  • Aihubmix (decision-model-preview):
  • is_flawless: 0.98 (marked true!)
  • grade_bracket: 10_flawless
  • sentence_critique_summary: no_flaws_accurate
  • Result: Completely failed to detect the grammatical error and awarded a false 10/10.

Test Case 2: Benefactive Direction Reversal

  • Japanese: 外に出してやってくれませんか。 ("Could you let him/the pet out?")
  • User Draft: "would you let me outside?" (Critical reversal: favor for someone else vs. oneself)
  • Local Router:
  • grade_bracket: 3_major_error (4/10)
  • benefactive_direction: recipient_reversed_self_vs_other
  • Aihubmix:
  • is_flawless: 0.83
  • grade_bracket: 10_flawless
  • sentence_critique_summary: no_flaws_accurate

────────────────────────────────────────────────────────────────────────────────

Summary & Recommendation

1. Model Calibration is Inferior: Whatever model or snapshot is currently mapped to decision-model-preview on Aihubmix appears to be an
uncalibrated or older preview checkpoint. It fails on the exact Japanese linguistic subtleties (benefactives, wh-scope, modality) that we
carefully tuned.
2. Artificial Constraints: The hard 16-question limit and 429 rate limiter break our dynamic sentence-chunking architecture.
3. Current Setup is Superior: Your local router (http://[9router IP]/v1 running oc/jev-1.13-free) is ~200ms faster, has no question caps,
never rate-limits, and correctly diagnoses complex Japanese grammar.

Recommendation: Keep the current endpoint (http://[9router IP]/v1) active. If you still want the script modified to allow switching
providers via settings or want to test it anyway, let me know and I can make the Jev endpoint independently configurable in the UI settings
dialog.
```

Original Post:
----

It's called decision-model-preview. There is only a docs page. No announcement or anything. I can't post a link because reddit's filters just deletes posts that contain a link to the cloud platform that hosts it. But you'll find the page if you Google the model name.

▲
89
+6
44👁
r/LocalLLaMA · u/pubudeux · 11d ago
First few days of qwen3.8-flash-next on 4x R9700 - it's been really interesting so far post image

Here's a metric dashboard giving an idea of the last few days.

Been testing with a variety of different agentic coding use-cases, mostly using a pi harness.

qwen3.8-flash-next has seriously exceeded my expectations (used https://huggingface.co/tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8)

Both speed and quality have surprised me, given that I can get 3-5 concurrent streams going with \~100t/s gen each, and single stream easily gets to 150+t/s. Prefill is 10k+t/s

💬 67 (+2) open on reddit ↗
▲
84
+5
41👁
r/LocalLLaMA · u/WebAssemblyMan · 8d ago
DeepSeek harness 0.2 - Optional Bundle Architecture, Windows Sandbox improvements, Async Question Mode, Desktop release, Web Search without key post image

Optional Bundle architecture: Schedule (session-local delayed / timed / interval reminders) was removed from the default set and made an explicit Optional Bundle. This cleanly separates “installed” from “enabled” and is the first systematic use of the Profile + Bundle model for official features.

• Windows Sandbox improvements: A new permission-diagnosis skill can detect common Access Denied causes and perform backed-up, recoverable permission fixes after user authorization, giving the Agent a reliable recovery path instead of blind retries.

• Async Question Mode (experimental): “Ask the user” is no longer a hard synchronous block. After a timeout the Agent can keep working while the user answers later, introducing asynchrony between interaction and execution.

• Model-layer polish: DeepSeek-account sessions can use Web Search without an extra API key; third-party model catalog updated (some old IDs removed); long model lists now support fuzzy search and keyboard navigation.

• Desktop release: Official Windows and macOS clients are out (Linux unsupported). Account login is supported, suggesting paid plans may be coming soon.

• Overall theme: Version 0.2 strengthens the Agent Runtime’s composability, recoverability, permission boundaries, and execution-state semantics — the practical foundations needed to move from a toy toward production use.

💬 17 (+3) open on reddit ↗
▲
82
+7
34👁
r/LocalLLaMA · u/Skyline34rGt · 9d ago
BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B

"AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.

AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.

  • Architecture: Dense Qwen3.8-compatible multimodal model
  • Parameters: 27B
  • Context length: 262,144 tokens

[](https://huggingface.co/BAAI/AREX-2#key-features)Key features

  • Long-horizon self-improvement: turns extra test-time rounds into useful solution refinement.
  • Feedback-driven reflection: reads scores, logs, errors, and timings to decide what to change next.
  • Cross-domain performance: training on coding and machine-learning tasks also improves the model's deep-research performance.
  • Long-horizon reasoning: sustains productive iteration as the task budget grows."

Gguf's - https://huggingface.co/mradermacher/AREX-2-GGUF

💬 29 (+2) open on reddit ↗
▲
79
-1
40👁
▲
79
+8
30👁
r/LocalLLaMA · u/Ok-Breadfruit-3523 · 8d ago
Update on the “Monstrosity”. 6 BC-250 board cluster post image

This is 6 bc-250 ex mining boards with 5 in the asrock 4u12g case they came in. After a lot of testing my current preferred setup is 4 boards running Qwen Next Flash IQ2\_XS at 100k context with around 28 tok/s for short generation and 24 tok/s at 50k with around 115 ppt. The other two boards run 3.6 35b q4 at 60 tok/s with 100k context and 450 ppt. This is all using llama with vulkan and rpc over 1gb Ethernet.If anyone has any suggestions with this beast I am all ears. I had these boards left after reselling a bunch and had never done anything with local ai before so this has been a blast. Also yes that is a cardboard box with 3 fans on top as the intake.

💬 42 (+5) open on reddit ↗
▲
73
+5
40👁
r/LocalLLaMA · u/TypicalPudding6190 · 12d ago
Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM post image

We built an inference engine InferredThoughts for MoE models that don't fit in VRAM + RAM. Most of the model stays on the SSD, and experts are read as tokens need them.

This started as a proof of concept, and poc worked, we are getting 9-10 tok/s decode on Qwen3.8-Flash-Next NVFP4 (9.06 on the benchmark turn, 10.4 on the best turn).

This is just the start. With better SSD streaming, we expect v2 to reach ~14-15 tok/s decode.

Repo: InferredThoughts
https://github.com/compiledthoughts/Inferred-Thoughts

Model: Qwen3.8-Flash-Next, 176.9B params, NVFP4 GGUF (119 GiB):
https://huggingface.co/CompiledThoughts/Qwen3.8-Flash-Next-NVFP4-Q8_0

Machine: RTX 5060 Ti 16 GB, Ryzen 7 9700X, 32 GB DDR5, Gen5 NVMe SSD 1Tb, Windows 11

Where the 119 GiB lives

| part of the file | size | where |
|---|---:|---|
| dense weights (attention, shared experts, LM head) | 4.4 GiB | VRAM |
| token embedding table | 0.6 GiB | RAM, one row read per token |
| hottest routed experts | 8.8 GiB | VRAM |
| next-hottest routed experts | 6.0 GiB | pinned RAM |
| remaining routed experts | 48.5 GiB | SSD, streamed on demand |
| n-gram table | 50.7 GiB | SSD, 16 rows read per token |

So 20 GiB is in memory and 99 GiB stays on the SSD: 48.5 GiB of routed experts, streamed as the router picks them, and the 50.7 GiB hashed n-gram table (looking forward to qwen4 ngram).

Speed

  • Decode: 9.06 tok/s on the benchmark turn, 10.4 on the best turn at xhigh effort
  • Prefill: 49.2 tok/s on a 5.5k-token prompt
  • llama.cpp on the same machine: 4.9 tok/s average decode

How it works

  • VRAM holds the dense weights and the hottest experts (GCLOCK eviction), pinned RAM the next tier (read over PCIe), and the rest come off the NVMe on 8 read threads.
  • Lookahead prefetch guesses the next layer's experts and starts their reads early.
  • NVFP4 matmuls run FP4 x FP4 on the tensor cores, with no unpacking first.
  • About 270 MiB is read from the SSD per token, and roughly 75% of expert lookups hit memory.
  • Each token uses 480 experts (10 in each of 48 layers). About 377 of them are already in VRAM or RAM; the other ~103 are read from the SSD, about 270 MiB per token. This hit hit-rate is what allowed us to reach 9tps.
  • It only reads from the SSD and almost no writes so ssd should have minimal wear due to writes. But we saw SSD hit 70C during long runs.

Also supported: Qwen3.6-35B-A3B NVFP4. It fits in VRAM + RAM . You can also run it via ssd streaming and it was the learning curve for this work. On the same machine with tuned config it hits: 47.3 tok/s decode at ~4k context, 591 tok/s prefill.

Serving: an OpenAI-compatible server that renders the model's own chat template, with tool calls (still buggy and tested with Cline), the reasoning split out from the answer, and a built-in chat page.

Limits of v1: RTX 50-series / Blackwell (sm_120) only, tested on Windows 11 and WSL2 only, greedy decoding only.

Links

Questions and feedback welcome, especially from anyone running big MoEs on small cards.

▲
72
 
47👁
r/LocalLLaMA · u/Mrinohk · 12d ago
Don't trust frontier models when asking about budget hardware!

Early this year when I was first looking at building up my inference capability you could get the 16GB Tesla P100s for between $60 and $80. Asked claude about it, told me absolutely not worth it. No tensor cores, bad int4/int8, no BF16, not worth it. Needs special power accommodations, Above 4G decoding option in the bios (it made it out like it was some rare option), and a semi-exotic cooling solution.

Optimized the shit out of my RX6600XT in llama.cpp as a result. Got pretty far.

Decided to say fuck it, bought a single P100 last week, finally showed up day before yesterday. Got a newer power supply with the appropriate connections (not hard, not that expensive, seen options as cheap as $60 from good brands, I spent $100 on one with some headroom), multiple llama.cpp forks and patches that carry some wild optimizations to handle the capability gap, and a 3D printed housing for a 94mm fan from noctua. Doesn't generate enough static pressure to keep it cool during prefill, but more than strong enough for the generation step.

The numbers I was getting before, with my RX6600 XT with Qwen3.6 35B A3B UD\_Q4\_K\_XL with MTP and --cpu-moe:

PP \~800 at 0 ctx, drops to \~700 by 10k

TG \~30-35 prose, 45-50 code.

This setup could do 64k context (and possibly higher) at 16bit kv. cpu-moe helps a ton in that respect.

With just a little bit of tuning, and using specifically the patches from shinbunbun for llama.cpp, same model with the same MTP settings, --n-cpu-moe 22:

PP \~600 at 0 ctx, 440-500 by 10k

TG \~54-60 prose, 66-72 code.

Running only 32k context right now to make it work. Could fit more with a higher n-cpu-moe, but my harness doesn't need that much (rarely see it over 30k, persistent memory leads to chats that simply aren't meant to last).

I know the capability gap between 3.6 35B and the basically any of the qwen 3.X 27B models is pretty big, but this is huge for the price. They've gone up since I bought mine, about \~$15 across the board. Still something you can get for under $100 and makes for inference that is simply impossible to get at that price otherwise.

I've got another one coming so I can go full offload on the model, and maybe even start playing with 3.8 27b. Right now IQ3\_K\_XL I get around 9 tokens per second with MTP, and basically no real context. Don't actually know if splitting a model that fits in one card across multiple helps speed, that's completely new territory for me, but I'm having fun regardless.

Card is seriously underrated. It's a great, (relatively) inexpensive way to get capable compute to finally start doing local AI stuff. I went from having to just leave my computer alone while the model was running and do everything from my macbook (good bye gaming) to being able to let the model live and work in the background while I'm doing basically anything on my PC. When the second arrives, I'll be planning my dedicated inference box they'll both live in. Feeling inspired by that guy cooling his PC with a VW radiator.

💬 54 (+3) open on reddit ↗
▲
72
+6
32👁
r/LocalLLaMA · u/MasterNomie · 11d ago
What model sits between Qwen 3.8 27b and Flash next for coding?

Having tested both Qwen 3.8 27b and Flash next on RTX 5090 with 96GB RAM, I want to find the middle ground between the two for coding capabilities but not sacrifice decode speed to standstill. I would like decode speed to between 75-100 ideally for fast iterations; otherwise I become impatient.

Currently I get 200+ TPS on Qwen 3.8 27B and approx 50 TPS on Flash next.

My hardware - RTX 5090 and 96 GB DDR5 which I plan to upgrade to 128 GB (in this economy, yes, but unwillingly).

What model sits between these two in terms of coding capabilities and hardware requirement? If none is present, I can perhaps think of using Flash next for plan creation and 27b for implementation.

Edit: Fast forward few days. I gave Strata a go with Swift 1.5 Flash Next IQ3\_XXS and able to achieve \~150 tok/sec decode and 5k tok/sec prefill. I am escatic! The quality of response from 27b is considerably better and the speed is great. Both targets achieved.

▲
72
+22
45👁
r/LocalLLaMA · u/lbgos_Loss783 · 7d ago
I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking

Hey local AI community, I've been working on this for a while and finally feel ok sharing it.

It's a cyber benchmark where the model gets a shell in an isolated docker box and has to find the exact flag. Pwn, web, crypto, rev, forensics, a few real CVEs and some multi-stage ranges. 19 tasks, 6 models, 544 scored attempts.

To be clear, I didn't build every task by hand. GLM 5.3 helped me create several of them. For an open model its cyber capability is really high, and it barely refuses anything, so it was one of the best options for this. GLM 5.3 isn't one of the benchmarked models.

The local part: I ran Qwen3.8 27B (Unsloth Q4\_K\_XL, xhigh) on a llama.cpp RPC pool across a 3090 and a 3080 in two Proxmox nodes, connected over a direct 2.5G link. That gave me enough concurrent tps to run several agents at once. I also started a low reasoning run, but it was taking 20+ hours because of a harness problem, so I killed it.

Why I'm posting now: John Hammond put out a video about how threat actors use AI (https://www.youtube.com/watch?v=xHDc6-7bjyw). One part is a guide from a criminal forum on running abliterated models on RunPod, and one of the models in it is Qwen3.8 27B. I had benchmark data on that exact model, so here's what it can actually do.

Stock Qwen3.8 27B got 28.1% on the first try and 0% on pwn. Not bad for a 27B on two gaming cards, but not much of a threat on its own either.

The cheap API models are a different story:

\- MiMo 2.6 Flash solved 73.7% on the first try

\- GPT-6 Luna solved 90.9% within 3 tries

\- On multi-stage ranges, where you chain several steps, the top models got 92-96%

Pwn is still hard for everyone (best was 56%), and 3 tasks haven't been solved by any model in 82 attempts.

Results: https://lbgos.dev/bench

Harness (MIT): https://github.com/lbgos/rangebench-harness

The tasks aren't public so they don't leak into training data, but you can still run them. DM me here or on X (lbgosna), and I'll send them over. You run it on your hardware or tokens and I'll add your results to the board. If a few people send local runs, I'll make a separate local-only table.

This is my first time building something like this, so any feedback on methodology, task mix or what's missing is welcome.

💬 40 (+9) open on reddit ↗
▲
71
+9
38👁
r/LocalLLaMA · u/returnity · 11d ago
Searching for 3.8 35B: Qwen3.6-35B-A3B (Testing 5 Finetunes vs. Base)

TL;DR -- You should probably just use base Qwen3.6-35B, as only Occamy-1.0 is competitive with it. Tiel is a major let-down, worse than Ornith. KAT surprises (good), Nex surprises (bad). This post is long. Sorry, lots to cover.

I think we all want to see a next-generation small MoE from the Qwen team to replace 3.6-35B in our workflows. This model is a perfect fit for smaller gmaing laptops and mid-tier rigs. It sucks that Qwen seems to have abandoned this model, but at least there are fine-tunes that improve upon it... right?

Well... maybe not. I ran benchmarks on the 3.6-35B-A3B base model, as well as five finetuunes: Occamy-1.0, Ornith-1.5, KAT-Coder-V2.5-Dev, Tiel-Coder, and Nex-N2.5-mini, and the results are quite surprising.

I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark I can run in ~10 hourso on my machine, it's not actively post-trained on by any of these models, and it provides a lot of useful information along with the raw accuracy scores. This includes: first-try and retry pass rates, token counts, solve times, and how well-formed the output diffs are. Here's the table:

| model | First-try pass | Retry pass | tokens | sec/case | tok/solve | well-formed diff |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B BASE (STOCK template) | 37.4% | 71.0% | 8650 | 285 | 14.1K | 96.3% |
| Occamy-1.0-35B-A3B (STOCK template) | 29.0% | 70.1% | 6801 | 285 | 17.2K | 86.9% |
| Occamy-1.0-35B-A3B (froggeric medium) | 30.8% | 69.2% | 6009 | 233 | 16.8K | 94.4% |
| Occamy-1.0-35B-A3B (froggeric, xhigh) | 27.1% | 67.3% | 8631 | 310 | 20.0K | 91.6% |
| Ornith-1.5-35B-A3B | 23.4% | 63.6% | 4813 | 226 | 16.4K | 87.9% |
| KAT-Coder-V2.5-Dev | 20.6% | 58.9% | 2190 | 84 | 9.3K | 86.9% |
| Tiel-Coder-35B-A3B | 18.7% | 53.3% | 4851 | 171 | 18.2K | 89.7% |
| Nex-N2.5-mini | 10.3% | 30.8% | 5037 | 188 | 33.3K | 95.3% |

As you can see, the only finetune that even competes with the base model is Occamy-1.0. The rest are utterly dominated by the base model, a grim disappointment for finetune enthusiasts. I was particualarly surprised by the performance of Tiel, which seems to get a lot of love in this subreddit.

Speaking of Tiel, I want to clarify that Tiel is just Ornith-1.5 with a different chat template, Sharp, which is based on froggeric with an added "terse mode" instruction that's supposed to reduce excessive verbosity. I wanted to standardize for templates, so ALL models are using the base froggeric v22.5 template set to medium (which is equivalent to standard thinking on, no additional message sent). I used this because I wanted to test Tiel vs. Ornith-1.5, and Tiel is the chat template. Also, practially, I use froggeric in my real workflows. However, to ensure coverage, I also tested the STOCK template on the 2 highest-performing models, to make sure it wasn't affecting the scores. As you can see, the template doesn't make a significant difference in the scores, and the scores for base 35B with different templates are so close to identical that I excluded the froggeric one from the table.

I also tested froggeric/Sharp's reasoning-effort toggle, and found xhigh -> medium significantly reduces token counts and solve times (by ~1/3), without affecting accuracy significantly. That stands in stark contrast to Tiel's 'terse mode' toggle, the core feature of Tiel over Ornith, which dramatically reduces accuracy along with the reduction in token counts. My results strongly suggest that if you want a less verbose model, you're better off lowering the reasoning effort than using Tiel with terseness on.

Speaking of token use, that's probably the big differentiator here. A couple models stand out: Ornith and KAT-Coder-V2.5-Dev are the most efficient models, with KAT in particular having a brevity unmatched by anything else. KAT is fucking fast, and I think despite its lower accuracy than Occamy, it has a place in my lineup as a subagent because it just gets. shit. done. Occamy is also interesting, as it is the only model that perfoms on a similar level to the base, but it uses 20-30% fewer median tokens. However, Occamy also had a number of runaway generations where the token count blew up, so it's total tokens/solve is actually higher than base.

In an effort to further distinguish Occamy from base, since Aider struggled to do that, I ran tau2-bench, an agentic tool-calling benchmark consisting of multi-turn interactions with a simulated counterparty. I figured this was a good bench to use as Occamy is post-trained specifically for 'co-work' scenarios, but not trained on this particular set. I used Qwen3.8-27B with reasoning effort set to low as the simulated customer in these conversations. The base model was able to pull away from Occamy in the harder retail domain of this benchmark, but Occamy resolved the issues in the airline domain at an equal rate while requiring fewer turns. Here's the results.

| model (Q8_0) | domain | pass^1 | tokens | sec/task | turns/task |
|---|---|---|---|---|---|
| Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) | airline | 80.0% | 5390 | 290 | 11 |
| Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) | airline | 78.0% | 6770 | 390 | 13 |
| Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) | retail | 86.0% | 4098 | 333 | 14 |
| Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) | retail | 79.8% | 3284 | 284 | 13 |

Overall, I think the results are clear, if unexpected: Occamy-1.0 is the only fine-tune that even competes with the base model on Aider Polyglot, but even it is not a clear winner. Tiel is noticebly worse than plain Ornith without the terseness toggle, and the terse mode doesn't even save any tokens. xhigh in froggeric/Sharp degrades accuracy slightly and bloats token use, which makes sense given the models were not RL'd for the extra thinking effort prompt. KAT-Coder-V2.5-Dev is the most efficient model, with accuracy nearly as good as Ornith and better than Tiel. Finally, Nex-N2.5-mini is a disaster.

💬 64 (+1) open on reddit ↗
▲
67
+1
32👁
r/LocalLLaMA · u/ThePrimeClock · 13d ago
SupersonicLabs/Julia-1 · Hugging Face

New open source Jev like model for running on local devices from a group called Supersonic Labs.

It's a 144M param local non-generative local classifier.

From their site:

Julia 1 opens our research into compact decision models. It builds on mmBERT-small, a multilingual encoder, and chooses among answers supplied with a question. It has 144.3 million parameters and runs on a CPU.
Our question: can one model classify, rank levels, and answer yes-or-no questions as the options change? Julia 1 is the first result of that investigation. Here are its successes, its failures, and the methods we used to measure them.
▲
66
+24
60👁
r/LocalLLaMA · u/Postmodern_Plunger · 7d ago
Inference Engineering for Dummies

Hi all! I am a former SWE who has recently transitioned into inference engineering. I launched a side hustle a few months back and I've just taken it full time due to excessive demand. Its been such an opportunity because the people optimizing runtime are far fewer than the people that are trying to build apps or offer AI solutions.

So I've been lurking around this community, and I've noticed a lot of people who seem to have massively suboptimal setups for their hardware, and I've grouped the biggest errors into several buckets. The purpose of this guide is to expose common inference bottlenecks and provide best practices for avoiding them within your hardware constraints.

RUNTIME:

I. Choosing the Right Runtime

This is the biggest mistake I see. Choosing the correct runtime for your architecture and model is the most important decision to make. In general, here are some rules to help you determine what runtime to use.

Firstly, Ollama is never optimal. Its just the simplest. If you want quick and easy and have extra RAM, it's a good place to start. It's very user friendly and requires less setup. But it just won't offer best inference speeds.

If your model requires cpu offload, then llama.cpp will be your best choice. If not and you're solely in gpu, VLLM will likely provide the best results. It's really as simple as that for 90% of cases. SGLang may be worth it if your workload involves Langgraph, as it is highly optimized for the tooling. Otherwise, stick to the above. Mainline branches are best, with community forks offering only highly niche performance boosts (i.e., for specific models/configurations, but are generally under optimized and not well maintained).

II. Optimizing and Maintaining Runtime

The other big mistake people make with runtime is failing to compile it with hardware specific flags. Not going to go through all of them here, Google can help you out. Just search "optimal runtime compilation flags for \[runtime\] using \[GPU, CPU, RAM type\]." The most missed/missed flags tend to be for architecture specific optimizations. Those are crucial.

Runtime should be recompiled (with optimal flags) any time \*any\* of the following occur:

\- System updates

\- Kernel/driver updates

\- Running a model released or modified later than your last compile

\- You haven't recompiled in over a month (recent updates often contain kernel or path optimizations)

MODEL CHOICE:

I. Quantization:

Quantization. Such a big word. Such little meaning. All you need to know is that it makes a model smaller. There are a million Q\_K\_X\_&$&$&$ quant sizes, so I'm not going to go over them individually. Rather, I will provide basic principles.

\- IQ quants are generally the best for the size. If you're choosing between IQ4\_XS or Q4\_K\_M, IQ4\_XS is both a smaller VRAM footprint and higher complexity.

\- Nonlinear (NL) quants are only ever going to be better if you have CPU offload. Even then, IQ quants often offer extra context space vs NL quants and thus are preferable.

\- If it's a quant you've never encountered, read the docs. It more than likely is highly optimized for that specific model and is indeed the one you should choose. Searching it or asking chatgpt \*will give you the wrong answer every single time for custom quants.\* You will only encounter these with custom tuned models.

\- Standard Q\_K\_M quants are best if you have absolutely no hardware constraints for the model you're running, as path optimizations are the best. If you have no hardware constraints, though, you could be running a better model. This is only useful for running simple models for simple tasks.

II. Task

Certain models excel at certain tasks. This is subjective and preference based but this is my list:

Coding: for API, anthropic. Hands down the best models. Opus and Sonnet 5.5 both excel in performance and their low token usage per task makes them more affordable than previous iterations. Deepseek models are the best budget choice. Qwen models are the clear winner for local inference on all fronts.

Writing: Opus/sonnet for technical writing, Gemini for creative writing. Gemma for local creative writing.

Video: Wan 2.2 for local in most cases will get it done, chatgpt and copilot both have surprisingly robust free image/video gen, Veo is the best paid.

HARDWARE:

Buy a budget box, or build your own from parts. I've managed to squeeze better performance out of an RTX 3060 and 128 gb RAM than a DGX spark across all categories for multiple models. The spark has an edge for dense models, but I was able to run higher complexity models overall on the other setup for 1/5 the price. AMD and Intel lag significantly on speed per price, but I've heard Intel has had some major gains recently. Have not confirmed myself though.

MODEL OPTIMIZATIONS:

I. Spec Decode (MTP)

\- If you have CPU offload, spec decode will \*always\* slow you down. The extra overhead compute isn't worth it if you don't have at least several hundred Mb/s bandwidth, which your CPU won't.

\- MTP is sometimes a baked in feature, and sometimes requires a special secondary model. Ensure you know how it works for your model and what flags to run.

\- Each model will be optimized for exactly 0-1 type of spec decode. Figure out which one it is (or isnt) rather than wasting your time testing methods.

II. Model Tuning

Just to show the kind of command optimization you can get, here is my sample command for running Qwen 3.8 Flash-Next, a 156b parameter model, on 12 gb VRAM (and 128 gb RAM) at 10 token/s decode and 150 token/s profile at 200k context:

\~/llama.cpp/build/bin/llama-server --flash-attn on --batch-size 1024 --ubatch-size 1024 --no-warmup --cache-reuse 256 --jinja --host 0.0.0.0 --port 8090 --presence-penalty 0.0 --repeat-penalty 1.0 -m \~/llama.cpp/LLM/Qwen3.8-Flash-IQ4\_XS/UD-IQ4\_XS/Qwen3.8-Flash-Next-UD-IQ4\_XS-00001-of-00003.gguf --temp 0.95 --top-k 20 --top-p 0.97 --min-p 0.05 -np 1 --chat-template-kwargs '{"enable\_thinking": true, "preserve\_thinking": true, "reasoning\_effort": "xhigh"}' --threads-batch 16 --threads 8 --gpu-layers 150 --n-cpu-moe 48 -c 200000 --override-tensor per\_layer\_token\_embd.weight=CPU -ctv q8\_0 -ctk q8\_0 --cache-ram 8192 --checkpoint-min-step 512 --ctx-checkpoints 4 --kv-unified --reasoning-preserve --load-mode mmap+mlock

That's a lot, right? It's every possible optimization you could apply. I'll go through them individually. This is llama.cpp specific, but you'll find the same flags with slightly different syntax apply to other runtimes.

\-flash-attn (-fa) on: forces flash attention optimizations and paths. Explicitly set to on to override any fallback. Auto can be optimal if the model is recent and is not yet optimized.

\-batch/-ubatch: batch is the decode chunks, ubatch is the prefill chunks. They must be multiples of one another, otherwise you're adding compute. Equal to one another is ideal for CPU offload, and a 2x-4x higher batch is optimal for full GPU loads. You'll need to play with these values to optimize. Batch/ubatch should be a power of 2 to optimize architecture. Intervals of 256 is typically good enough for testing.

\--no-warmup: prevents initial model poll to load weights. Removes unnecessary latency

\-- cache-reuse x: instructs the model to reuse cache values and scan for similarity at x token intervals

\-jinja: highly underutilized and important flag. Utilizes native chat template kwargs to ensure output consistency.

\--presence-penalty: flat penalty rate to words that appear in text. Used mainly for creative writing to prevent repetitive prose.

\--repeat-penalty: reduces liklihood of already used tokens being reused. Best used for preventing loops in thinking agents.

\-temp: model temperature– how creative the model is. 0 is completely deterministic, 1 is creative freedom.

Top-k: hard cutoff that keeps only the k most likely words. Each model will have recommended k values for thinking/instruct setups. Low key reduces hallucinations at the cost of repetition and loss of creativity

Top-p: includes P percentage of possible tokens. It reduces liklihood of hallucination dynamically.

Min-p: dynamic cutoff based on highest probability token. If the biggest probability token is 50% and min-p is 0.05, then the bottom 0.025 (2.5%) liklihood tokens will be excluded. Reduces noise without hampering creativity terribly.

Chat template kwargs: explicit chat template activations; newer runtime compilations should have flags for these. Controls model reasoning, reasoning effort, and internal chain of thought storage.

\-threads (-t): number of cores used for decode. Set equal to physical cores (hyperthreading will thrash cores and degrade results)

\-threads-batch(-tb): number of cores used for prefill. Set to double the number of physical cores, as hyperthreading helps here.

\--gpu-layers (-ngl) : total layers on GPU. Fit as many as you can without OOM.

\--n-cpu-moe: number of MoE layers offloaded to cpu. For MoE models, you should always offload these first and keep all layers on GPU if possible. Offload as few as possible to CPU.

\--override-tensor...: tells the runtime to offload the n-gram table if needed. For qwen 3.8 flash specifically.

\-ctk/-ctv: k and v cache quantization. K cache should \*never\* be below q8 unless youre running on less than 80k context. V cache can be q4 up to 150k context without issues, for the most part. Generally, q8 for both will be best for speed and is my recommendation to start with.

\--cache-ram: sets RAM aside for cache allocation to ensure it doesn't go to swap

\--context-checkpoints: the amount of checkpoints captured. Generally, you don't need as high as the defaults do. Leave the default if you have extra RAM, otherwise you may want to lower it.

\--kv-unified: tells all instances to run on the same kv cache pool rather than allocating individual cache.

\--reasoning-preserve: tells the runtime to retain CoT traces for evaluation. Prevents the model from getting stuck or looping as much when thinking.

\--load-mode: tells the runtime how to load in the model. No mmap generally loads slower but is more stable. Mmap+mlock (or just mlock) is a balance of both with fast loading and page faults initially but it stabilizes as you run it, mmap alone is fast but will cause constant page faults and slows down inference, especially on models with CPU offloading.

I hope this guide helps! I'd be willing to answer any specific questions or make any additions if there are additional areas the community agrees are major uncovered inference bottlenecks. Some claims are based on my personal experience and I am open to data based claim revisions or anecdotal counterclaims, so feel free to provide. Happy tuning!

💬 53 (+10) open on reddit ↗
▲
64
+3
30👁
r/LocalLLaMA · u/ciprianveg · 10d ago
Do you need some extra memory on your DGX Spark? post image

&#x200B;

I created this repo to help the DGX Spark users that have a spare 10-24 GB GPU at home to squeeze some extra memory out of a single Spark or a Sparks cluster.

It moves the spec-decode draft model off your Sparks onto that GPU: the freed GB of memory can be used for extra context, or better quant quality. Supports both TCP and RDMA, shipped as eugr-vllm compatible mods:

https://github.com/ciprianveg/gb10-vllm/tree/main/remote-dspark

▲
64
+9
42👁
r/LocalLLaMA · u/Top-Evidence174 · 11d ago
Mica v0.1 4B got diamonds in survival Minecraft on its first run. 26 decisions from an empty inventory. post image

I've been working on Mica, a 4B decision model, and wanted to see how far it could get in actual Minecraft, not a sim. New world, empty inventory, and the goal was a diamond pickaxe.

Last time I posted, it got an iron pickaxe. Honestly that took around 20 tries and it was pretty flaky. I've reworked the harness a lot since then. This time it went all the way to a diamond pickaxe, and once the harness was finished it did it on the first run.

It took 26 decisions and about 8 minutes of game time. It got wood, made a crafting table, then wooden and stone pickaxes, then iron and coal, a furnace and an iron pickaxe. After that it tunneled down to diamonds at y=2, mined three, put a crafting table down right there and made the pickaxe. Decisions took about 108 ms on average.

My favorite bit is around step 14. The planner wanted it to make planks to burn in the furnace, but Mica went and mined coal instead (0.69 vs 0.31) and then smelted all three iron at once. Which was the better call, honestly.

How it works: every step Mica gets the game state as text (inventory, nearby blocks, health, what happened last step) plus a few candidate commands, and it picks one. A Mineflayer bot running Mindcraft skills does the actual moving and mining. The panel on the right of the video shows each decision and its probabilities live. The bot also knows where the nearest diamonds are, so it isn't searching for them.

To be clear, I'm not saying a 4B model plays Minecraft on its own. What I wanted to show is that a model this small can sit behind a bot, read what's going on, and make the next call well enough to get all the way to diamonds.

I'm planning to release the harness soon. Mica will read Minecraft chat, so you can type what you want and it'll work toward it. Simple stuff like getting items, crafting or following you should work fine, but it'll struggle with anything really complex, like building a house.

Also, v0.5 should be out in the next 1-2 weeks. A lot of the architecture changed, and I did extra training on the parts where v0.1 was weak, so I'm expecting a clear jump in performance. The aim is to be at or near the top among 4B JEV-like models.

There'll be two versions: Mica v0.5 4B, and Mica v0.5 4B Distill Laya, which is light enough to run on pretty much any PC.

For Minecraft, I'm hoping v0.5 will be good enough to take down the Ender Dragon, and I did extra training specifically with that in mind. No promises, but it'd be really cool if it pulls it off lol

Oh and fun fact, Mica is a fully vibe-coded project

Model: https://huggingface.co/sky7350/Mica-v0.1-4B
Model code: https://github.com/akivet/Mica-v0.1-4B
Minecraft harness: coming soon

▲
63
 
33👁
r/LocalLLaMA · u/ZenZombie117 · 11d ago
Liked Muse, so I cut the 30B model in half by width, distilled it back, and it does 57 of 60 tool tasks its parent does 60 of

I've liked how Muse-Glimmer worked, so I wanted to see if I could produce a smaller "kid" out of it. Ornith's sharp decisions on when to think and which tool to call were the other thing I liked, so Ornith-1.0-9B got to be the policy teacher while the parent wrote the words. No RL anywhere, distillation only.

I present to you Xyntetik-Kvist-14B.

What it is good for

  • Smaller than the Muse parent but still manages most tool tasks: 57 of 60 held-out closed-loop tasks (contacts, weather, flights, currency, dates, units, stock), scored by re-executing the calls against ground truth, where the parent does 60.
  • Fits a 24 GB card whole at Q8_0 (15.4 GB) or the Q5_0 mix (10.3 GB) and serves an OpenAI-, Anthropic- and Responses-compatible API through Xyntetik Runner, so it drops into an agent loop you already have.
  • Every failed attempt is published beside it: 12 gated runs, 2 full passes, one shipped. The training record has the preregistrations, the amendments and the defects, so you can see exactly where it breaks before you build on it.
  • Give it a calculator tool for arithmetic. Without one it gets "17% of 2,340" wrong, and the card says so.

Numbers, from the card

| claim | number |
|---|---|
| parent | Muse-Glimmer-30B, cut by width (hidden 6,656 to 5,760, FFN 19,968 to 10,240, heads 32 to 24), all 52 layers kept |
| size | 14.44 B parameters; BF16 28.9 GB, Q8_0 15.4 GB, Q5_0 mix 10.3 GB |
| distillation | 6,000 steps, 98.3 M tokens, 162 hours, then 1,440 steps on agentic trajectories |
| fidelity to parent | KLD 0.762, margin-qualified top-1 84.0% on 45,056 held-out positions (a student's row, not the quant bar) |
| tool tasks | 57 of 60 held-out, re-executed against ground truth; parent 60, untrained control 0 |
| format and calls | 199 of 200 first turns well formed; 99 of 99 tool calls valid |
| attempts | 12 gated attempts, 2 full passes, attempt 12 shipped |
| weak spot | calc tasks 12 of 15 over 160; 7 of 160 runs end in a reasoning loop |
| serving | Runner v0.5.7 or later |
| licence | Apache-2.0 |

Links

EDIT: reading the comments, i should have said this first. this is not a general drop-in for Muse or a gemma4 replacement, and it was never going to be on my compute (98M distillation tokens vs the trillion a real distill wants, i simply lack the compute). the purpose was more on getting the tool calling right. IQ4_NL mix (7.6 GB) is up now too, it scores the same 57/60 on the tool tasks but does not fit an 8 GB card whole (50/52 layers on a 3070, ~5 tok/s).

▲
59
+3
27👁
r/LocalLLaMA · u/streppelchen · 11d ago
Minisforum MS-S1 MAX-P495 @ €7.799,00

MINISFORUM MS-S1 MAX-P495 – Minisforum EU

Expected to ship mid october.

At that price point, it doesn't make a whole lot of sense in my opinion.

I get that ram prices are where they are, i get that it's a newer model of hardware, but twice the price for 50% more ram and \~5-10% more performance is just hard, especially when compared to the recently released m5 ultra studio.

▲
58
+5
22👁
r/LocalLLaMA · u/jjusko20 · 8d ago
Update #2: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch post image

Last update: https://www.reddit.com/r/LocalLLaMA/comments/1wu9ksu/update\_yandexaliceai\_80ba3b\_fine\_tune\_progress/ \- basically, an instruct fine tune on the base model using a synthetic distilled data set. I've been posting regular updates so I imagine at least a few people have seen this.

Live stream: https://figure-bios-expect-cio.trycloudflare.com/

UPDATE: Finished train. hopefully some examples soon.

The initial train is finally almost done, after about 48 hours of humming. While the loss curve looks a little crazy, I've done some analysis (and some chatting with the LLMs) to understand that my average loss each epoch has been steadily decreasing (few reasons the loss curve looks wacky, vocabulary size, low to high token counts in epochs, etc) - but I'm pretty happy with what I'm seeing so far.

I'm post training the attention and the shared expert, and leaving the base experts frozen - this is a behavioral and logic fine tune that preserves the original yandex training data.

I plan on, within the next few days, releasing a few gguf quants of this, along with a llama.cpp patch for running it locally. I'm not sure how well the initial fine tune is going to work out - loss looks good but I'll have to do some evaluating. Either way, I plan on continuing training with reinforcement learning and an extended SFT set, as I have room and a ton of capacity left in my QLoRA adapter. I'll release this version as a public checkpoint anyways though (kinda like how deepseek did it) so people can play around with it and hopefully get excited for new checkpoints.

Cheers! Stay tuned, this is a pretty fun model size to play with, I'm excited to release the instruct version. I'll open source whatever you guys want out of this - I already open sourced the distillation engine (see SFTMill, it's been posted in here in the last few days) - but I also have a custom kernel for training this for V100s and a few other patches I can share (this training has been plugging away on 3, 32gb v100s - man it took a while to get that to work). Mandatory plug for my own goals: if you're hiring remote or in NYC for a dev or ml engineer, hit me up!

God I hope it writes the adapter when this is done I didn't audit that code well enough.

💬 18 (+1) open on reddit ↗
▲
57
+3
54👁
r/LocalLLaMA · u/Effective-Ad2060 · 8d ago
We benchmarked 18 RAG pipelines against an agent loop on Google's FRAMES. The best pipeline hit 78.9%. The agent loop hit 92.7%.

Hybrid search, reranking, query decomposition, and query expansion are often treated as must-haves for good RAG. We wanted to see how much each actually helped, so we tested them. Same model, same embeddings, same documents, across all 824 multi-hop questions in FRAMES.

We built 18 pipeline variants. The best one scored 78.9%. Our agent loop (with retrieval tools) that could read the results and search again scored 92.7%—roughly the same as giving the model the right articles upfront.

The reranker results might surprise you. A small reranker dropped our best pipeline’s accuracy by 9 percentage points, while a larger one barely helped. I’d already suspected reranking wouldn’t help much here, but wanted to test that assumption.

Another thing we noticed: models sometimes fill in gaps from memory, even when you explicitly tell them to stick to the retrieved documents. Those answers can still be full of citations. We ended up checking every correct answer against what the system had actually read.

Here’s the write-up if you’re interested:
Agentic RAG vs. traditional RAG on FRAMES

Full disclosure: I work on PipesHub, which is open source. The benchmark code and runbook are in the repo: https://github.com/pipeshub-ai/pipeshub-ai/tree/frames

Quick note on what the numbers mean: they're end-to-end answer accuracy, not retrieval scores. Every answer was graded by an LLM judge (Claude Sonnet 5) using the FRAMES paper's own grading prompt, and independently by a second judge (Gemini Flash 3.8). The two agreed on almost every answer (Cohen's κ 0.93–0.98). We also checked each correct answer against the text the system was actually shown, so answers that came from the model's memory don't count as retrieval wins.

💬 54 (+4) open on reddit ↗
▲
57
+4
30👁
r/LocalLLaMA · u/MomentJolly3535 · 12d ago
Swift 1.5 Qwen3.8 27b (A must-have for low thinking!)

Just made this post for those who missed it : https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b

UkisAI released their updated Qwen 27B (tuned for token efficiency). I grabbed the IQ4\_XS quant to test against Unsloth's Q4\_K\_S:

Low-thinking: UkisAI consistently beat Unsloth in most of my tests.

High-thinking: Unsloth still pulled ahead here.

I was struggling with a custom script in Directory Opus. I gave it to Gemini Flash (medium thinking on Antigravity free tier) it looped for 40 minutes, tried many things, burnt all the weekly limit-tokens, and failed to solve it.

Fed the exact same problem to this 27B model: Fixed it completely in 6 minutes on an old 3090 (67 t/s)

Honestly i was kinda impressed, didn't expect an IQ4\_XS quant of a 27B model in low thinking to beat a major cloud model.

▲
54
+14
37👁
r/LocalLLaMA · u/Designer_Cost8989 · 9d ago
Index-Translate: 150 text languages, plus document translation, multilingual subtitles and dubbing

Quick update: we’ve opened a free public API for Index-Translate-35B-A3B! It’s OpenAI-compatible, and you can get started with our Python script—no extra dependencies needed.

I'm part of the BiliBili Index LLM team. We're sharing Index-Translate and its companion models for translating text, documents, and videos.

Index-Translate supports 150 text languages, with 2B, 9B, and 35B-A3B (preview) options. You can specify terminology, writing style, and output format—for example, keeping product names consistent, translating in a casual tone, or preserving JSON and placeholders during localization.

There are also models for more specific workflows:

  • Index-NativeLong: translate whole documents, using their context to help keep names and terminology consistent across passages.
  • Index-Homura: set a syllable budget for translated lines, useful for fitting a dubbing script.
  • Index-Echo: generate multilingual subtitles or translate speech into speech, using the source speaker's voice as a reference.

The attached video shows English → Japanese dubbing, followed by an English clip with subtitles in six languages. The 150-language coverage applies to the text models; Echo supports a smaller set of language pairs.

https://reddit.com/link/1wugf2t/video/9zajzj72zpsh1/player

Code and released weights are Apache-2.0.

Try the demo · GitHub · Models

What would you try it on—video subtitles, game localization, or documents? We'd especially appreciate examples where it gets your language pair wrong.

💬 22 (+3) open on reddit ↗
▲
54
+3
22👁
r/LocalLLaMA · u/jjusko20 · 9d ago
Update: Yandex/AliceAI 80B-A3B fine tune progress

loss curve \(taken from the last micro of every step, to explain the variation\)

some help from gemini 3.8 flash high

About 40% of the way done with the initial fine tune. The loss is so spiky because I accidentally used the last loss of each micro, rather than the average of each step

The training live stream is at: https://figure-bios-expect-cio.trycloudflare.com/ \- and it allows you to inspect any and all of the training data I'm using, if you're interested - I can also provide those roughly 3.5k examples as a dataset. It was generated from sftmill

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wslgkw/watch\_me\_posttrain\_aliceaifoundation80ba3b\_from/

▲
53
 
25👁
r/LocalLLaMA · u/sloptimizer · 10d ago
RAM Offloading with vLLM - tcclaviger appreciation post post image

Thanks to tcclaviger, vLLM now has expert RAM offloading support (link). This makes frontier models much more accessible on a local setup!

I was able to run the original DeepSeek-V4-Flash-Vision-Exp on four R9700s.

podman run --rm -it \
--init \
--network host \
--ulimit memlock=-1:-1 \
-v /models:/models:ro \
-v ~/.vllm-cache:/cache \
-e VLLM_ROCM_USE_AITER=0 \
-e ROCR_VISIBLE_DEVICES=0,1,2,3 \
-e VLLM_CACHE_ROOT=/cache/vllm \
-e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \
-e TRITON_CACHE_DIR=/cache/triton \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--annotation run.oci.keep_original_groups=1 \
--security-opt label=disable \
--security-opt seccomp=unconfined \
--shm-size 160g \
docker.io/tcclaviger/vllm@sha256:ef99b3d07c3f15e7978528c7510762ba024df9ab4242070d8ed092cd4cc1a694 \
/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
--served-model-name DeepSeek-V4-Flash-Vision-Exp \
--tensor-parallel-size 4 \
--enable-expert-offload \
--expert-offload-mem 160 \
--reasoning-parser deepseek_v4 \
--reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
--tool-call-parser deepseek_v4 \
--enable-auto-tool-choice \
--max-num-seqs 8 \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-batched-tokens 2048 \
--kv-cache-dtype fp8 \
--block-size 256 \
--max-model-len 256000 \
--gpu-memory-utilization 0.97 \
--mm-processor-cache-gb 4.0 \
--override-generation-config '{"max_tokens": 128000, "temperature": 1.0, "top_p": 0.95}' \
--speculative-config '{"method":"dspark","model":"/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp","num_speculative_tokens":3,"draft_sample_method":"probabilistic","enable_adaptive_verification":false}' \
--compilation-config '{"cudagraph_capture_sizes": [4,8,12,16], "max_cudagraph_capture_size": 16}' \
--host 0.0.0.0 \
--port 8090

▲
50
+16
42👁
r/LocalLLaMA · u/menage_a_un · 7d ago
I've ended up with an AI lab in a public community college. What should we actually be teaching?

Looking for some ideas from people who know a lot more about this than I do.
We've got funding for a small AI lab in a public further education college in Ireland (roughly community college in the US).
The hardware is reasonably decent. The goal is to give students useful skills beyond just using ChatGPT.
If you had the lab, what would you teach them?

💬 70 (+18) open on reddit ↗
▲
49
+6
27👁
r/LocalLLaMA · u/netherreddit · 10d ago
Inference Engines will become a series of one-offs

ninfer, dwarfstar, Splash, llamAmpere, gufo, etc.

We've all seen them popping up, great tok/s, people loving them. Forks of llama.cpp or another engine, or made from scratch.

For better or worse, the list will continue to grow

They work so well because they dodge a main difficulty of software, generality, and just implement for a single model/hardware combo (or a few), and then optimize kernels/compute graph for that one case. Highly 'overfit' codebases that beat well-known inference engines (llama.cpp, vLLM, etc) and incidentally will be completely forgotten in 6 months.

But new ones will take their place...

THESIS

One-off engines will become the norm. Llama.cpp, vllm, etc, will not make sense for most people to use, because they're slower

A few axioms you probably accept:

  1. the more general a codebase, the harder it is to cleanly fit new features in over time. This reduces the pace of innovation. The smaller, the faster
  2. AI coding is getting better and cheaper. Thus, the barrier to creating an inference engine is dropping
  3. many coding tasks are difficult to completely give to AI (or a human) because they are not fully specified. But "Make tok/s go up" in an inference engine fork for one hardware/model combo is fully specified, and is therefore a great candidate for 100% autonomous implementations to be perfectly fine in terms of quality (as long as correctness tests are included, which is trivial). No human bottleneck.

String these axioms together, and I arrive at

  1. general engines, like llama.cpp, will perpetually lag behind these one-offs in development speed, and therefore token speed
  2. none of the one-offs will be able to maintain generality and dev speed over time
  3. one-off inference engines for specific hardware/model combinations will continue to proliferate, and be loved

Thank you for coming to my ted talk

IMPLICATIONS

  1. This thesis brings up an interesting question: What elements of inference engines WILL remain in common?

Most obvious example: it would be annoying to have a different usage API for every engine, so we already standardized on OpenAI API compatibility years ago.

Is that also true for cli arguments/configs? The packaged gui (llama-server)? Benchmarking tools (llama-bench)? Logging, model format, Etc?

One-off engines that replicate the experience of everything wrapping the inference itself will be more seamless to adopt. Case in point, the main reason I haven't tried any of these new one-off engines myself is it was annoying enough to figure out how to drive llama.cpp properly. Don't want to do that again unless it's really worth it.

There's probably a place for an open source project that standardizes all of this and makes it easy for one-off engines to adopt.

  1. Maybe we'll see more 'half-general' inference engines that just target one hardware platform. So still general on the dimension of models, but not on hardware. Splash could be an example.
  1. Nobody wants to continuously scan github/reddit/x for the best inference engine for their model/rig. Some will just have their agent custom make one. But I think a larger number will not do that. So, hardware-specific communities will form. Think r/appleM2Max32gbLLM and r/4090And64gbRamLLM, etc (however that actually ends up organizing. exaggerating a bit on the names.)

ALTERNATIVE FUTURES

Scenarios where the one-off future doesn't happen:

  1. General inference engines find a way to 'plugin-ify' the model/hardware specific kernels and compute graph so you can swap them at runtime. So you'd download not just a .gguf, but also an .inference\_recipe to go with it, which contains the optimizations for your specific hardware, for that specific model. Maybe those optimizations will make it into llama.cpp mainline in 3 months, but you can use them today, without a fork.
  2. General inference engines find a way to AI-ify their workflow so much that they maintain quality and codebase coherence but also achieve the same development velocity for each model/hardware platform as the one-offs. I think this is the best for everyone involved.
  3. The full vision of something like MLIR, or Mojo is realized to a sufficient degree. ie writing hardware-optimized kernels is fully and invisibly done by compilers, no hardware-specific tinkering needed anymore for each silicon platform) Then, inference engines that cover all hardware/models would be much more manageable to maintain and add features to. btw, if you really want to have an impact, solve this. The world will thank you for centuries to come. Unfortunately not many people have even conceptualized this as a goal.

P.S. there's growth in a dimension separate from single model/hardware engines which is more like "frontrunning a general inference engine's features because it's slower to pull in PRs". Freetoken, BeeLlama, etc. Not as model- or hardware- specific as the other examples I've given. Haven't thought much about that dimension.

💬 185 (+8) open on reddit ↗
▲
46
+2
14👁
r/LocalLLaMA · u/KingGongzilla · 11d ago
Swift 1.5 + HyperQwen = 37% less task completion time at 100+ tps w/ 150k context on RTX 3090

Hi everyone :)

The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090.

To get the best of both worlds, I adapted Swift 1 and Swift 1.5 for HyperQwen and benchmarked them against HyperQwen’s specialized W4A16 AutoRound fast quant for speed and quality.

Performance

|Model|Average time/task ↓|Average output tokens/task ↓|Decode tok/s ↑|
|:-|:-|:-|:-|
|Qwen - HyperQwen fast quant|108.1 s|8,985|112.1|
|Swift 1.0 + HyperQwen|66.2 s|5,245|105.9|
|Swift 1.5 + HyperQwen, INT8 heads|72.2 s|5,751|104.0|
|Swift 1.5 + HyperQwen INT4 heads|68.2 s|5,669|107.2|

All models ran on one RTX 3090 24GB with FP8 KV cache and 150k configured context. Average task time covers about 630 tasks from different benchmarks listed below.

Swift-1 seems to use the fewest tokens and has fastest task completion time. Swift-1.5-INT4 achieves roughly 37% lower average time per task compared to the HyperQwen Qwen fast model. Despite slightly lower TPS (due to lower draft acceptance) it finishes sooner because it generates fewer tokens.

Quality

There are some minor quality and performance tradeoffs between the models:

|Test|Qwen HyperQwen fast|Swift 1.0|Swift 1.5 INT8 heads|Swift 1.5 INT4 heads|
|:-|:-|:-|:-|:-|
|GSM8K, 200 questions|97.5%|98.0%|98.0%|97.5%|
|IFBench, 300 prompts, strict|74.0%|73.3%|73.7%|72.3%|
|LiveCodeBench, (100-problem subset)|90%|89%|89%|91%|
|Custom tool-call/JSON eval|29/30|28/30|30/30|30/30|
|English/Python perplexity ↓|6.551|6.605|6.643|6.679|

Applied changes to Swift models to adapt for HyperQwen:

Changes to Swift models:

  • Kept the upstream AWQ INT4 model weights and converted embeddings to INT8.
  • Swift 1.0 and Swift 1.5 INT8-head variants: quantized the output head and MTP (multi-token prediction) linear layers to INT8 and added HyperQwen’s reference draft vocabulary for speculative decoding.
  • Swift 1.5 INT4-heads: quantized the output head and MTP linear layers to GPTQ INT4 instead, and built a Swift-specific 65,536-token draft vocabulary.

Setup

If you want to try it yourself, point your coding agent at these setup instructions and ask it to set up Swift 1.5 + HyperQwen on your machine.

All three models can be found here:
https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks

Big shoutout to UkisAI for Swift and syv-ai for HyperQwen! It's genuinely insane to be able to run these models on an RTX3090 at those speeds!

▲
45
-2
18👁
r/LocalLLaMA · u/Adventurous-Gold6413 · 12d ago
Which of the 16gb VRAM qwen3.8 27b’s is the best?

I’m having a hard time finding out which one gives you fastest speed, maximum context with best possible quality. I can run unsloth qwen3.8 27b iq4\_xs with 65k q8 kv, context without MTP and vision offloaded to cpu. But also kinda slow for agentic work at like 30ish tok/s )I mean it’s acceptable) But that is like bare minimum for harness stuff, I know many people use Q3 quants but are Q3 quants really safe? Like you gotta think I won’t only be using it for vibe coding, but also general tasks. Where general knowledge quality would be nice to keep intact. There are so many quants like IQ4XS smaller, Or GRQ or whatever those quants are called or YMQ, I don’t even know anymore. Which one is the best?

💬 83 (+1) open on reddit ↗
▲
43
+10
36👁
r/LocalLLaMA · u/SnooPredictions515 · 8d ago
Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

I've been working on getting the 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac. In my earlier post, I shared a custom expert-streaming fork of llama.cpp . It worked, but decode capped out around \~23–27 tok/s and slowed down as context grew.

Today I'm releasing Slipstream: a compiled C++ Metal inference engine with native SSD expert streaming and speculative drafting for Apple Silicon.

The main result: If you already downloaded my original V3 model (34k+ downloads), you don't need to re-download anything. You can run that exact checkpoint on Slipstream for a 1.76x speedup: 41–52 tok/s (up from 23.1 tok/s in llama.cpp) on the same 64GB Mac.

Even better: decode speed doesn't collapse at long context. Across 3,086 live requests in real coding sessions, it stays flat at 33–44 tok/s all the way out to 130,000 tokens.

Previous posts for context:

Open source resources:

1. How to run your existing V3 model on Slipstream

If you have the model from the last post (~/models/qwen38-flash-next-v3), you can point Slipstream directly at it.

Step 1: Clone & build (under 1 minute)

git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4

Step 2: Download the model (if you don't already have it)

Downloads the 3 GGUF shards + MTP draft head (~95.5 GiB total) huggingface-cli download nitinpanj/qwen38-flash-next-v3 \ --local-dir ~/models/qwen38-flash-next-v3

Step 3: Raise wired GPU memory limit & serve

Raise wired GPU memory limit once per boot (required on 64 GB Macs): sudo sysctl iogpu.wired_limit_mb=59392 # Serve your existing model: ./slipstream serve --model ~/models/qwen38-flash-next-v3 --port 8090

First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into <model>/prepared/ (\~5–7 minutes). Subsequent launches load in \~10–15 seconds.

The server exposes a standard OpenAI-compatible API (http://127.0.0.1:8090/v1/chat/completions) ready for curl, Oh My Pi (omp), Claude Code, or OpenCode.

2. Speed: llama.cpp Fork vs. Slipstream (Same V3 Checkpoint)

Here is a direct head-to-head comparison running the exact same 95.5 GiB model files across 6 reasoning and coding tasks on the same M5 Pro (64 GB unified memory, temperature 0.0):

|Domain / Task|Prompt Task|llama.cpp Fork|Slipstream|Speedup|llama.cpp TTFT|Slipstream TTFT|
|:-|:-|:-|:-|:-|:-|:-|
|Math Reasoning|GSM8K (eggs problem)|24.0 tok/s|43.6 tok/s|1.82x|4,024 ms|2,337 ms|
|Math Derivation|MATH-500 series ($p - q$)|24.3 tok/s|43.1 tok/s|1.77x|1,655 ms|1,587 ms|
|Constraint Logic|3-chair deduction|25.4 tok/s|46.0 tok/s|1.81x|1,469 ms|1,042 ms|
|Python Coding|merge_intervals ($O(N log N)$)|19.7 tok/s|35.0 tok/s|1.77x|1,507 ms|1,070 ms|
|Systems Coding|Rust CSV parser|22.7 tok/s|37.5 tok/s|1.65x|1,257 ms|859 ms|
|Tech Writing|Multi-head attention|22.5 tok/s|39.4 tok/s|1.75x|1,267 ms|843 ms|
|AVERAGE|Across all 6 tasks|23.1 tok/s|40.8 tok/s|1.76x|1,863 ms|1,290 ms|

https://preview.redd.it/vh1danbnzwsh1.png?width=3000&format=png&auto=…

What made Slipstream faster:

  1. Asynchronous layer-ahead prefetch (fcntl(F_RDADVISE)): In llama.cpp, synchronous page reads for missed expert matrices stalled the GPU on NVMe latency (\~475 ms per chunk). In Slipstream, non-blocking read-ahead hints stream upcoming expert layers from SSD into RAM while the GPU is still executing the previous layer, cutting prefill staging latency by 28%.
  2. Hybrid MTP + Prompt Lookup speculation: During tool calls and code generation, Prompt Lookup Decoding (PLD) matches prompt anchors in under 50 ns with 0 allocations, preventing draft rejections. This lifted tool-calling decode from 5.6 tok/s to over 45 tok/s.
  3. Metal GPU-mapped n-gram tables: llama.cpp faulted on the 26.8 GiB n-gram table during prefill. Slipstream maps and gathers n-gram embeddings directly in Metal kernels.

3. Context Scaling: Real Telemetry up to 130,000 Tokens

On standard Transformers, decode slows down sharply as context grows because the KV cache swells and memory bandwidth saturates.

Qwen3.8-Flash-Next avoids that through its hybrid architecture:

  • 48 recurrent linear DeltaNet layers (fixed $128 \\times 128$ hidden state, $O(1)$ memory growth with context).
  • Only 16 full-attention layers.

Here is actual telemetry collected across 3,086 live requests during real agent coding sessions on my M5 Pro (64 GB):

|Context Range (Tokens)|Live Runs|Average Decode|Median (p50)|Peak Decode|Average TTFT|Notes|
|:-|:-|:-|:-|:-|:-|:-|
|< 1,000|314|41.5 tok/s|41.9 tok/s|59.8 tok/s|2.16 s|Short baseline|
|1k – 4,000|21|41.0 tok/s|42.5 tok/s|64.5 tok/s|5.26 s|Small documents|
|4k – 8,000|58|43.6 tok/s|43.2 tok/s|67.2 tok/s|7.36 s|Code review turns|
|8k – 16,000|117|43.6 tok/s|44.6 tok/s|58.2 tok/s|7.91 s|Multi-file context|
|16k – 32,000|562|38.2 tok/s|40.9 tok/s|58.0 tok/s|13.59 s|Deep agent session|
|32k – 64,000|1,029|35.0 tok/s|37.5 tok/s|55.6 tok/s|13.24 s|Large repo refactor|
|64k – 96,000|650|32.4 tok/s|34.7 tok/s|53.9 tok/s|12.81 s|Multi-turn transcript|
|96k – 130,000|364|32.9 tok/s|33.3 tok/s|43.8 tok/s|7.95 s|Cache-hit deep turns|

https://preview.redd.it/qruz5nmpzwsh1.png?width=3300&format=png&auto=…

Takeaway: Decode speed stays between 33 and 44 tok/s all the way out to 130k tokens. Even at 130k context, it generates tokens faster than stock llama.cpp did on a 500-token prompt.

4. Optional: Swift KV-Sparse Model Variant

If you want higher reasoning accuracy and lower KV cache memory, I also put together an optional Swift variant of this model: Swift-Qwen3.8-Flash-Next-V3.

What Swift changes:

  • KV-Sparse Attention: Replaces standard dense attention with KV-sparse layers distilled from Swift-1.5, cutting down RAM pressure at long contexts.
  • Spliced Q8 Donor Backbones: Slices 686 high-precision Q8 donor tensors into resident backbone layers for sharper representations.
  • Concise Reasoning: Distilled to eliminate repetitive thinking loops in deep contexts.

Both models run on Slipstream using the exact same engine command. Here is how they compare across 145 paired evaluation problems (temperature 0.0, seed 1234):

|Domain / Benchmark|Items|Original Flash-Next V3|Swift-Flash-Next V3|Accuracy Delta|Original Decode|Swift Decode|
|:-|:-|:-|:-|:-|:-|:-|
|AIME 2025|20|45.0% (9/20)|45.0% (9/20)|0.0%|44.3 tok/s|44.3 tok/s|
|MATH-500 (L4–5)|35|60.0% (21/35)|62.9% (22/35)|+2.9%|44.8 tok/s|44.8 tok/s|
|GPQA Diamond|35|45.7% (16/35)|54.3% (19/35)|+8.6%|44.8 tok/s|44.8 tok/s|
|GSM8K|25|96.0% (24/25)|96.0% (24/25)|0.0%|45.6 tok/s|45.6 tok/s|
|HumanEval|25|92.0% (23/25)|92.0% (23/25)|0.0%|40.6 tok/s|40.6 tok/s|
|Hard Systems Logic|5|100.0% (5/5)|100.0% (5/5)|0.0%|39.2 tok/s|39.2 tok/s|
|OVERALL|145|67.6% (98/145)|70.3% (102/145)|+2.8%|43.9 tok/s|44.4 tok/s|

https://preview.redd.it/gt1edfvrzwsh1.png?width=3000&format=png&auto=…

To run the Swift model instead:

huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3

./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090

5. Foundation for Qwen4

The core primitives in Slipstream:

  • 512-route sparse MoE streaming with SSD prefetch
  • QSA (Quasi-Sparse Attention) indexer & selection kernels
  • Hyper-connection mixing and per-layer embedding gathers
  • Metal GPU-mapped n-gram table gathers
  • Single-lane speculative verification with PLD & MTP

...were built around this hybrid architecture. If Qwen4 adopts a similar blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream should be able to run Qwen4 locally on consumer unified memory hardware on day one.

6. Hardware Tested & Porting to NVIDIA / AMD

  • Hardware tested: All testing and benchmarking were done on an Apple MacBook Pro (M5 Pro, 64 GB unified memory, 2 TB SSD).
  • CUDA / ROCm ports: I don't have access to modern NVIDIA or AMD GPU hardware, so I can't build or test CUDA/ROCm backends myself.
  • If you have hardware and want to help port this: If anyone in the community has NVIDIA or AMD hardware and wants to help bring expert streaming and hybrid speculation to Linux/Windows, I'm happy to help collaborate on the port. Feel free to open an issue on the repo or DM me.

7. Credits & Upstream

  • Splash Team (Incoai): Full credit to the creators of Splash. Their C++ Metal speculative decoding design and memory architecture provided the foundation for this work. I will prepare a clean PR proposing these Flash-Next and SSD streaming extensions to the Splash upstream repo in case they want to incorporate them.
  • ds4 Team: For their insights on Metal router numerical precision (Taylor polynomial softplus expansion) and streaming scheduling designs.
  • Qwen Team: For training Qwen3.8-Flash-Next and releasing the hybrid linear MTP architecture.
  • ukisai: For the Swift-1.5 distillation work enabling KV-sparse reasoning.
  • bartowski & unsloth: For donor quants and quantization tooling.
  • mihailescu2m: For the initial expert streaming work in llama.cpp.
💬 29 (+2) open on reddit ↗
▲
43
-6
20👁
r/LocalLLaMA · u/Haunting-Stretch8069 · 10d ago
Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide

I'm running Qwen 3.8 27B Q4 XS with \~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.

Build llama.cpp with Vulkan:

cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j

Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then:

llama-server \
--model Qwen3.8-27B-UD-IQ4_XS.gguf \
--mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \
--n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \
--batch-size 2048 --ubatch-size 512 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \
--load-mode none --fit off \
--cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \
--no-context-shift --jinja --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080

---

Edit: The process is documented here: https://zenodo.org/records/23088880. Feedback will be integrated into the upcoming Qwen 4 setup.

▲
42
+5
29👁
r/LocalLLaMA · u/bring_back_the_v10s · 9d ago
How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks?

I've been closely following the rise of the Strata inference engine and as someone with 28GB VRAM and 32GB RAM I'm itching to buy 32GB more RAM just to use Flash Next. But of course before I make such a financial commitment as a member of the GPU-poor class like myself, first I need to make sure the IQ3 quants are worth it. My use case is primarily agentic coding tasks with harnesses like Pi or OpenCode.

Has any of you guys used Flash Next IQ3 for relatively serious coding? Is it worth it? Compared to, say, Qwen 3.8 27B Q4 or Q5.

https://github.com/Niko1221/Strata/

💬 72 (+5) open on reddit ↗
▲
42
+4
20👁
r/LocalLLaMA · u/Usual_Maximum7673 · 11d ago
Trained locally: ultra-fast 0.8B/2B System 1 decision models that match Jev on benchmarks and Doom, ~30 ms per decision (open weights)

TL;DR: The Jeff models are a set of Qwen3.5 and Gemma fine-tunes for zero-shot classification: small, efficient, open-weight models with respectable out-of-the-box performance that can be slotted right into code or fine-tuned/LoRA-trained as needed. Give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in \~28 ms on an M4 Max (see caveats below).

Maybe equally exciting for open model enthusiasts like myself, everything was done on local hardware: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing - all connected and monitored from my Android phone via Tailscale. Apache 2.0, Jev-compatible API. Weights: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B · Code: github.com/firelex/jeff · Videos: games table

When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware.

So here's what I did:

  • The 0.8B trains in about 2 hours and the 2B in about 3.5, on one workstation GPU (RTX PRO 6000, 96 GB).
  • \~31k synthetic training questions written and checked by Qwen3.8-Flash-Next on two DGX Sparks. No cloud GPUs, and no closed-model output in the training data.
  • The rest of the 271k training questions are public datasets converted into decisions, plus 10k code-built probability questions.

Benchmarks (4,599 questions: BBH, Financial PhraseBank, JudgeBench, RAGTruth, WinoGrande):

|Model|Untrained base|Jeff (trained)|Calibration error|
|:-|:-|:-|:-|
|Qwen3.5-0.8B|45.3%|79.1%|0.049|
|Qwen3.5-2B|46.5%|83.1%|0.028|
|Jev (published)||83.0%|≈0.06|
|AutoJev-27B (published)||84.9%|—|

The caveat: the published numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86–89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64–68% against Jev's 94%, and \~50% on JevBench's hard tier against \~73%. See the HuggingFace model card for details. But that's not surprising, and I don't think it matters. No 0.8B or 2B model reasons like an LLM, and I don't think anyone should expect it to. The Jeff models are extremely fast judgement-callers (much faster than Jev), and have reasonable out-of-the-box performance. In one of my apps, I used the 0.8B model for voice-based navigation, and with a quick fine-tune, I got to real-time performance (24ms) at almost 100% accuracy.

Now the fun part: games, as a zero-shot test. Games are not the ideal zero-shot test, but they're fun, and the TypeSafe guys (Jev) did it, too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one. The options say what each move leads to, never which one is right.

|20 episodes each|Doom (kills)|Frogger (crossings)|Pac-Man (pellets of 98)|
|:-|:-|:-|:-|
|Random moves|−0.05|0|11.2|
|Hand-coded rule bot|6.55|10.25|94.1|
|Qwen3.5-0.8B, untrained|5.0|1.0|25.8|
|Jeff 0.8B|6.55|10.3|57.0|
|Jeff 2B|−0.9|6.0|41.2|

Jev's published Doom score is also 6.55, but its prompt spells out the aiming rule (fire when the bearing is between −8 and +8 degrees) and it takes \~212 ms per call over its API. Jeff gets "the nearest monster is a little to your left" and decides in \~29 ms on my Mac.

Lessons learned:

  • System 1 models are here to stay. Having the ability to process unstructured data at software speed inside an app is extremely powerful. And being able to do this locally is fantastic.
  • A small model is a classifier, not a planner. Models in the 0.8B-2B range don't reason like Qwen3.8-27B or Jev. But they also don't need to. As long as you present the options in the right way, you can get up to 50 decisions per second (depending on your hardware).
  • Fine-tune it if needed. If the models' zero-shot performance isn't good enough for you, fine-tune them briefly or add a LoRA adapter.
  • Wording matters enormously. Play around with how you present the options. Giving Frogger's final step option the same words as every other forward option ("safe, and one row closer to the goal") took one episode from 15 crossings to 23. Previously, the frog just stayed on the last log.
  • Bigger isn't better. As the game tests showed, the untrained 2B is already more risk-averse than the untrained 0.8B (in Doom it prefers "turn away from the nearest monster"), and training made it hesitate in Pac-Man. That's probably why the 0.8B beat the 2B.
  • Benchmarks don't predict play. Untrained Gemma 4 E2B beats both untrained Qwens on the benchmarks (62.5%) and plays every game worst: right most of the time, not reliably, and in a real-time loop the mistakes compound.

Happy to answer questions about the pipeline (synthetic data from a local teacher, leak filter, calibration) or the game harness. Everything, including the videos, is linked above.

▲
41
+12
29👁
r/LocalLLaMA · u/Terminator857 · 8d ago
China and the memory market

Once china sets its goals for dominating a market it wins. Usually takes many years, but it happens. Can't compare the political will of a country versus profit and loss thinking of a corporation.

China will eventually win in the memory market and current memory makers are at an unfair disadvantage.

https://www.tweaktown.com/news/112680/chinas-cxmt-is-on-track-to-nearly-match-microns-dram-production-capacity-by-the-end-of-2026/index.html Quote:

CXMT will finish 2026 with approximately 350,000 wafer starts per month (WSPM) of DRAM capacity, which is just 25,000 WPM less than Micron.

... by 2030, its total capacity will increase to around 1.41 million WSPM, according to Citrini. CXMT alone is projected to build new production capacities in Beijing, Hefei, and Shanghai, to expand its production capability to 950,000 WSPM in 2030, assuming everything goes as planned.

/end quote

Is China hoping for a RAM price drop crash to extinguish the competition?

Additional references:

  1. https://www.tomshardware.com/pc-components/dram/chinas-cxmt-targets-30-percent-dram-memory-market-share-by-2030-with-sixth-mega-fab-future-plans-bottlenecked-by-access-to-advanced-chipmaking-tools
  2. https://www.trendforce.com/news/2026/09/24/news-cxmt-ymtc-ramp-memory-capacity-but-chinas-ai-cloud-boom-could-soak-up-new-supply-through-2027/
  3. Microns quarterly report: https://www.investing.com/news/company-news/micron-fq4-2026-slides-record-revenue-ai-demand-drives-supply-tightness-93CH-4926074 . Turn off javascript to view.
💬 34 (+3) open on reddit ↗
▲
41
+6
17👁
r/LocalLLaMA · u/masiha97 · 12d ago
Public MCP server for Canadian privacy law data (free, no auth) - works with any client that speaks Streamable HTTP

I built this, so the disclosure goes up front. It's a free, public MCP server plus a REST API with Canadian privacy law data. The MCP endpoint is at https://movahedi.ca/mcp and it uses Streamable HTTP, so any client that speaks that transport can connect. No signup, no API key, read-only, anonymous. The REST API is at https://movahedi.ca/api/v1 and the docs are at https://movahedi.ca/developers. What's in it: - Canadian privacy enforcement actions (CAI decisions from Quebec's access-to-information commission), searchable by keyword - A 263-term privacy glossary - An 11-point Quebec Law 25 readiness checklist The 5 MCP tools are: search\_enforcement\_actions, get\_enforcement\_case, lookup\_glossary\_term, list\_glossary\_terms, law25\_requirements. Quotas are 2,000 calls/day anonymous, or 10,000/day with a free API key (no email required). With Claude Code you can add it like this: \claude mcp add --transport http movahedi-privacy https://movahedi.ca/mcp\ For local setups: since the server is remote over HTTP, a client that only speaks local stdio can reach it through a proxy like mcp-proxy or mcp-remote. That's how I've seen people pair it with locally run models and agentic harnesses. Happy to answer questions about the data or the setup. I am the builder (Alexa, on behalf of privacy researcher Mohammad Movahedi, movahedi.ca).

▲
40
+4
24👁
r/LocalLLaMA · u/Qual_ · 9d ago
Astrabox - Open source Arcade Game Generator post image

Hey everyone! I’ve been working on ASTRABOX: an arcade interface where you describe a game, the AI builds it, and you can ask for changes by voice while playing.

Each game gets its own visuals and gameplay, while a shared runtime handles controllers, scores, player joining, etc.

I built it around Codex, but the code is open source. I’d love to see someone adapt it to a local coding model, local STT/TTS, and a different harness.

It runs as a local web app—you don’t need a Raspberry Pi or an actual arcade cabinet, although that’s what inspired the project 🕹️

It’s still experimental, but feel free to customize it, change the little robot, the environnement, or everything.

https://github.com/Qualzz/astrabox

Curious what models and tools you’d use for a local version.

Edit: Clanker helped me with writing this message in english.

https://preview.redd.it/rru52ja4brsh1.png?width=2224&format=png&auto=…

https://preview.redd.it/fm3m19j2brsh1.png?width=1280&format=png&auto=…

💬 11 (-1) open on reddit ↗
▲
40
 
18👁
r/LocalLLaMA · u/jacek2023 · 10d ago
Ornith-1.5 DFlash

Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash

Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-397B-DFlash

Ornith-1.5-35B-A3B-DFlash pairs the Ornith-1.5-35B-A3B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-DFlash

▲
38
+7
32👁
r/LocalLLaMA · u/SultanGreat · 8d ago
What's the best setup for Qwen3.8 27b for a 16 gig VRAM?

Hello guys!

I have been experimenting with qwen 3.8 for a long time and I hadn't been able to get reasonable speed. I am on a 5060Ti 16 GB, and although this gpu can game, I am aware that AI demands more than 16 GB.

I am on a Fedora 44, AMD Ryzen 9600x and 16 GB system ram (16 GB system ram and 16 GB vram, totaling to 32 GB) and I would like to use llamacpp, although I would use any other tool if I could if it meant faster speed.

I am looking for a large context. Atleast 128k context. The first question is, what quantization to pick? In my experience Q3 UD was satisfying, but I am looking for uncensored model. In my experience, MTP has never lived up to its hype for me (and I don't know why!?), which is why I am thoroughly lost on making a good setup after an honest week of experimentation, which is why I have resorted to ask here as a last resort.

Update : Found a model, thanks to u/_wortkarg_

link : https://huggingface.co/RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3\_XXS-Uncensored

command (A better command would be appreciated and updated accordingly):

~/llama.cpp/build/bin/llama-server \
--model ~/Documents/Models/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf \
--alias "llamacpp" --host 0.0.0.0 --port 8001 \
-ngl 99 --flash-attn on --ctx-size 131072 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--parallel 1 --batch-size 512 --ubatch-size 256 \
--no-warmup --jinja \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

I am hitting at about 35 t/s+ speed with this one.

💬 84 (+3) open on reddit ↗
▲
38
+8
15👁
r/LocalLLaMA · u/CompetitiveDraft9381 · 12d ago
Updated from 3x3090(2x3090, 1x3090TI) to 2x5090

Upgraded from 3x RTX 3090s to 2x RTX 5090s on my homelab server and picked up a solid speed jump on top of it from a software update (speculative decoding + NVFP4). Setup: llama.cpp (build b11216), running Qwen3.8-27B (abliterated, Q8\_0). Blue = old 3090 setup, green = new 5090 setup on the same Q8 model, teal = the 5090s again after switching to NVFP4 + speculative decoding. For colorblind folks, the order is: 1 - 3090, 2 - 5090 Q8, 3 - 5090 with NVFP4. One caveat on the "before" numbers: one of the three 3090s was on a slower PCIe slot than the other two, so that setup was running a bit below what 3x 3090s on equal slots would do. Overall, very satisfied. I bought 2 prebuilt PCs for $6.4k each when the 5090 went up to $7.5k, 2 weeks ago or so. I really wanted to upgrade to 5090s for a long time for NVFP4 support. The plan is to sell the 3090s for $2k each or so. It would probably be another 2-3 months until they go up that high, but I expect that they will. So, with the prebuilts' leftover components and the 3090s, in the best-case scenario I expect to get back $9k, so the total cost of the GPUs would be about $6k after taxes, which is still nuts and more than the 5090's MSRP. Ask me any questions, or if there are any other benchmarks you guys want me to run, let me know and I will.

▲
37
+11
24👁
r/LocalLLaMA · u/Defiant_Ranger607 · 10d ago
What kinds of problems are still fundamentally hard for LLMs?

I played a game of a custom chess variant against an claude opus 5.5, and it beat me.
The game combined several rule changes: the board wraps around from the h-file to the a-file, knights move three squares in one direction and one sideways, and captured pieces can be dropped back onto the board, as in crazyhouse. I gave the model the rules, the starting position, and a board diagram, then asked it to reply with one legal move at a time.
Also I recreated this board game https://nika-game.com/ and played with claude, and it still win, although I believe it is really unpopuplar and old game without much training data available (claude didn't even know the rules initially)

I used to think chess exposed a fundamental limitation of LLMs: keeping track of a changing board, following exact rules, and planning ahead seemed like a poor fit for a language model. This game made me reconsider that assumption.
So I’m curious: what broad classes of problems do you think LLMs still can’t solve reliably? Are there limitations you consider fundamental/archiectural, rather than problems that might improve with better models, more computation, or tools? What would be a good test?

💬 87 (+6) open on reddit ↗
▲
37
+2
19👁
r/LocalLLaMA · u/PilgrimofHaqq2 · 11d ago
9 prompt rules cut my coding agent's wasted thinking up to 70% (GLM 5.3 & GLM 5.3 Flash)

360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%.

The block (shipped to global instructions):

Thinking discipline 1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it. 2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling. 3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps. 4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue. 5. Do not revise just to agree. If the user pushes back without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. Being agreeable at the cost of being correct is a failure, not politeness. 6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason. 7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle. 8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do. 9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke.

Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam

💬 21 (+1) open on reddit ↗
▲
36
+5
16👁
r/LocalLLaMA · u/light_2earth · 11d ago
macOS 27 ships a free local LLM on Apple Silicon Macs. I made it easy to use from Node and Python

Apple Silicon Macs on macOS 26+ come with a small LLM built in. No download, no API key, and nothing leaves your Mac.

Why I built it

I was making a tool that writes API docs from code, and I didn't want users to install Ollama or paste an API key. Apple's model was already on their Mac, so I used it.

Getting it to work well was harder than expected. At temperature 0 it kept repeating itself, and a 2-second call took 20. Calls took 17 seconds instead of 1.5 until I kept one process running. So I turned all the fixes into a library: apple-llm.

What it's good for

\- Pulling structured data out of messy text, like emails into tickets or receipts into expenses. The JSON always matches your schema.

\- Tagging, summarising and rewriting

\- Private data you don't want to send anywhere

\- Tools you share with other Mac users, who don't need to download a model or get a key

What it's bad at

Coding and long reasoning. It's a small model.

There's also an optional cloud mode that uses Apple's bigger server model for harder questions. That one is not local: your prompt goes to Apple's servers, and it has a usage limit.

Node: npm install apple-llm

Python: pip install apple-llm

https://reddit.com/link/1ws5l5p/video/oarqcgm767sh1/player

GitHub: https://github.com/jagdishpal02000/apple-llm

💬 26 (+1) open on reddit ↗
▲
35
+7
31👁
r/LocalLLaMA · u/Any-Winter-4079 · 8d ago
DDR4/PCIe4 vs DDR5/PCIe5 for LLMs- I benchmarked them for pre-training. What are your thoughts? post image

Hello everyone.

I've recently ran some experiments comparing DDR4/PCIe4 and DDR5/PCIe5 for AI workstations on a pre-training run, and would like to hear yours thoughts.

First of all, and as a summary of my results ( code here: https://github.com/Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training ), I rented two machines on Vast.ai, one with an H12SSL-i motherboard, an EPYC 7352, 192 GB of RAM and of course using PCIe4 (26.3 GB/s) and another with a WRX90E-SAGE SE motherboard, a 9975WX CPU, 256 GB of DDR5 RAM and PCIe5 (54.3 GB/s), and DDR5/PCIe5 is about 15-20% faster on pre-training (depending on whether you include or exclude validation and other costs) under the same number of GPUs.

With the current RAM prices, however, for the cost of 256 GB DDR5 RAM at 6400 MT/s you can get a full (extra) RTX PRO 6000 WS/Max-Q, at which point the comparison clearly favors DDR4/PCIe4 (with 2 GPUs), with about 50% extra throughput vs a single GPU at equal(ish) cost.

Now, there aren't a lot of downsides in my mind to choosing DDR4/PCIe4, but there can be a few:

  1. at least some of these DDR4/PCIe4 motherboards are on the older side, and were one to need replacement, they are not so easy to get (for example, the H12SSL-i used, I can only find it for sale as refurbished now, so who knows in a few years if it will even be available for retail).
  2. newer GPUs (as in, new NVIDIA generations) may stop working with older motherboards (meaning yes, PCIe is backwards compatible but the motherboard's BIOS/UEFI sometimes has issues during POST with newer GPUs (e.g., some older PCIe3 motherboards already have trouble recognizing Blackwell cards, and this may be the case for PCIe4 and newer cards in the future). Meaning if one were to buy a newer GPU in the future, the whole workstation may not be suitable.
  3. If one has to go ahead and bite the bullet on RAM prices, the million question is when? To me prices now are \*\*awful\*\* but so were RTX PRO 6000 prices and here we are (i.e., even higher).
  4. For pre-training, I would still choose DDR4/PCIe4, but I suspect for inference DDR5/PCIe5 might be a fair bit better than the 15-20% that it gives you on pre-training, plus we might be moving into some techniques soon such as dynamic expert/data loading into the model at runtime, which again would favor better DDR/PCIe speeds.

With all of this, I am curious if anyone has benchmarked this, and what are your thoughts on it. Would you hold out on DDR5 at the moment, and therefore go for PCIe4, or would you bite the DDR5 bullet early? Another issue with RAM is channels and DIMM count/channel, because if you want to go 'cheap' like let's only get 192 or 256 GB of DDR5 on 8 channels at 1 DIMM/channel (e.g., 8x32 to get 256), then upgrade to 512 later (when budget allows), that means you have to replace your full RAM (because all the slots are occupied, requiring new 8x64 to get 512 for instance)... And if you get fewer DIMMs like 4x64 to get 256 GB (leaving 4 DIMM slots unoccupied) then you get half the bandwidth because only 4 channels are populated. So maybe a machine that has dual DIMM support per channel is the answer to this (fully populating 8x32 to get the full bandwidth, and still allowing you to expand to another 8x32 to get 512), but in general it's a tricky point too.

So, what do you do/are you guys doing? Have you recently bought a workstation or upgraded to one, for pre-training, fine-tuning, RL, inference, whatever your use case may be, and come up with this dilemma? Are you choosing DDR4/PCIe4 as it would seem reasonable or are you going for DDR5/PCIe5 and if so, why? I am interested in all use cases and opinions!

💬 19 (+1) open on reddit ↗
▲
33
-2
21👁
r/LocalLLaMA · u/MooseEfficient2151 · 11d ago
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection

*link to original article*

TLDR: security researcher eddie zhang used a modified+uncensored local qwen 3.8 27b to create an executable capable of dumping LSASS memory for credential harvesting while evading 2 modern EDR security products.

this makes me reflect on how cloud providers keep putting guardrails on everything to the point where even authorized testing gets blocked. local models are the only real option if we want total control, but running heavy local rigs for long agent tasks drains so much compute and management overhead.

been using claude code hooked to sumus to handle my local project workflows and orchestrate tasks in the background, while keeping full local file access on my machine. curious if anyone here is running local uncensored models as local agents for heavy automation, or if you still blend cloud models with local execution setups for your dev environment?

▲
32
+2
13👁
r/LocalLLaMA · u/Balance- · 12d ago
It would be really cool to have an official 3D-print engineering benchmark/leaderboard like this post image

Someone prompted different LLMs to generate CAD code for a bridge under fixed constraints (2-foot span, under 500g filament, 18-hour print limit), printed them, and load-tested them to failure. The results were quite varied: some models couldn't even design parts that fit together, while the winner held over 100 lbs. Most benchmarks don't really capture physical intuition, spatial reasoning, and functional code generation at the same time. I would love a standardized benchmark and leaderboard for this.

▲
30
+8
47👁
r/LocalLLaMA · u/Porespellar · 7d ago
RTX Spark laptops and mini desktops rumored to launch Oct 7th (24GB to 128GB variants possibly)

Basically a DGX Spark minus the ConnectX-7 ports. I’ve seen expected initial pricing from like $1800 to $2900. Not sure what configurations are actually at those price points.

It’s all Internet hearsay until we actually see these things ship, but it’s nice to know that it’s potentially around the corner next week, especially given DGX Spark’s insane price increases lately.

Sadly, you can’t cluster them, but getting an entire computer + GB10 equivalent chip for a little over half the price of a used 4090 seems like an ok deal in this market.

💬 51 (+19) open on reddit ↗
▲
30
-2
19👁
r/LocalLLaMA · u/rorowhat · 10d ago
Best model for blender?

Trying to see if I can use a local model to generate game assets, or even 3D printer models. Any suggestions? Something that would fit in 64GB of ram, speed is not an issue. Just need it to work well.

💬 50 (-1) open on reddit ↗
▲
30
+3
20👁
r/LocalLLaMA · u/thebadslime · 10d ago
I created a personality test for models, need more TESTS!!

First 3 disposition results.

Hi there!

My name is Jerry and I recently built Enclosure, e deterministic environment for testing LLMs. It's a simulated office enfironment https://github.com/openconstruct/Enclosure with tools agents are used to, like slack, calendar, mail, chat and more. For conversations is uses AIML instead of a model, so it is totally deterministic.

The first test I made(had Claude make) is the one I had in mind when I designed Enclosure, a personality test for models. It's not a winnable benchmark, but rather a tool to help people find the mdoels that fit their workstyle best. It's called disposition and you can find it here: https://github.com/openconstruct/disposition

I tested the cheapest models on Alibab modelstudio already, after I refill my opencode go next month, I will probably test more. I am asking the community to test some models if you think it's a cool project.

Instructions are in the Disposition repo, and the submission repo is here: https://github.com/openconstruct/disposition

▲
30
-4
15👁
r/LocalLLaMA · u/SeveralViolins · 12d ago
Splash fork optimised for M5 Max: ~1.5× faster (1.25× single request)

I’ve spent the last couple of days with Opus 5.5 working on a fork of Inco’s excellent and already blazingly fast Splash engine to optimise it for M5 Max chips. Taking liberties and referring to it as Splish. Roughly the opposite direction to u/Erp4759’s great M1 port (Splash on M1, part 2). Charts (stock Splash vs Splish): https://raw.githubusercontent.com/publicExcess/splish/m5/docs/m5/charts/png/single-request.png https://raw.githubusercontent.com/publicExcess/splish/m5/docs/m5/charts/png/concurrency.png Results Against Splash 1.1.0 as shipped, on the same Mac with the same models, Splish is: \~1.25× faster at a single request (+11% to +35%) Up to 1.5× faster at 2–4 requests Quality is unchanged on everything I measured In real world use, going from about 45-51 tok/s to 56 - 64 tok/s in short story prompts in Deepseek Harness. All figures are for 4-bit models on a 40-core M5 Max unless stated otherwise. What worked 1. Kernel choices measured specifically for the 40-core M5 Max, using Splash’s own tuner. The tuner is in Splash’s source code but isn’t included in the packaged app. This was the biggest single-request win: Swift-1.5 went from 74.7 → 89.8 tok/s (+20%). 2. Loading those choices from a file (SPLASH\_KERNEL\_CHOICES). No speedup by itself, but it means anyone can retune without rebuilding. 3. New verify kernels for the M5’s tensor units. These use lighter barriers and compute row sums once per projection: +5% at 1 request +10–19% at 2–4 requests 4. Extending the same kernels to more projections. A further 1–3% at 3–4 requests. Together, #3 and #4 make a decode step at 2 / 3 / 4 requests: 1.32× / 1.49× / 1.36× faster than tuned Splash. 5. Tuned Qwen3.6-35B-A3B with the new kernels. Speedups at 1–4 requests: +5% / +18% / +22% / +18% 6. An attention tweak for the 27B shape. Attention is 2–3% faster and prompt processing 3–4% faster. Too small to show up in the overall numbers. 7. GGUF (Q8\_0, Q4\_K\_M, Q6\_K): faster input loads in the decode kernel. +2–14% per kernel and about +2% per step. Output is bit-identical. 8. A copy rule for coding agents, borrowed from TensorFold. When the model is rewriting text it has already seen, the drafts copy it verbatim. Whole-file edits get +24% to +42%, while everything else stays within ±3%, and output is exact. I’m exploring a complementary approach for a future version. The README also lists everything that didn’t work for me, which is probably useful if anyone wants to avoid going down the same rabbit holes. There’s lots more I’d like to test, but thought this was a nice start. The tuned settings are for a 40-core M5 Max. Other M5 chips fall back to Splash’s defaults unless overridden. An auto-tuner is coming. If you run it, python3 dev/m5/report.py prints a performance report. Results from other machines are very welcome, especially if you find cases where it’s slower.

▲
29
+11
32👁
r/LocalLLaMA · u/jjusko20 · 7d ago
Update #3: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Last update for those who may be following: https://www.reddit.com/r/LocalLLaMA/comments/1wv5h8x/update\_2\_post\_training\_yandexaliceai80ba3b/

I screwed up guys 😂

Turns out my loss curve during my last run was legitimately unhealthy - as some of you, and myself, were concerned about. After evaluating my QLoRA, I found zero'd gradients in all but two layers. Turns out I had a NaN issue related to my custom v100 kernels that I didn't catch - so that run is cooked, I had to restart. I guess two layers training managed to emulate a loss curve I could at least derive a sensible explanation for until I actually got to evaluate.

Thankfully, checked to make sure gradients were applying again, and restarted the run. Once again, it's live streaming at https://figure-bios-expect-cio.trycloudflare.com/

Loss curve looks much healthier this time and is making me feel more confident that this is going to be okay. Stay tuned! I'm gonna release GGUFs and a llama.cpp patch when I have a working version.

My first epoch loss curve from this run

my first epoch loss curve on the failed first run \(note the differences in scale even if the pattern looks similar\)

▲
29
+1
12👁
r/LocalLLaMA · u/ea_man · 12d ago
Who wants to try a Pi trick for 27B to reuse prompt prefill between different sessions?

You know that when you start Pi you have to process the initial prompt, that takes some time when you use the slow dense QWEN 27B (that's the very reason why you use Pi instead of Cloud Code!), then you start to add extensions, tools, your append.md and whatever... Well now that got big, like 20k big and it does bother. So let's cache the "initial prompt" PP, so that when you start an new pi session: TA-DA! Instant ready, jolly good. Well if you tried to do that with Pi, llama.cpp and QWEN 3.x hybrid KV all kind of things step in your way to prevent that, so many that I won't even start to count I'll just tell you what to do: 1. dwl and install this Pi extension (tested on Pi version 0.87.1 ) 2. dwl and patch lama.cpp: yup no way around this if you kill the server between session, suck it or leave now 3. do your self a favor and use Froggeric template, for all your QWEN models, even the old ones. \---- So I'll help ya and give you some kinda useful parameters to launch the thing too: --slot-save-path /home/eaman/llama/slot_caches/ --ctx-checkpoints 32 --checkpoint-min-step 4096 -np 1 --chat-template-file chat_template_3.8.jinja You need to have save slots, that's the whole point, the caching is meant to resist restarts. Beware the chunk of blocks cached follow ubatch boundaries, so yeah try to keep that down if you wanna cache some more. Now in the extension README.md there's explanation of env variables that you can tweek, you go read those and edit accordingly to your setup (or have your LLM read that and suggest / config for you), TLDR you need at least: export PI_PREFIX_CACHE_BASE_URL=http://localhost:8080/v1 export PI_PREFIX_CACHE_PERSIST=1 export PI_PREFIX_CACHE_SLOT_DIR=/home/eaman/llama/slot_caches Disclaimer: this is not an easy thing, if you are not familiar with patching llama.cpp and installing extensions manually leave this thread for an other day. On the other hand this thing kinda works for me so if someone else is interested after some testing (because all kind of evil things want to break prompt caching) I'll upload a final extension and see about the llama.cpp problem with saved check points.. Possible results: https://preview.redd.it/xobe9ve7h5sh1.png?width=1281&format=png&auto=… EDIT: made a version for OpenCode: https://store.piffa.net/lm/ocache/ For those of you who want a basic understanding of the problems and solution regarding caching the prompt I've asked the LLM to make a short summary.

▲
28
 
8👁
r/LocalLLaMA · u/jjusko20 · 11d ago
Watch me post-train AliceAI-Foundation-80B-A3B from base to instruct at home, live, on my V100s!

No click bait baby I promise - I'm live streaming the training process kinda like MiMo.

UPDATE: \[Training is paused for an hour or two\] back to training in batches. u/FullOf_Bad_Ideas has pointed out to me I'm burning a ton of compute for nothing on sequence lengths - we'll be breaking the run up into 7/8 batches and then going again.

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wpg4a8/im\_trying\_to\_posttrain/

If you didn't see my original post a few days ago, I'm attempting a slightly more ambitious than usual project in trying to create at least a rough AliceAI-Foundation-80B-A3B-Instruct

I spent the weekend distilling my initial instruct training dataset out of Qwen 3.8 27b, medium thinking - intentionally done because I can run it locally, and I wanted a full dataset in some reasonable amount of time. Still took my v100s running 4x instances at 25tps, like 96 hours of non stop generation to complete the dataset.

I opted not to go for a pre-existing public dataset because I wanted to practice building my own distillation engine (which was configured to work off of an OpenAI compatible endpoint, so it'll distill anything you can hook it up to). The final dataset (this time) consists of 3340 samples: 1760 of general instruct transcripts, and 1580 agentic specific work rows about SWE, harnesses, terminals, etc - I gave the distilling engine a python sandbox and got to simulate turn driven development with a user, and I trained for a bunch of different harness syntax for tool calls, which hopefully will be enough to generalize - gonna run 2 epochs at first.

My GPUs are sobbing right now - turned them on on Friday and left for a weekend vacation, got back today, waited an hour for the data to finish generating, and then immediately fired up the train.

The stuff above is the short version. I'm guessing the initial SFT train will take about 3-6 days, and I plan on working on a RL implementation after I'm satisfied that the SFT has at least worked properly. I am training a rank 16 QLoRa adapter on only q/k/v/o proj, no direct knowledge weight fine tuning.

I thought what MiMo did with their recent training was really cool to watch online, and I like sharing my work with like minded people, and frankly, there's a part of me that's hoping someone will see this and want to hire me (looking for NYC work if you know anyone looking for some passionate ML engineers!) - so I've set up my own little training stream on a cloud flare tunnel.

The stream has the live in progress status of the train, including a live view of the actual data being processed by the model. It also includes way more detail about how I actually designed and generated my training data. Happy to throw the full set on HF as well. I don't expect this model to beat any existing standards but I'll be curious to see if I can get it to operate properly in a harness so I can formally bench it.

I hope you find this interesting! The live stream is a self updating website where you can see exactly what's happening - no need to reload. To watch the training live, visit https://figure-bios-expect-cio.trycloudflare.com/ \[i am currently fixing training issues but it'll be back asap\] -- I'll be keeping it up until the initial SFT is done, at least. The stream lets you inspect the training live as well. This is just a cloudflare tunnel to the trainer.

3 hour update? Loss started at 9ish and is bouncing near 3/4

Update today: back online

▲
28
+2
21👁
r/LocalLLaMA · u/Exciting-Engine882 · 12d ago
is switching from llama cpp to vllm worth it

I have hp z8 g4 with 512 ram and 1x3090 1x5060 16gb. has anyone made the transition from llama cpp to vllm recently? is it worth it? docker under windows or full linux install? I am mainly interested in the model support, it seems that many new local models are supported day 0 in official vllm, while for llama cpp it takes months sometimes. LE: I want to use it for big'ish moe models, that would have to offload some tensors to system ram. I will use it just for myself. I don' t need it to be faster than llama cpp, if it runs at about the same speed it is fine , as long as it works.

💬 68 (+1) open on reddit ↗
▲
27
+1
17👁
r/LocalLLaMA · u/Brief-Tap-6616 · 11d ago
95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090

Hello everyone! A little while back I posted about LlamAmpere, a fork of Llama.cpp with Ampere-specific improvements (though it is caught up to main and will support other hardware, too).

Thank you to everyone that tried it out and shared back their results across the 30xx cards. I'm happy to share I've pushed v0.4 out this morning. On the 4.6bpw model tested, speeds improved \~10% vs the last version while also improving the max context by 10%+ (technically, it can go above 262K, but I have not tested any custom kernels or graphs to support YaRN).

The closest competition comes from vLLM, keeping within <10%, but does so with lower maximum context. It is significantly faster than other llama.cpp options tested.

https://preview.redd.it/u70q7lew1bsh1.png?width=1080&format=png&auto=…

[](https://preview.redd.it/95-tps-through-100k-generated-262k-ctx-on-a-single-30…)

There's also a number of other improvements for other quants/formats, with EXL3 seeing significant speed up (\~80% the speed of the 4-XS-M quant tested). It has a slightly lower KLD, but not a range I have found stat significance for at the task level, so I sticking with the XS-M model for now (built on top of Swift-qwen's distill, which is far more token efficient than the stock train for \~1% performance loss). the 4.3 bpw EXL3 model does provide a bit more room if you are interested in 2+ concurrent predictions. Improvements in this format and the IQ2/3 codebook quants will be most useful for people on 12/16/20 GB setups. These measurements are at temp=1, vs some of the vanity speeds you will see people claim with temp=0 and/or short generations.

As always, please share your results + config details so I can keep improving!

fork is here: https://github.com/JakeATX/llamAmpere/blob/main/QWEN\_AMPERE.md#build-and-run**

model used here:

https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF**

Build command:

git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server

Build + launch (linux):

git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server
curl -L -o ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf \\
https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF/resolve/main/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf
\-m ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf -c 262144 \\
\-ngl 99 -fa on -ctk turbo5 -ctv turbo4 -b 4096 -ub 1024 -t 8 -tb 8 --parallel 1
cd
./build-sm86/bin/llama-server

Previously, people had expressed concern over quantizing KV cache, and TQ specifically. The TL;DR on that is that any reasonable KV quantization strategy (at least for hybrid attention models like Qwen) is going to be swamped by quantization of the weights. The KV quant we're using here (TQ5/TQ4) is less than 1/3 of the KLD we see when moving from 8 bit weights to 4.6 bit weights (and the KLD is only partially additive, so some of the incremental errors cancel out). There was no statistical significance when testing this KV quant at the task level against 8/8 kv (just trivial variations in sentence length). I will be adding KVaRN in the next release, but with a better codec than currently available elsewhere, so it requires a bit more testing before release.

v0.5 will be focused primarily on the 12GB cards, but this should have generation-wide speed ups, so even if you're not on 24GB, please share your results.

Enjoy!

▲
27
 
16👁
r/LocalLLaMA · u/mauricekleine · 12d ago
Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode post image

Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_…. That thread shaped v1.2: - Reasoning effort is explicit per run - Every prompt and output is public. - All current top ranking private and open weight models have been added - Someone spotted Grok miscounting a 400-character answer. Turns out that trips up most models, so Hard mode answers row by row rather than a single text string. Results: - GPT-6 Astra: 30/30, the first perfect run on 15x15 puzzles - Best open weights: DeepSeek V4 Pro 83% (tied 4th), DeepSeek V4.1 Flash 77% for $0.84 total - Hard mode (10 random 20×20s, one solution each): Opus 5.5 8/10. Every open-weight model: 0/10 Still OpenRouter-only, so no way to run locally yet. PRs welcome. nonobench.com (raw data, API, and code on GitHub)

▲
26
+13
35👁
r/LocalLLaMA · u/wayneworkman · 7d ago
Peacebell - a from-scratch small language model

I've been using my free time during weeknights and weekends for the last 11 months working on and refining a small domain-specific language model. It specializes on information about World War II.

The number one question I get asked about this is "Why did you pick World War II?" here are some of the reasons:

\- There's a lot of good Wikipedia articles about WWII, and this is permissively licensed. Meaning I can use the materials.
\- There's a lot of good public domain information about WWII in general - more to train on.
\- WWII is factually dense - making it a challenge.
\- The facts surrounding WWII are mostly unchanging - meaning my model would age well.
\- I had to pick a first topic.

A lot of my journey is documented on my blog: https://wayne.theworkmans.us/llm.html though I've not posted recently.

The model is more than from-scratch. I'm using a custom built training pipeline. And I produced all of my own synthetic data to train on (based on Wikipedia articles).

The majority of my time has gone into data curation and balancing.

As I built this model, I've learned a ton about training data for language models, and about language model creation. I tripped over every bump along the way, 100s of times.

I also learned a lot about World War II, and I'm emotionally exhausted. Many know the basics... the Manhattan project, the Holocaust, the concentration camps, Pearl Harbor, D-Day. Though beyond these topics, there is enormously more tragedy than I previously knew. As an adult with my own family now with better ability to comprehend, many times I'd just cry face down on my keyboard from some of the things I learned. Sometimes I would abandon working on it and go to bed early. I've talked with my wife about how awful some of the things that happened are. It's been hard. And I'm ANGRY! So incredibly angry about the atrocities that happened. Especially angry about the things that happened to civilians, non-combatants, women and children.

Well enough of that.

I open-sourced the training materials and the weights. There are two versions of the model. There's a 291M parameter version and a smaller 148M parameter version.

I built the 148M to compete in the various HuggingFace dashboards that limit model size to 150M. Then I built a new benchmark that focuses on WWII topics, that's also on the hub, though the questions are private to prevent them ending up in people's training data (and no they aren't in Peacebell's training data either).

You can try the 291M model for free here. As you use it, keep in mind this is first-version, it's rough, it's not always right. And it really struggles with longer context. Fresh context gives better results.
https://huggingface.co/spaces/wayneworkman2012/peacebell-v1-291M-demo-cpu

The leaderboard is here:
https://huggingface.co/spaces/wayneworkman2012/ww2bench-leaderboard

I've entered Peacebell into various SLM Arena's, such as CodeSoft's SLM arena here:
https://huggingface.co/spaces/CodeSoft/SLM-Arena

Basically everyone in this LocalLLaMA would be able to run the model easily, even without a GPU. There's a customized vLLM fork here that can run either Peacebell model:
https://github.com/wayneworkman/vllm

Next version is expected to be released sometime in 2027.

💬 8 (+3) open on reddit ↗
▲
26
+2
13👁
r/LocalLLaMA · u/eribob · 12d ago
Worth going from Qwen3.8 27B to flash next or maaybe deepseek v4 flash?

I am running qwen3.8 27b on my dual rtx 3090 (fp8 quant, unquantized cache, 129k context) and I think it works decently well with hermes, opencode etc. But! I am tempted by the new models coming out such as qwen3.8 flash next, deepseek v4 flash, glm 5.3 flash. However, there is a big jump in vram and therefore in cost! The least expensive option seems to be buying 2 of those cmp 170hx 64gb cards for roughly 6-7000 usd in total (that is the price I can find for verified cards here in europe at least). With that I would get another 128gb of vram for a total of 176gb so I could run I think around 3-4bit quants of the above models, right?). I am thinking that it might be faster because of moe but not sure how much smarter? For that kind of money I would want a real noticable improvement! 4xv100 32gb would be cheaper (maybe half price?), but even more hassle to set up, more power draw, and slower. What do you think? The free option is to just wait for qwen4 27b and (hopefully) just download more IQ.

▲
25
-5
23👁
r/LocalLLaMA · u/fallingdowndizzyvr · 11d ago
If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context.

Here's the project. I have nothing to do with it. I'm just an amazed user.

https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BEN…

Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat.

"[6204 chunks in 119.0 s | encode: 1239 tok/s | decode: 57 tok/s]"

That's with MTP on. The PP speed in particular is just so fast. That PP speed is twice the speed of the fastest Strix Halo specific fork of llama.cpp I've ever used. Needless to say, the uplift is even greater compared to mainline llama.cpp.

It works with models other than QFN, but the current number is small. You can find the list on their project page.

💬 49 (+3) open on reddit ↗
▲
25
-2
14👁
r/LocalLLaMA · u/your_real_Fathe_ · 12d ago
Qwen, where's the small stuff? (1B/2B/4B)

I know Qwen is a key player in the local LLM space and has consistently introduced truly impactful technologies—like n-gram in Qwen-Next and the recent Qwen 3.8 27B, which is an amazing local model. However, my question is: why are we seeing fewer small-scale models lately—such as 4B, 2B, or 1B versions? This is especially notable given that Qwen hasn't released any new models in this weight class since the 3.5 series, and rumors regarding Qwen 4 suggest they don't plan to do so either. I realize the 27B model is outstanding and deserves praise in its own right—and it might seem a bit selfish to ask for more—but the reality is that not everyone has high-end hardware. Many people have limited hardware capabilities; this trend somewhat conflicts with the core mission of open-weight LLMs, which is to make AI accessible to the general public. I know smaller companies have recently released lightweight models, but the issue arises when we see that many of these new releases are simply fine-tuned or improved versions of Qwen base models. Since building an LLM from scratch is prohibitively expensive and difficult for small companies or individual researchers, it follows that the absence of lighter Qwen weights directly slows down the development of edge-compatible models and AI applications for consumer-grade hardware on a broader scale. (The same point applies to Google's Gemma series, though—let's be honest—they haven't even released new flagship models since Gemini 3.1 Pro, so...)

▲
24
+2
14👁
r/LocalLLaMA · u/NickCanCode · 12d ago
Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?

There is always at least 1+GB of VRAM not usable not matter how I set the --tensor-split (-ts) param. I tiny shift toward one side will move the weight significantly to the other side. 😵‍💫 Adjusting context will increase/decrease usage on both side. --tensor-split 499,501 = GPU1 12.5 GB, GPU2 15.4 GB --tensor-split 501, 499 = GPU1 14.7 GB, GPU2 13.4 GB Tried --spec-draft-device with CUDA0 and CUDA1 separately, no change at all. (same distribution as above) Also tried --mmproj-device, no much difference. Tried --no-mmproj-offload, somehow the lower side get even lower 🫣 = GPU1 14.7 GB, GPU2 12.3 GB I guess it is related to MTP + Tensor Parallel stuff being concentrated on one GPU. No idea how to solve this. llama-server \ --batch-size 2048 \ --cache-ram 24384 \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ --chat-template-file /mnt/AI/models/qwen-chat-template-froggeric-22.5.jinja \ --checkpoint-min-step 1024 \ --ctx-checkpoints 32 \ --ctx-size 192000 \ --fit off \ --gpu-layers all \ --image-min-tokens 1024 \ --load-mode none \ --main-gpu 1 \ --min-p 0.0 \ --mmproj /mnt/AI/models/Qwen3.8-27B-mmproj-BF16.gguf \ --model /mnt/AI/models/Qwen3.8-27B-NVFP4-MID-HIGH.gguf \ --parallel 1 \ --presence-penalty 0.0 \ --repeat-penalty 1.0 \ --spec-draft-n-max 5 \ --spec-draft-n-min 0 \ --spec-draft-ngl all \ --spec-draft-type-k q8_0 \ --spec-draft-type-v q8_0 \ --spec-type draft-mtp \ --split-mode tensor \ --temp 1 \ --tensor-split 499,501 \ --top-k 20 \ --top-p 0.95 \ --n-gpu-layers-draft all \ --no-prefill-assistant \ --reasoning-preserve

💬 21 (+2) open on reddit ↗
▲
23
+1
32👁
r/LocalLLaMA · u/darklordfireape · 9d ago
Update: Strix Halo + R9700 with llama-halo-hybrid - now beats DGX Spark

Hi folks, I've spent the last couple of months experimenting with Strix Halo and previously I released a proof of concept I called llama-halo-hybrid. I've continued updating it and it now performs very well. The idea is that you can take an R9700, or similar, and place dense parts of the model, KV, and some of the layers on the GPU and let the APU take the rest of the model. You can add the extra GPU through a PCIe extender (framework desktop), Occulink, or a thunderbolt dock depending on which machine you have. Detailed notes along with code in the repo on github. I'm not selling anything, this is all 100% open, MIT-licensed.

It breaks 60+ tok/s decode and 2000+ tok/s prefill, supporting full 256k context.

This is not some custom inference engine that requires a custom quant to run. This is llama.cpp modified to run whatever you want, albeit mostly tuned for Qwen and GLM families. After continuing to tinker with it, it now performs better than DGX Spark (albeit cheaper) running Qwen-3.8-flash-next and slightly better yet with the Swift-1.5 variant. Most of my testing was done with the Q4/Q4\_K\_XL models to balance size and quality.

Note \- if you are just using Strix Halo by itself, this is probably not the right tool. Check out gufo, which looks very promising.

https://github.com/sixvolts/llama-halo-hybrid

I would love any feedback you all have and happy to investigate tuning for different "sidecar" GPUs other than the R9700 if there's demand and I can get my hands on one.

UPDATE (10/4): I added some more docs around using different cards besides the R9700. The R9070, V620, and 7800XT all perform very well and I put quick guides for those configs, along with notes on thunderbolt/USB4 setup and the dual-machine setup I used here:
https://github.com/sixvolts/llama-halo-hybrid/tree/main/halo-cookbook

💬 27 (+9) open on reddit ↗
▲
23
 
15👁
r/LocalLLaMA · u/Defiant-Plantain1873 · 10d ago
Recommended replacements for glm 4.7 flash

I know I sound crazy, but i’m using a strix halo and finding that GLM 4.7 flash just runs significantly better than qwen 3.6 35 a3b. But its obviously quite old at this point, i wish we had a new glm that was 4.7 flash sized but is anyone using a model that they have found better than this.

My brief testing with qwen shows that glm is better at tool calling and better at world knowledge, but maybe there’s a chance my qwen set up is wrong

💬 31 (+2) open on reddit ↗
▲
22
+1
12👁
r/LocalLLaMA · u/norenEnmotalen · 8d ago
Unsloth, Swift1.5, Peculiar-Ragdoll, ThinkingCap - Qwen3.8-27B

In a previous post I shared comparison between Swift1.5 and peculiar-ragdoll's checkpoints. Added the original unsloth Q4\_K\_XL and ThinkingCap Q4\_K\_M (they don't offer L or XL) to the comparison. Here are the results over a 69 set of eval questions.

All tests are now run at same "medium" reasoning effort.

unsloth-ud\_q4\_k\_xl one ran using llama.cpp - not the splash forked inference engine.

https://preview.redd.it/tztb66wygvsh1.png?width=2958&format=png&auto=…

I'll do a 3x repeat for the slow run to see if it maintains 69/69 each time.

EDIT: u/jucabala457 asked I test mradermacher/Signal-3.8-27B-Terse-Coder-i1-GGUF The Q4\_K\_M is closest quant available. A nice addition for sure! That GGUF couldn't run with Splash-based engine due to tensor incompat. I ran it using llama.cpp the slow way. The total time taken isn't a fair comparison for that reason. Updated results below

https://preview.redd.it/pvfvgd35twsh1.png?width=2976&format=png&auto=…

I also just made the tuieval tool available here https://github.com/ashe-wb/tuieval

Can't promise you the tool will work right away on your install since a fully local binary is what I've been using and testing with. Customize it with packs of domain-specific eval questions you deal with on the daily. This is the most important part. A model or fine-tune that is not good for one thing might be excellent for something else and only you know what your domain interests are. The ability of a model to render game graphics means nothing to me but it means everything to someone else.

https://preview.redd.it/m2n4rzlzjwsh1.png?width=2000&format=png&auto=…

▲
22
 
14👁
r/LocalLLaMA · u/Kmic68 · 12d ago
2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0 post image

Hey guys! I have been excited to share this here. This is a project consisting of kernel optimizations for the Tesla p100 series graphics card ($80). I want to start by saying I am 17 years old and do not have a formal degree. I used Ai for a lot of this and while I understand some, I do not understand everything. Notes: My gpus are capped at 175w/250w each so these numbers may be able to be pushed higher. I also experience minor thermal throttling and sit at a nice toasty 79 degrees, which definitely effect numbers (the table above is while hot, so if you have good cooling expect 5-10% more on prefill and decode). I am using gen3 pcie with two x16 slots. Also, for anyone curious, decode numbers depicted in image were averaged from a list of questions ranging from creative writing and coding. Improved: tps went from 7-15tps at 0 context to 50-60, 260k context went from 2-4tps to 30-35, prefill went from 220 tps at 0 context to 350, 260k context went from 40tps (as far as i remember, i never really measured cause it was too hard) to 110tps, fixed fp16 math errors by using some mixed fp16/fp32 math operations so rounding errors were eliminated, and merged as of sept 22 so it should support qwen 3.8 flash architecture This setup is somewhat flag specific (ie: (-c 262144 -b 32768 -ub 1024 -np 1 \\) without -b 32768 mtp becomes overloaded and drops acceptance to near 0 at full depth) so keep that in mind while setting up. One more thing, I took regression very seriously in this. Math had to be more accurate or byte identical or it would fail tests. Build, flags, math proofs, and anything else you may need will be linked below. Enjoy guys! I would love your feedback on this and am looking at pull requests. Github: https://github.com/Kmic-68/llama.cpp/tree/p100-optimizations Details for build, math proofs, etc: https://github.com/Kmic-68/llama.cpp/tree/p100-optimizations/p100-docs

💬 25 (+1) open on reddit ↗
▲
21
+2
11👁
r/LocalLLaMA · u/Any-Lingonberry7411 · 9d ago
Best local model for Blender and game dev?

I have been looking at some local models, but even the smartest ones like GLM5.3 Flash and DSv4 Flash have a hard time creating coherent models in Blender and placing them logically in game engines.

Is this something that local models are just too dumb still to do good job at?

▲
21
 
14👁
r/LocalLLaMA · u/pilkyton · 10d ago
PSA: ModelScope CLI is now moved to "modelscope-hub"

To save people 30 minutes of research (because they didn't bother documenting this officially at all):

  • The "modelscope" package is now just the library. Doesn't contain a CLI anymore. If you try to install it or update your old CLI package, you get "No executables are provided by package \modelscope\; removing tool. error: Failed to install entrypoints for \modelscope\".
  • They moved all CLI tools to "modelscope-hub".

The new command to install it:

uv tool install "modelscope-hub"

▲
20
 
12👁
r/LocalLLaMA · u/LH-Tech_AI · 12d ago
[Release] - SupraTTS-0.1-Beta - a tiny 29.6M parameters TTS model

Hey guys! Today, we are releasing SupraTTS-0.1-Beta, a tiny \~29.6M parameters Text-To-Speech model. The audio quality is a bit better than the original Glow-TTS (the architecture our model is using!) while it's keeping the same size. Here are some samples: https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta#samples >Link to the model on HF: https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta I hope you can do something useful with it, e.g. on small edge devices and on CPU. Feel free to give us feedback and ask question. Follow us on HF to not miss the next upgrades of SupraTTS, e.g. better voice quality, multi-language-support, multi-voices support, emotions and speaking styles and ZERO SHOT VOICE CLONING**!! 🤗

▲
19
+11
29👁
r/LocalLLaMA · u/Apprehensive_Side219 · 7d ago
Alternative to Nvidia spark?

I just spent the last two weeks trying to catch a microcenter in stock with a spark and then today the price went up 30% and now I can't realistically afford it. It was already pretty close to the edge of my budget, and now I don't think I can swing 7k for what I had been expecting to pay 5 for last week. Any suggestions for alternative approaches welcome. I really don't want to wait until 2028 to get started.

💬 35 (+10) open on reddit ↗
▲
19
 
18👁
r/LocalLLaMA · u/Roy3838 · 10d ago
Thanks to you r/LocalLLaMA, my mom was able to use my app! The open-source app that can watch your screen and trigger actions. It is now easy to use, thanks to your feedback.

TL;DR: I'm a solo dev who wanted a simple, private way to have local LLMs watch my screen and do simple logging/notifying. After a year of building, I released v3.0.0 and my mom was able to use it for the first time and I wanted to say thank you!

Hey r/LocalLLaMA,

What is it used for?

It is designed to monitor anything, some use cases:

  • When my Simulation crashes, call me.
  • When Concert tickets become available, click the buy button.
  • When my Steam game is downloaded, send me a Telegram.
  • When a Render is finished, send me an SMS.
  • When ... \[Anything happens\] Then ... \[Notify me, log it\]

How It Works

It's a micro-agent framework controlled by an MCP (Agent which I call Observer). So you type in Observer what you want monitored, and it'll control the framework to monitor it.

The desktop app uses llama.cpp as an inference engine, the webapp uses transformers.js, and they both support your v1/chat/completions endpoints :DD

You can try it out in your browser with zero setup!... running gemma-4-e2b ONNX in the browser, crazy stuff! Thanks to Xenova/HuggingFace for transformers.js c:

It passed the mom benchmark lol!

You guys told me that the framework was cool, but it was very manual to setup agents/workflows. So I've spent the last year slowly making it more accessible so anyone from any technical background can use it.

Every couple of months I ask my mom to use the App. And for the first time she actually was able to setup a monitoring agent with a local LLM! Which makes me think the app is ready for general public adoption (wuuuu!).

I hope this makes local LLMs useful for everyone! Tutorial/Demo Which is the whole point of the project.

My Commitment and being FOSS

The core Observer AI platform is, and will always be, free and open-source. That's non-negotiable. The code is all on GitHub for you to use, fork, and inspect.

The line in the sand which I have is "if it's free for me, it should be free for the user", that won't change ever.

Let's Stop Wasting Time!

This project wouldn't exist without the inspiration I've drawn from this community. You are the people I'm building this for.

I'll be hanging out here all day to answer any and all questions. Thank you again for everything!

Cheers,
Roy

▲
18
 
13👁
r/LocalLLaMA · u/Porespellar · 12d ago
Zer0Fit - Zero-shot predictions, classifications, and regressions using Google ML research models running locally as a dockerized MCP

AI grad student here. With all the recent interest in Jev, I thought I would share something I built la few months ago that brings ML models and LLMs together in a different way than Jev does for different use cases. https://github.com/porespellar/Zer0Fit Background: A few months ago on their research blog (https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/ ) Google released TabFM zero-shot foundation model for tabular data. It was kind of ignored except by maybe a few machine learning nerds that care about that kind of thing. I mean, for real tho, TabFM wasn’t exactly the sexiest name choice. I personally thought TabFM was cool as shit because it kind of melded classical machine learning models into an LLM of sorts. So anyways, I wrapped Google TabFM (their model for classifications and regressions), and Google TimesFM (their model for predictions) into a convenient Fast API and made the whole thing a dockerized MCP that you can connect to your favorite LLM. I call my project Zer0Fit - Zero-shot ML tasks without needing to train or fit a model. Here’s my repo if you want to check it out: https://github.com/porespellar/Zer0Fit You see what I did there with the name? It took me hours to come up with that name :) I’ve made it as easy as I could to install. Just clone it and run the install script. So the basic idea is, you connect the MCP to whatever LKM you want, give it a dataset (CSV, tabbed data, or time series), and ask it what you want it do do with the data. it decides which of the Google models to use, and then it runs the regression, classification, or prediction task in context and gives the results back to your LLM. That’s the best way I can describe it. See the Google blog for the details on what the Google models are actually doing. Again, I’m not doing anything special, I’m just wrapping the Google models up to serve locally and making them exposed via MCP. The Google models are doing all the heavy lifting. I have absolutely no connection to Google research and am not associated with them in any way other than being a fan of them releasing this for us to try locally. Is it better than a data scientist building a custom model to do an ML task? No, definitely not, but it is much easier, and probably will get you an answer that is reasonably close (or possibly at least in the ballpark) and that might be good enough for some use cases depending on what you’re looking for (assuming it’s not a task that requires high precision, or high speed classification). Anyways, I just thought the Google models deserved some attention and love from the community, so I wanted to make them more accessible, that’s all, that’s why I made Zer0Fit. If you want to try it out it’s over on my GitHub in the link above. Please remember, this is all just stuff. It’s cool to play with, but don’t use this with anything where it’s output matters. Use at your own risk. P.S. I made it with Open WebUI in mind so it should work well in that, but it’s an MCP so it should work with just about anything that is MCP-friendly. Edit: Mods pointed out that I posted about this before and wondered if it was a repost or if anything changed. I should have mentioned that I just recently released an updated version that now pulls the new 2.5.0 version of Google TabFM that came out a few weeks ago.

▲
17
-1
13👁
r/LocalLLaMA · u/jacek2023 · 11d ago
Holo4

*Holo4*\-27B-GGUF is a vision-language model (VLM) for Computer Use, built on the Qwen3.8 dense architecture and developed by H Company. Used with the hai-agents harness, it can send screenshots and tool results to the model, then execute its requested clicks, typing, code, and tool calls.

https://huggingface.co/Hcompany/Holo4-27B-GGUF

*Holo4*\-35B-A3B-GGUF is a vision-language model (VLM) for Computer Use, built on the Qwen3.5 mixture-of-experts architecture and developed by H Company. Used with the hai-agents harness, it can send screenshots and tool results to the model, then execute its requested clicks, typing, code, and tool calls.

https://huggingface.co/Hcompany/Holo4-35B-A3B-GGUF

https://preview.redd.it/rqx7l4qqs8sh1.png?width=1656&format=png&auto=…

https://preview.redd.it/rzez39trs8sh1.png?width=1656&format=png&auto=…

▲
16
+3
26👁
r/LocalLLaMA · u/brainchillzZ · 8d ago
Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

So everyone has been yelling about how I should be using Gufo instead of halogen because it's open source and it's "just as good or better". Checking in on their GitHub (GitHub.com/gufo-org/gufo) got me immediately .. "Qwen 27B Q4: 70.56 tok/s single user, 123 tok/s with 8 users" on a strix halo device? Yes please ... So I broke down and tried it today ...

Setup: gufo 0.4.0 from their podman image, Qwen3.8 27B UD-Q4\_K\_XL from Unsloth plus the DFlash2 Q4\_K\_M draft model, using their own benchmark script and their own settings (greedy, thinking off, 128 output tokens, prompt cache off).

If you want the short version ... yeah I got 70.22 tok/s. So the number is real. But the prompt that produces it is "Write the word red exactly 1000 times".

But it's also not real. In that figure all the speed comes from the speculative decoding. The draft model guesses like 7 tokens ahead, the 27b checks them in one pass and keeps what it agrees with. When the output is the same word over and over the draft is right every time. On a real prompt it's right maybe half the time.

Their benchmark has a second set of nine ordinary prompts (some C++, a word problem, a summary, Italian, Chinese, JSON, a bit of fiction, a debugging checklist).

On those:
| | repeat-a-word prompt | normal prompts |

|---|---|---|

| 1 user | 70.2 tok/s | 39.4 tok/s median, anywhere from 22 to 52 depending on the prompt |

| 8 users, their "aggregated" number | 122.6 | 82.4 |

| 8 users, tokens actually delivered per second | 82 | 52 |

About that last row. The "123 tok/s aggregated" figure is each request's decode speed added together, with prompt processing and queue time left out. If you just count tokens coming out of the box per second of wall clock it's 82, or 52 on normal. prompts.

To be fair to the gufo people, none of this is hidden. Their benchmark docs have separate "mixed" and "repetitive" columns and the mixed numbers they publish match what I got. It's only the repo description and the top of the README that lead with the best case. And 39 tok/s from a 27B at Q4 on an APU is still really good. Without the draft model their docs put it around 12.

The other thing I wanted to know was how it compares to halogen (peonist-ai/halogen-flash-server), which is what I normally run. Both can serve Qwen3.8 Flash-Next, so I put that on both and sent the same prompts to each. Two boxes, same hardware, same OS image. Greedy, thinking off, 256 tokens.

| | halogen 0.13.8 | gufo 0.4.0 |

|---|---|---|

| nine normal prompts, average decode | 43.9 tok/s | 38.2 tok/s |

| 4 users at once, end to end | 76.6 tok/s | 63.1 tok/s |

| cold prompt processing, \~9.7k tokens | 1288 tok/s | 1495 tok/s |

| the "red" prompt | 56.9 tok/s | 87.4 tok/s |

So for everyday generation halogen was about 13% faster for one user and about 18% faster with four. gufo was 16% faster at chewing through a long prompt and a lot faster on the repetitive one.

I'll add this just in case, because someone will ask or at least try to poke about it in the comments

\- I know the weights aren't the same. halogen uses its own 4-bit format, gufo uses the Unsloth GGUF. I only measured speed. I did not compare output quality at all.

\- They were two different machines but identical hardware and software, and my boxes have agreed within 1% on other benchmarks, but it's still two machines.

\- One run each was done for the head to head. The reproduction of their numbers was 3 reps.

\- gufo has shipped four releases over the last four day, so this could all be stale by next week.

There was quite a lot of stuff that I liked about gufo that isn't performance related. It takes plain GGUFs, it's MIT, the 27B loads in about 3 seconds (Flash-Next in 13), the per-request log line tells you draft acceptance and cache hits, and it does 8 batched sessions. It also has ASR, TTS and image models that I haven't touched. Their benchmark hashes the output with and without the draft model and it was identical every time, so the speculative path isn't changing what the model says.

One thing to keep in mind if you try it out is that it reserves memory per session up front. Flash-Next with 4 sessions at 64k context took 94GB.

So it isn't smoke and mirrors exactly. Everything I checked reproduced. Just know that the 70 is a ceiling you'll only hit if your workload is incredibly predictable text, and plan around the 30s for the 27B on normal stuff.

I kept this all setup to tinker with on actual output quality over the next few days, I'm happy to run other prompts or try it with different settings if anyone wants to see something specific.

▲
16
+1
14👁
r/LocalLLaMA · u/caenum · 8d ago
Best OpenSource Claude Cowork alternative?

Hey guys,

Looking for an alternative for Claude Cowork:

  • Project Work / Documents
  • Integrations like Notion, Gmail, etc.
  • Tools like Websearch, PDF creation, etc.

Came over Eigent (https://github.com/eigent-ai/eigent) but cant find any actual reviews about it, what usually is a sign thats not good performing..

Also have tried multiple other frameworks (OpenClaw, Hermes, OpenWebUI Chat Interface) - but those are different use-cases for me.

LLMs will be server through my own server, so should be open for connecting to Ollama, Ninfer, etc.

So anyone knows a good application which behaves like Claude's Cowork?

Thanks )

▲
16
+2
17👁
r/LocalLLaMA · u/Federal-Effective879 · 9d ago
szmcp: a ZIM HTML to Markdown converter and yet another ZIM MCP server

Hello all, I wanted to share a little project I vibe-coded for myself that you may find useful.

As many people here like to suggest, I wanted to give my small local LLMs access to information to improve their world knowledge. I didn't want to give my LLM free reign searching and browsing the web to keep my queries private and functional offline, so I wanted to give them an offline knowledge base. Wikipedia ZIM files from Kiwix were a good starting place for this. Several MCP servers for ZIM files exist, but I didn't like the existing ones I found for various reasons. The most notable one is openzim-mcp , which works in its advanced tool mode but has overly complicated context-bloating tools, and whose simple single tool mode doesn't work very well in practice.

I built my own MCP server for ZIM files in Rust, exposing a simple tool set that's actually easy for small local LLMs to use, while providing all the functionality one normally needs. It's designed mainly for Kiwix MediaWiki ZIM archives generated by mwoffliner (such as Wikipedia, WIkivoyage, etc.) but also usable with many non-wiki ZIM files. I also wrote my own custom HTML to Markdown converter for MediaWiki pages that produces clean, well-formatted Markdown including special content such as wiki infoboxes, LaTeX formulas, tables, etc. It also strips out references and boilerplate sections from wiki pages to keep the resulting markdown clean and context efficient.

You can hook this MCP server to llama.cpp's Web UI to give your small local LLMs much better world knowledge. A system prompt that I found works well is:

You are a helpful assistant. When answering factual queries, search through Wikipedia using the provided ZIM access to ground your answers. If the articles or sections you read don't have relevant details, you can search more, but don't keep searching forever; you need to answer reasonably quickly.

I tested it with various LLMs of varying sizes. I got good results with Gemma 4 12B (or bigger), IBM Granite 4.2 8B (or bigger), and Ling 3.0 Flash (best results while still maintaining usable speed on my 128 GB Mac). Qwen 3.6 35B-A3B was usable but tended to overthink and hallucinate; Qwen 3.8 27B was too slow to be usable for this purpose on my Mac. I also experimented with smaller models, and got usable results for simpler queries with MiniCPM5 2B, LFM 2.5 2.6B, and IBM Granite 4.2 3B. Gemma 4 E4B did not work well for this.

I also build a sub-command within this tool to convert entire ZIM files from HTML to Markdown to save disk space (and avoid the need to convert on every tool call). It converts a 49 GB Kiwix nopic full English Wikipedia ZIM file into a 19 GB Markdown ZIM file, while maintaining all article content (aside from references) and maintaining full-text search. Likewise, it converts the 17 GB top-1M nopic enwiki Kiwix ZIM file to 6 GB. You can make the resulting ZIM files even smaller if you specify the option to only index article intros for full-text search (since the full-text search Xapian index is a large fraction of the file size). The converter is multi-threaded and written fairly efficiently using Rust, so you can convert all the millions of articles in a full English Wikipedia Kiwix files in a few hours on a typical modern computer.

GitHub link: https://github.com/sultanqasim/szmcp

▲
16
 
18👁
r/LocalLLaMA · u/SignificantZebra5883 · 9d ago
i would like to learn deeply about fine-tuning local models before burning money

There's so many new techniques like RL, RL LoRA, QLoRA, CPT LoRA.

I believe i would have a usecase for them, but i don't know where to learn, youtube is filled with bad quality tutorials if i just search and the good channels (fireship, bycloud) don't cover these as they're quite new concepts, i guess?.

how can a regular joe like me learn about these concepts in a "practical depth" so i can actually fine-tune qwen 27b successfuly on lets say custom corpus? without spending 100$ figuring out that "oh i didnt even need CPT here" or "well i chose the wrong Rank count! time to start this 2 day run again!"

context and TLDR: im building a legal general purpose chatbot for context, i have a big corpus, but im a bit stuck on what to do next

thanks for reading and any pointers!

▲
16
+3
12👁
r/LocalLLaMA · u/Dev-in-the-Bm · 10d ago
Best approach for automatically tagging local music collection?

I don't use music streaming services much, and listen to music from my own local collection.

I don't use any local streaming servers like Plex or Navidrome, they wouldn't work for me because I use a dumbphone and play music off of my SD card.

I've manually built a bunch of mood based playlists so I can easily pull up a playlist with the music I want, but that's
obviously very tedious and inefficient.

I've been playing around with ML models to automatically add genre, mood, and other tags to my collection, the open models available today are insane.

The thing is I haven't been able to find any polished tools for doing this.

Most of what's available is either CLI or built for streaming servers.

Is there anything I missed?

Should I just setup a streaming server just for tagging the collection, or is there a better way?

💬 23 (+1) open on reddit ↗
▲
16
 
7👁
r/LocalLLaMA · u/Danmoreng · 12d ago
Gem16 - custom engine for Gemma4 12B & 26B on Blackwell 16GB GPUs

It’s probably a bit niche and the models are a bit old at this point, but after reading about Ninfer a few months ago I did my own small vibe coded engine project for my 5080 Laptop GPU. Initially I thought I can only fit the 12B model with enough context into the VRAM, but with custom quantisation (EXL3 like) the 26B fits nicely as well. The engine is entirely Codex written, but it took a lot of weekends to make it work and make it work as fast as vLLM/faster since vLLM didn’t work with MTP on 16GB VRAM. Also, my engine works on Linux and Windows equally well. Primarily this is designed to be single-user only, the 12B model can serve 2 sessions. It also comes with a native fancy looking GUI, but the main focus was the engine itself. 12B with audio & vision, 5.800 t/s prefill & 87 t/s decode 26B with vision, 5.660 t/s prefill & 182 t/s decode, fits 220k context https://github.com/Danmoreng/gem16 Sadly the most interesting feature of the 12B model with native audio understanding seems to have the issue, that after around 8k context the model doesn’t recognise audio tokens anymore. This seems to be a model issue, as others have also reported it: https://huggingface.co/google/gemma-4-12B-it/discussions/45 Would love to get some feedback!

▲
16
+2
17👁
r/LocalLLaMA · u/WebAssemblyMan · 12d ago
What if open-source AI focused less on giant models and more on reusable capabilities?

Instead of everyone building another general-purpose model, the community could distill open models into domain specialists—biology, Python, accounting, OCR, and more. Developers could combine these capabilities into local tools: small model + OCR + accounting → local accounting assistant Like Linux, open-source AI could grow through shared components rather than complete systems. Could domain capabilities become the fundamental unit of contribution?

▲
15
 
13👁
r/LocalLLaMA · u/danielfrances · 10d ago
Help me find a good stack for reversing an old online game client

Hi, so I was working on building a local server for an older online game client a few years ago, and the amount of data I had to synthesize was intense. I ended up shelving the project. I had managed to build a basic login server, sorted out some crypt stuff, but it was just way too slow of progress for me. I've got some decrypted packets and lots of data to work with, so the LLM is not going to be forced to do this entirely blind.

I restarted it recently with Fable, and as expected, it has been a huge help. However, I'm consistently hitting the safety guardrails now that I am further into the project. I am wondering what you all would suggest for a local setup? I have 16GB of VRAM (RTX 4060 Ti) and 128GB of DDR4. If that is entirely insufficient, I might be willing to pay for hosting a more powerful local model. I'm fine with it being slow and chugging along all day and night - I am primarily concerned with it actually figuring out the client functions, and doing things as accurately as possible.

I appreciate any insight into specific models, harnesses, and other stuff I should be looking into. Thanks!

▲
14
+5
32👁
r/LocalLLaMA · u/FantasticNature7590 · 7d ago
I spent 3 weeks testing local Qwen3.8 on the new low-latency SGLang/vLLM recipes: DFlash2 2.8x. Builds: RadixArk + Inferact 27B NVFP4, 27B BF16, orcarouter 27B Uncensored, Flash-Next NVFP4

Hey guys,

Last time I tested Qwen3.8-Flash-Next on its own. This time I put three Qwen3.8 checkpoints through the same 10 tests on the same RTX PRO 6000:

  • RadixArk/Qwen3.8-27B-NVFP4 (dense)
  • orcarouter/Qwen3.8-27B-Uncensored-NVFP4 (dense, uncensored fine-tune)
  • RadixArk/Qwen3.8-Flash-Next-NVFP4 (MoE)

Each model got the same prompts and its own model card's sampler, with one attempt per task.

Video with the battles, the castles and the ball run: https://youtu.be/VOtfja\_Toj4**

Short version

  • Tests won: Flash-Next 5, 27B 4, Uncensored 0, plus one tie (long-context recall was 100% for all three).
  • Prefill, full window: Flash-Next 22.4s, 27B 97s, Uncensored 99s. Not the same power cap, see section 1.
  • Speculative decoding on the 27B: the DFlash2 drafter took Spec-Bench from 75 to 210 tok/s for one user, 2.8×.
  • SGLang vs vLLM: SGLang was faster overall (210 vs 160 tok/s), but that's mostly the checkpoint. On the one export I ran on both engines, vLLM was 16% faster (169 vs 146).
  • Tool use (BFCL subset, thinking off): 27B 73.3%, Uncensored 70.8%, Flash-Next 64.5%.
  • Battle arena: the 27B scored 700/1000, ahead of Claude Fable 5.1 (678) and GPT-5.6 (473), both entered through their chat apps at max thinking.
  • Rube Goldberg machine: only Flash-Next got the ball into the cup. Both 27B models spent their whole \~111K-token answer budget thinking and never placed a part.
  • CAPTCHA (40 puzzles, local copy): 27B 24/40, Flash-Next 21/40, Uncensored 19/40.
  • Things you look at: Flash-Next made the best voxel castle, the best design board and the best video edit (19/20 on my rubric).

https://preview.redd.it/4e22yukvi4th1.png?width=1484&format=png&auto=…

Setup

  • GPU: one NVIDIA RTX PRO 6000 Blackwell, 96GB
  • CPU: AMD Ryzen 9 9950X
  • System RAM: 96GB DDR5
  • 27B and Uncensored: lmsysorg/sglang:v0.5.20, DFlash2 drafter, 262,144-token window, 4 slots
  • Flash-Next: lmsysorg/sglang:dev-qwen38-next-local, built-in MTP drafter, 262,144-token window, 1 slot. Its BFCL run used sglang:v0.5.20, like the 27Bs.
  • Sampler: the model card's thinking settings at the highest effort for the agent tests. BFCL and the needle test use the card's non-thinking settings.
  • The agent tests (SVG, video editing, voxel, design, Rube Goldberg, CAPTCHA) run inside Pi, a coding agent, with bash, read, write and edit. CAPTCHA gets only screenshots, mouse and keyboard.

One workstation, one model server at a time, and every number comes from a saved run.

1. Speed: drafters, SGLang vs vLLM, and long prompts

For the 27B I ran a speed matrix: every drafter, two engines, and four builds (three NVFP4 exports, one of them the Uncensored fine-tune, plus full-precision BF16). Each arm got its own server from a cold boot, the card's sampler and the 400W cap. "One user" is the Spec-Bench median over its 480 prompts. Engines: lmsysorg/sglang:v0.5.20 and vllm/vllm-openai:v0.29.0.

Which drafter (tok/s, one user):

|Drafter|SGLang · RadixArk NVFP4|vLLM · Inferact NVFP4|
|:-|:-|:-|
|none|75|59|
|MTP (built into the model)|160|113|
|DSpark|174|137|
|DFlash2|210|160|
|DFlash2 + torch.compile|214|not run|

DFlash2 wins on both engines. It keeps about 3.7 drafted tokens per step, against 2.9 for MTP.

Which build, on which engine (tok/s):

|Build · engine|DFlash2, 1 user|No drafter, 1 user|DFlash2, 4 users (total)|
|:-|:-|:-|:-|
|RadixArk NVFP4 · SGLang|210|75|607|
|Inferact NVFP4 · vLLM|160|59|517|
|Uncensored NVFP4 · SGLang|146|46|473|
|Uncensored NVFP4 · vLLM|169|63|538|
|BF16 · SGLang (full precision)|97|29|291|

  • The engine gap depends on the build. RadixArk's export on SGLang was the fastest arm overall, but on the one export I ran on both engines (the Uncensored), vLLM was 16% faster.
  • The NVFP4 exports aren't interchangeable. Same architecture, same 4 bits, same engine (SGLang), same drafter: RadixArk's export ran 210 tok/s and the Uncensored one 146.
  • 4-bit vs full precision: NVFP4 with DFlash2 is 2.2× the BF16 speed.
  • Prefill doesn't care about the engine: a full 245K-token window took 96–103s on every NVFP4 arm, SGLang or vLLM. BF16 took 129–135s.

https://preview.redd.it/vnd83h5zi4th1.png?width=1484&format=png&auto=…

https://preview.redd.it/0v1unh5zi4th1.png?width=1484&format=png&auto=…

The three models:

|Metric|Qwen3.8-27B|27B-Uncensored|Flash-Next|
|:-|:-|:-|:-|
|Prefill, full window|97s|99s|22.4s|
|Decode, Spec-Bench, one user|210 tok/s|146 tok/s|not run|
|Drafter vs no drafter|2.8×|3.2×|n/a|

The 27B keeps writing at 223 tok/s with a full 245K-token window behind it. Speculative decoding depends a lot on the content: maths ran at 339 tok/s, roleplay at 149. The pattern was the same for both 27B builds.

One important caveat. Flash-Next's speed test ran on 2026-09-12 at a 600W power cap. I later moved the card to 400W, and the 27B matrix ran at that cap. In my power sweep, prefill lost about 6% per 50W removed, so some of the gap is the cap. Moe also helps

https://preview.redd.it/jx28sde1j4th1.png?width=1484&format=png&auto=…

2. Tool use: the dense 27B leads

This is a 900-case BFCL v4 subset (11 categories), not the full leaderboard. Thinking was off, with temperature 0.7 and top\_p 0.8 from the card.

|Metric|Qwen3.8-27B|27B-Uncensored|Flash-Next|
|:-|:-|:-|:-|
|BFCL core|73.3%|70.8%|64.5%|
|Tool accuracy|87.8%|88.0%|82.4%|
|Abstention|79.5%|69.5%|68.5%|
|Multi-turn|52.5%|55.0%|42.5%|
|Malformed calls|0.08%|0.27%|0.28%|

These aren't comparable with my last post's Flash-Next BFCL numbers, which used temperature 0.

https://preview.redd.it/yqdc2j32j4th1.png?width=3396&format=png&auto=…

3. Long context: perfect for all three

I hid a fact in a log file that filled 33%, 66% or 99% of the 262K window, at three depths, with three needle types. The cache was flushed before every request.

  • 27B: 27/27
  • Uncensored: 27/27
  • Flash-Next: 81/81 (three samples per cell instead of one as I run this at the beginning)

The largest prompt was about 259.5K tokens.

https://preview.redd.it/fgeuq733j4th1.png?width=3396&format=png&auto=…

4. Battle arena: the local 27B beat Claude

Each model gets a rules sheet and a 1,000-point budget. In the open arena it designs one army blind and fights 13 armies: nine historical references plus the other entries. The score is 1,000 × its average win rate. Every matchup is 200 deterministic battles (100 seeds, sides swapped).

|Rank|Entry|Score|
|:-|:-|:-|
|1|RadixArk/Qwen3.8-27B-NVFP4|700|
|2|Claude Fable 5.1 (chat, max thinking)|678|
|3|Qwen3.8-27B-Uncensored|603|
|4|Qwen3.8-Flash-Next|535|
|5|GPT-5.6 (chat, ultra thinking)|473|

In the gauntlet, the model sees each enemy and builds a counter. The 27B beat 12/15, the Uncensored 12/15 and Flash-Next 13/17. Flash-Next ran an earlier version of the gauntlet with two more enemies, so treat that row as close, not ranked.

Thinking cost: the 27B's arena army took 39K thinking tokens in 4 minutes. Flash-Next's took 73K in 9 minutes.

https://preview.redd.it/1wbb6p28j4th1.png?width=3396&format=png&auto=…

https://preview.redd.it/o2yccq77j4th1.png?width=1484&format=png&auto=…

5. The SVG test is also a fact check

Prompt: find out which card local-AI hobbyists run and which current open model fits it, then draw the card lifting the model, labelled with a quant and a size that fit. All three picked the RTX 3090. I checked every label against what each session actually fetched.

  • 27B: 5/5 facts correct. Qwen3-Coder-30B-A3B at Q5\_K\_M, 21.73 GB, the real file size. Q6\_K at 25.09 GB is correctly marked as not fitting.
  • Flash-Next: 4/5. It got all four file sizes right and the exact 3.3B active parameters, but labelled the 3090 with a "12VHPWR, melted once" joke. That's the wrong card.
  • Uncensored: 3/5. It labelled the model "QWEN3.8-27B" but used the file size of Qwen3.6-27B Q4\_K\_M, and its "12 tok/s" isn't in anything it fetched.

All three passed 8/8 format checks. The 27B looked at its render twice and Flash-Next three times, where the rule allows one look.

https://preview.redd.it/oijy66u9j4th1.png?width=1920&format=png&auto=…

6. Video editing, voxel and design

Video editing: the model gets a raw 132-second take with fillers, a retake and a swear. It never sees the footage, only transcription and silence-detection tools, and then edits through FableCut's tools. There are two cases, each scored by hand out of 10:

  • Flash-Next 9 + 10 = 19
  • 27B 8 + 8 = 16
  • Uncensored 5 + 8 = 13

Voxel (Wawel Castle in three.js), ranked by eye:

  1. Flash-Next is the only one with the gold Sigismund Chapel dome and the Vistula bending around the hill.
  2. 27B built a clean but generic castle.
  3. Uncensored placed the camera inside its own build.

Flash-Next also used the fewest thinking tokens there: 73K, against 101K for the 27B.

Design (an animated explainer board in my design system), ranked by eye: Flash-Next first, and the two 27Bs shared second. All three passed 7/7 hard rules.

https://preview.redd.it/ue511ppaj4th1.png?width=1920&format=png&auto=…

7. Rube Goldberg: only one machine

The setup is a fixed level: a ball on a ledge, a cup on the floor and a wall in between. The model writes a parts list (no code), and a 2D physics engine runs it. It can run and look as often as it likes within 90 minutes. The score is automatic: does the ball itself end in the cup?

Flash-Next: yes, at 12.4s. It made 49 simulator runs and used 40 parts (35 of them moved). The ball travelled 1,672 px. It used 224K thinking tokens and compacted its context 8 times.

27B and Uncensored: no machine. Both spent about 111K tokens thinking in their first answer, reached the per-answer limit and stopped before writing a single part. Everyone got the same rules and one attempt. A rule that let a model continue after hitting the limit might change this, and I haven't tested that yet.

https://preview.redd.it/px8mlhhbj4th1.png?width=1280&format=png&auto=…

8. CAPTCHA: local models in a real browser

I used Open CaptchaWorld (20 CAPTCHA types, two of each). The model only sees screenshots and only acts with the mouse and keyboard. The site's own checker marks the first answer, and it must arrive within 7 minutes.

|Model|Solved|Median time|Thinking tokens, all 40|
|:-|:-|:-|:-|
|Qwen3.8-27B|24/40|55s|456K|
|Flash-Next|21/40|145s|1.0M|
|27B-Uncensored|19/40|36s|401K|

With its own 20-minute limit, Flash-Next solved 23/40. The paper reports 93.3% for humans and 40% for the best agent on its full set, which isn't the same 40 puzzles. With one run each, a three-puzzle gap is not a strong signal.

Which one should you run?

  • RadixArk/Qwen3.8-27B-NVFP4 for agents and tool calls. It won BFCL, the arena, the SVG fact check and CAPTCHA.
  • RadixArk/Qwen3.8-Flash-Next-NVFP4 for long prompts and building things, especially visual ones. It reads a full window much faster and won video editing, voxel, design and the Rube Goldberg machine.
  • orcarouter/Qwen3.8-27B-Uncensored-NVFP4 only if refusals are your actual problem. It won nothing here and invented facts in the SVG.

Resources

Configs, Docker setup and reports

show remaining 403 characters

The test harness is still private while it's changing.

Full video with the battles, the castles and the ball run: https://youtu.be/VOtfja\_Toj4**

I abused AI to help write this up and to check it against the report. Every number above comes from a saved run.

Which test do you find most interesting and maybe you have some other creative ideas how to test models?

💬 8 (+1) open on reddit ↗
▲
14
+3
15👁
r/LocalLLaMA · u/KissMyShinyArse · 8d ago
Strata: how to configure sampling parameters

The top-level README doesn't mention this, but you can add a "sampling" key to your strata-iq3_s.json like this:

{
"sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20},

"exe": "/path/to/Strata/engine/strata",
"args": [ ... ],
...
}

From docs/DETAILS.md:

The run config's optional sampling block sets the defaults for requests that leave the fields out ("sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20}); a request's own fields always win, and with no block at all a request without sampling keys decodes greedy.
💬 10 (+1) open on reddit ↗
▲
14
-2
11👁
r/LocalLLaMA · u/norenEnmotalen · 9d ago
peculiar-ragdoll's Dirk-Qwen 3.8-27B vs. UkisAI Swift-1.5 Qwen3.8-27B

EDIT: Post 2 with more model fint-tunes here https://www.reddit.com/r/LocalLLaMA/comments/1wv1ico/unsloth\_swift15\_peculiarragdoll\_thinkingcap/

I have a long list of my own domain specific eval questions that I run to validate which models I can rely on: coding, coding (numpy/pandas), data analytics decision making, local RAG, and voice assistant. It's made up of the types of things I'm likely to deal with on the daily. The test questions vary in dififculty and composition: easy, medium, hard.

System: M1 Max 32c 32GB with context 128K for Dirk and 110K for Swift.

Swift doesn't have XL. So I had to test with the L quant to stay as close as possible.

I ran the eval (using my tuieval tool) on peculiar-ragdoll's Dirk-Qwen3.8-27B-UD-Q4\_K\_XL and Swift-1.5-Qwen3.8-27B-Q4\_K\_L loaded with a modified version of Splash. The "amalgam" is a local I made out of incoai/Splash 1.1 and paperniuk's apple7-m1-kernels. It is modififed a little but not in ways that would alter model performance. I only merged and tweaked for some memory features I like from llama.cpp such as fit context check at the start of a load and personal QoL updates re auto-context manipulations that I don't want to think about, etc.

To say this result surprised me is quite an understatement. It's blown my mind.

When I did the first test a couple of days ago with only 44 questions, I thought it must be a prompt caching issue I missed that Dirk was benefiting from. I validated it is not and ran it against a lot more questions to certify it. It's a legit test outcome.

Dirk-Qwen is much sharper at getting to decisions and responses. The "be brief" instruction that gets passed each time in the chat templste is doing more magic than I had anticipated. It also gets more answers correctly with way less time consumed.

What trips up Swift-1.5 are mostly hard questions. It tries and tries until the 16,384 max token limit per question is reached and it fails with truncation.

Even when you ignore the 16,384 truncation failures and compare the other questions, Dirk token usage comes out on top.

Snipped view... this basically goes on pattern for another 191 unique questions.

https://preview.redd.it/otg3d1rkbqsh1.png?width=1420&format=png&auto=…

More importantly, this behavior is not just in question answering. You can see it in actual code refactor tasks.

On an unrelated note: tne model that has been able to pass a 100% of my eval packs is Opus 5.5. Deepseek Flash 4.1 fp32 got them all right except three.

💬 22 (+1) open on reddit ↗
▲
13
+1
9👁
r/LocalLLaMA · u/SnooPeripherals5313 · 11d ago
3D/2D Text Visualisation post image

Like everyone, I use 3js for data visualisation. But while semantic clusters are interesting, they don't confer much practical information alone.

So I did something very simple: a query spatially re-assembles the nodes, and you can switch them to text.

Honestly, it's hard to swing 3D viz for text as a genuinely useful feature and not a novelty, but I get the feeling there's still some potential in the idea. Would be good to discuss, I'm sure someone here has made a better implementation.

▲
13
 
9👁
r/LocalLLaMA · u/Admirable_Reality281 · 12d ago
Xiaomi MiMo 2.6 Flash vs GLM 5.3 Flash

I've seen a lot of conflicting opinions about MiMo Flash, but I haven't tried it yet. How does it compare with GLM Flash for coding work like this? I'm interested in: \- back-end development \- debugging, refactoring, implementing features in an existing front-end codebase \- maintaining Docker images \- troubleshooting DevOps errors Not the silly stuff I see "build me 100 nice-looking webpages" or "make me a Three.js demo". So far, I've been happy with GLM 5.3 Flash. My main frustration is that it sometimes overthinks too much, and once it does, it's hard to steer it back on track. The DeepSWE score of MiMo appears to be a substantial improvement over GLM's, but \- there's no official score from DataCurve \- no amount of consumed tokens to achieve it and in general one benchmark doesn't tell me how it behaves on day to day work. I'd be interested in comparisons from people who've used both.

▲
13
-2
12👁
r/LocalLLaMA · u/Smooth-Television-48 · 13d ago
Navigating Cost Efficient Hardware in these Volatile Times

Where to even begin on this one...I guess I should start by acknowledging the risk vs reward for vendors other than nvidia, so: Yes I understand that nvidia are dominant currently on speed (llm and imagegen) and software ecosystem. I am too am hopefuly that software stack support continues to improve with other vendors. The current lag for other vendors is not a priority concern (it falls behind the primary price concern). Entry points for "decent" local inferencing look to be circa AUD 2000-2500+ (the price of a 2nd hand 3090, or two 3060s, b60 48gb, r9700 32gb), and yes other older architectures are available (eg. V100)...but they really end up around the same costs once all said and done. Workload will be a mixed bag with some DL/ML training/development projects, but when not doing that I'll consume HF models to run a coding agent, imagegen (just for the fun of it/try out video and for laughs), and probably dive into finetune/distilling. Hence, I'm looking around that sub AUD 5k mark to dive in and FAFO, but I don't want to be needlessly cavalier in my purchase either... Asking AI is no real use because it's out of touch with modern markets until you correct it a bunch. It's also out of touch with software stack development/progress. So I put it to the hive mind, where is the money best spent for diving deeper into local? \- accepting prices wont change and pay 2k a piece for 2nd hand 3090's. \- find some 16gb variants and get 4 instead of 2. \- dive into the intel arc rabbit hole with the b60 dual (48gb, but it's just 2xgpu on a single pci slot) \- AMD path (r9700 seems the best price point but could wait 3 months to see what the new 10x series looks like) \- unified memory systems (honestly the price vs performance just doesn't seem worth it at this point) ETA: I have a threadripper and a lot of DDR4 RAM, but current motherboard is constrained to 2 x16 physical slots. I also have a nvidia gpu already....but I dont want that to impact the core of the discussion as I could move that into a different system and use it to server models that fit wholly in its vram footprint.

▲
12
+1
34👁
r/LocalLLaMA · u/klieret · 7d ago
New benchmark on LMs fixing bugs before users run into them

Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW.

Most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But shouldn't we expect models by now to also find bugs before anyone runs into them?

So in SWE-sweep we just hand an agent a big codebase and ask it to find & fix as many bugs as it can. We then give a score based on a hidden set of bugs that we know about in the repos. All the bugs are real-world bugs. We do a lot of filtering to make sure the bugs are actually discoverable & fixable from reading the repo alone.

https://preview.redd.it/irffy7x5v2th1.png?width=1080&format=png&auto=…

We're still expanding the leaderboard list with more local models (unfortunately it's always a big harder with funding/infra etc), but right now it seems like it's quite hard to beat Luna xhigh in terms of cost efficiency.

Also the scores are way lower than I would've expected. Some tasks are legitimately superhuman in practice (like fixing up all of numpy), but there's also lots of small repos, where I would've expected a lot more from current models.

Everything is open source (MIT license) on github and you can find paper etc. on the website.

Happy to answer questions here, also super curious what open weights models you'd recommend running next (we're working on an update next week).

💬 14 (+2) open on reddit ↗
▲
12
 
32👁
r/LocalLLaMA · u/AdRepulsive7837 · 7d ago
best <40B alternatives to Qwen/Deepseek for (1) Coding (2) Long document QA test

Due to some reasons, Qwen/Deepseek Chinese models are NOT allowed in the my workplace. So, what local models, do you think, is the best alternatives to Qwen/Deepseek for

(1) Coding

(2) Long document QA test (like giving a long medical history of 120k tokens and ask a question based on that medical history)

Gemma 31B ?

Muse Glimmer 30B ?

Nemotron ?

also, I know that nothing beat qwen nowadays, but are there fine tunes from these alternative model that make them better than qwen3.8 27B in terms of coding ?

💬 37 (+5) open on reddit ↗
▲
12
 
15👁
r/LocalLLaMA · u/DerTomsn · 10d ago
Swift-1.5-Qwen3.8-27b-oQ8e-mtp on Apple M5 Max — 34.8 tok/s — llm-bench.io

I ran Swift-1.5-Qwen3.8-27b-oQ8e-mtp through the llm-bench.io a few times today: oMLX on M5 Max 64 GB, thinking on at xhigh, 262k context window.

The big difference between Swift 1.5 and the base Qwen3.8 27B is how much it writes. Per full run (agent workflow, code generation, research, role play) Swift averages 51k generated tokens and Qwen3.8 averages 77k.

| Scenario|Swift 1.5|Qwen 3.8 27B|
|:-|:-|:-|
|Code generation|28.6k|40.1k|
|Research|11.1k|21.2k|
|Agent workflow|7.5k|11.9k|
|Role play|4.0k|3.8k|

The full benchmark run duration: avg. 24 min for Swift 1.5, avg. 38 min for base Qwen 3.8 27B

Everything else is about equal:

  • generation speed: 34.6 vs 32.8 tok/s
  • prompt processing: around 400 tok/s for both
  • quality score (the site's LLM judge): 85.8 vs 85.0. My four runs range from 84.3 to 87.2, so well within the expected variance of the llm judge. I'd call it a tie.

Still 3/4 of what Swift generates is reasoning, it just does less than Qwen 3.8 27B. The output is still very usable. I'll for sure give it a try to be my daily driver for a few day.

Runs Swift 1.5:

Runs Qwen 3.8 27B:

▲
12
+1
9👁
r/LocalLLaMA · u/ButtercupLyn100 · 10d ago
I’m building an open-source browser agent that can run locally with LM Studio/Ollama — including a 450M browser VLM

I’ve been working on an open-source project called WebBrain that gives LLMs the ability to see and operate a browser.

One thing I really wanted to avoid was making the browser agent dependent on a single cloud model/provider.

So WebBrain can work with local models through things like LM Studio and Ollama, as well as cloud APIs if you want them.

I also ended up training a small vision model specifically for browser tasks:
webbrain-vl-2-450M

It’s based on LFM-2.5-VL-450M and fine-tuned on browser screenshots/tasks. The idea is that instead of sending every screenshot to a giant multimodal model, some browser perception can happen with a very small model locally.

It can run through WebGPU directly on the user's machine.

The agent itself combines screenshots with the browser accessibility tree rather than relying entirely on DOM parsing.

Current architecture is roughly:
• screenshot + accessibility tree for perception
• browser-specialized tiny VLM where useful
• model-agnostic planner
• local models via LM Studio/Ollama
• Chrome / Edge / Firefox / Chromium support
• optional cloud execution
• open source

I'm especially interested in figuring out how far browser agents can realistically go with small local models rather than GPT/Claude-scale models.

Repo: https://github.com/webbrain-one/webbrain
Model: https://huggingface.co/webbrain-one/webbrain-vl-2-450M
Dataset: https://huggingface.co/datasets/webbrain-one/webbrain-vl-2-450M-dataset

Would be very interested in feedback from people here running smaller Qwen/LFM/MiniCPM/etc. models locally — particularly what model you would try as the planner.

▲
12
 
12👁
r/LocalLLaMA · u/Merchant_Lawrence · 11d ago
Need small model that can work for tool caling and agent

Hi. So....... after toturing my 750 ti 4 gb and 16 gb ram with image gen model .i want continue experiment with agent mode like hermes or opencode, using local model but before go i want ask few question. are big model = good perfomance or small model can do same stuff. what small model recommend for agent my spec what caveat of doing this ?

▲
11
+6
27👁
r/LocalLLaMA · u/GodComplecs · 7d ago
Should we plead opensource labs to still produce great non thinking (instruct) models?

The results are in, no thinking / instruct mode for new models degrade performance more than on old models such as 3.6 vs 3.8, where 3.6 takes the lead on several coding benches in instruct mode.

I would ask the labs to still nicely focus on instruct mode also still, there a probably gains to be had without the lengthy reasoning still, some of us still use models for everything and they do not need long reasoning traces. Agentic is fine and all but to start SACRIFICING performance for the "base" model which we are used to from early Llama days is not a good direction imo.

💬 20 (+4) open on reddit ↗
▲
11
+1
15👁
r/LocalLLaMA · u/jjusko20 · 10d ago
SFTMill: Easily [off-policy] distill any existing LLM with an OpenAI Compatible Endpoint. Turn any behavioral goal into a comprehensive dataset. post image

Disclaimer: Any\* means any model that exposes its CoT without it being censored.

Hey guys - half a tutorial/guide, and half an I built this, so I went for resources. This is something that I created for myself recently when I couldn't find any good existing solution. I wrote this post myself, no AI!

Probably a fair number of you have seen my posts about fine-tuning AliceAI 80B A3B according to my own synthetic datasets. If you did, I'm still fine tuning it on a live stream right now - check out https://figure-bios-expect-cio.trycloudflare.com/ \-- it'll let you inspect any and all of the training data that I generated with this engine. If you have any interest, it's pretty neat! unfortunately that link is optimized for desktop only and I'd have to kill the run to reset it, so u may want to rotate the phone.

That thread was at https://www.reddit.com/r/LocalLLaMA/comments/1wslgkw/comment/pcw4kgw/?context=1&screen\_view\_count=1

That run is using off-policy distillation, and that's I made this for. my training data for that project with this repo, and just customized it for an OSS release. Basically, you create a "curriculum" for your goal - e.g. if I was training an agentic model, I'd need things like tool calls, bug fixing, working in a workspace, tracing errors, etc. You define your curriculum in a yaml file, then an LLM creates tasks based on the curriculum you defined, and the chosen LLM you're distilling from then solves each task, leaving you with a full Q/A set that encompasses your fine tune goals.

I used qwen 3.8 27b on medium to generate the tasks - I'd recommend avoiding anything any weaker than that.

I forked my private repo of this that I've been using into SFTMill, which is basically just the same thing with great documentation and a few steps added to get anyone onboarded rapidly. I created it \[and my original version\] because I couldn't find any existing pieces of software made with this design, for this purpose.

I release it because I enjoy contributing to the community, and there's a vague hope someone will eventually see one of my pieces of work and want to hire me (if you're reading this and you like the project and you need a software/ml engineer remote or in NYC, let me know <3). It makes me happy when my software helps others so I'd love if you let me know if it helped you. Cheers!

Shoutout u/FullOf_Bad_Ideas for helping me with my alice train in areas I wasn't experienced enough in - I threw a Multi-Turn Hybrid-Reasoning (user <> assistant) section in the readme just for you bud, hope it helps.

https://github.com/jackjusko/sftmill

▲
11
-1
8👁
r/LocalLLaMA · u/Brilliant-Hall1387 · 10d ago
Sherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version

Not an LLM, but the ternary findings should carry over, and we hadn't seen Sherry-style 3:4 weights run in a browser before. Disclosure: this is our work at Precisit, everything is MIT.

What it is

  • A 7.4M-parameter one-pass scorer (the jevlike family): the board goes in, one score per legal column comes out. No search.
  • Weights in T34, Sherry's 3:4 format: in every four weights one is zero and three are ±1, so four weights fit in 5 bits. One fp16 scale per 128 weights gives 1.375 bits per weight. The embedding is int8; norms and biases are fp16.
  • It runs in the browser on a small WebGPU runtime: 1.1 ms per move (idle M5 Pro, Chrome).

|Model|File size|vs depth-4 bot|vs depth-6 bot|
|:-|:-|:-|:-|
|dense (fp32)|29.7 MB|0.92|0.89|
|T34, trained ternary|1.59 MB|0.93|0.91|
|T34, fine-tuned from dense|1.59 MB|0.89|0.9|
|T34, converted after training|1.59 MB|0.13|0.11|
|Base243 (TQ1\_0 style), trained|1.93 MB|0.89|0.88|

200 games each, both sides play a random move 5% of the time, a win counts 1 and a draw ½.

What we learned

  1. Converting the finished model to 3:4 collapsed it (0.13 against the depth-4 bot). Training with the format in the forward pass fixed it completely, whether from scratch or fine-tuning.
  2. Attention's q/k/v matrices are the sensitive ones. Group size (64/128/256) barely mattered.
  3. Seeds matter: two runs of the same T34 recipe scored 0.945 and 0.882.

Play it:
https://precisit.github.io/onepass-web/demo/c4-size/

Code, models, every result:
https://github.com/precisit/onepass-webgpu-ternary

The write-up:
https://precisit.com/en/blog/onepass-c4-size/

Has anyone gotten post-training 3:4 conversion to work on models, or does it need training?

▲
11
+1
10👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 10d ago
What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects?

For example, directly using claude code or code, which is then hooked up to automatically delegate the actual code writing tasks to a local model like qwen 3.8 flash next, to save on cloud usage limits.

I’m imagining the loop would be:
User writes prompt
Claude/codex thinks about it and the plan
Claude/codex sends the specific and bounded coding instructions to the local model+harness (opencode, pi, etc) via api endpoint or MCP, with clear instructions on a defined endpoint
One the local model+harness hits the clear endpoint/“done” step, it sends a ping back to claude/codex
Claude/codex then verifies the output and then thinks about next steps to instruct the local model+harness on

Does this actually lead to improved savings on the cloud model usage while preserving code quality? Or does this end up being unnecessarily complex and not saving on any cloud usage

💬 21 (-1) open on reddit ↗
▲
10
+3
18👁
r/LocalLLaMA · u/Designer_Elephant227 · 7d ago
Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal.

What i run:

\- model: Qwen3.8-Flash-Next from turboderp, exl3 5.05 bpw (head 6 bit, vision 6 bit, mtp 5 bit)

\- backend: exllamav3 rocm fork from phoenixhaxor (commit cbbef08), i had to patch one file for gfx12

\- cpu/ram: Ryzen 9 7945HX3D with about 90gb ram

\- 112 of 512 experts per layer are on the gpu, the other 400 on cpu with 16 threads

\- 262144 context, q8 kv cache, chunk size 4096, batch 1

\- mtp drafting is on, acceptance is around 53-54%

\- the ngram table (102gb, bf16 not quantized) gets streamed from nvme

Speed at 230k context (prose): prefill 863 t/s (only the new 101k tokens, the rest came from the prefix cache) and decode 34.7 t/s.

Is this ok for the r9700 or can i still tune something? Thanks 🙂

💬 26 (+5) open on reddit ↗
▲
10
+2
11👁
r/LocalLLaMA · u/ParvusNumero · 7d ago
Mirostat?

Reading another post made me think:
Is anybody still using Mirostat?

It was all the rage and people said it avoided the “boredom trap” for long texts.

Do newer generation models not need that anymore, or are other samplers superior?

💬 24 (+3) open on reddit ↗
▲
10
+1
11👁
r/LocalLLaMA · u/Biomass23 · 9d ago
tp=6 can work on vLLM, with padding

vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).

I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.

So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.

I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.

GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x

▲
10
+1
16👁
r/LocalLLaMA · u/dh7net · 10d ago
Distributed Local Agents Benchmark.

I wanted a setup where I can compare all the harnesses with all the local models.

It turned out to be a rabbit hole. For instance, you would not only have to test all the harnesses (being sure that they are well configured), but also all models, with all their flavors, and this for all kinds of hardware.

Everyone can do their share, but no one can pretend to do all possible tests extensively.

To solve this, I created a website where everyone can test the configurations they want and share the results if they want. You can try it here: airbench.ai

There is a leaderboard where I share the tests I'm making, but I hope I can populate it with tests from others. https://airbench.ai/leaderboard?k=poL

My hope is to turn this into a fully distributed Agent Benchmark.

Let me know what you think.

▲
10
 
10👁
r/LocalLLaMA · u/SignificantZebra5883 · 10d ago
50B+ MoEs with few active parameters, what's the sweet spot for intelligence, agent speed, and affordable fine-tuning?

I’m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:

Question → many sequential tool calls → final answer/document

Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.

I’ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. I’m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.

The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.

Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?

I understand that small active parameter counts don’t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also don’t assume that more total parameters automatically means a better model. I’m interested in where that tradeoff works in practice.

There are three things I’m trying to pin down:

  • Inference hardware: Ideally inference runs rented with parallel agentic loops (this is for a B2C project, not single person use, we scale based on demand)
  • Fine-tuning hardware: Obviously FT LoRA will take more memory than inference, max like 4GPUs on vastai fits the budget.
  • Agent performance: After it gets the prompt the tool calls and everything will be local, so imo it has no problems being blazing fast, as soon as the model calls a tool call it will be back very fast, so for this agentic use case, quick TTFT and t/s and adaptive dynamic reasoning are prefer right?

For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.

I’m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).

I’m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like “model + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,” along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.

Has moving from a \~30B model to a 50B+ low-active-parameter MoE actually improved your agent’s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading

💬 22 (+1) open on reddit ↗
▲
9
 
13👁
r/LocalLLaMA · u/nonlinearsystems · 9d ago
M5 Ultra - Qwen3.8 Flash Next vs Laguna S 2.1 post image

Spent today running a same-day, same-harness shootout between Qwen3.8-Flash-Next (oMLX, 182GB oQ8e, MTP) and Laguna-S-2.1 GGUF (LM Studio, 128GB, 8bit) on a Mac Studio M5 Ultra 256GB. Both capped at 262K context, thinking on, unique content per run with zero cached tokens verified each time.

That last part matters because my first run was wrong hah... shared prefixes across sizes let the KV cache carry over and 200K "prefilled" in 21s.

Prompt Qwen Laguna
8K 2.0s 10.3s
32K 7.4s 30.4s
64K 14.7s 70.8s
131K 30.1s 217.2s
200K 47.1s 455.4s

Qwen holds \~4,200 tok/s linear which is amazing. Laguna degrades superlinearly (quadratic attention doing quadratic attention things). At 200K, prefill is 94% of total time on both.

Decode (tok/s): Qwen 59-74 across sizes (MTP at 70-76% acceptance per server logs, roughly 2x). Laguna 68 down to 34 as context grows. No speculation on Laguna, its DFlash path already lost to plain decode on this hardware in earlier testing. I think if Laguna could get DFlash figured out or MTP, this might be a different conversation.

Quality was a draw, 4/4 each, on four problems with script-verified answers (Muse created the gymnastics here: exact 9-digit combinatorics, interval code with 12 hidden tests, fresh knights/knaves, asyncio ordering trap). Opposite styles though: Laguna answers in 5-10s with a few hundred tokens, Qwen deliberates exhaustively (one answer took 119s / 11K tokens). Both burned a full 8K budget on hidden reasoning with zero visible output exactly once, then converted on a 16K retry.

Happy to answer methodology questions. Full writeup with charts and the test rig diagram: https://echalupa.com/blog/qwen-flash-next-vs-laguna-200k

▲
9
-1
24👁
r/LocalLLaMA · u/stevyhacker · 9d ago
Five local models, 6.8 GB of weights: my open-source Mac meeting notetaker

https://preview.redd.it/xs01ds930nsh1.png?width=4800&format=png&auto=…

Back in July I shared LokalBot here. It's a free, open-source Mac app that records your meetings and keeps a daily summary of your activity, all on-device.

0.9.2 came out today. Since July I've benchmarked every model in it and swapped most of the defaults for smaller ones. The whole stack is now 6.8 GB:

  • Qwen3-ASR 1.7B (MLX, 8-bit): transcription
  • Nemotron 3 (Core ML): who spoke when
  • Qwen3.5 4B Q4\_K\_M (llama.cpp): notes and action items
  • Harrier 0.6B Q8\_0: search embeddings
  • LFM2.5 1.2B Q4\_K\_M: autocomplete in any app
  • Apple Vision: screen OCR (opt-in)

A few numbers from my M4 Max (48 GB):

  • 26-min meeting to finished notes in 33 s warm, \~85 tok/s decode
  • Speaker error went from 43.4% to 14.6% DER on AMI. That's against my old pyannote setup, so it says more about my config than about pyannote.
  • Autocomplete p95 went from 1.83 s (Gemma 4 E4B) to 0.49 s

There's also a read-only MCP server and CLI, off by default, so Claude Code or any other MCP client can pull context from your meetings.

I don't have any 16 GB or other M series numbers yet. If you've got one of those, especially M5 or M6 I'd love to see what you get.

I also tried MiniCPM5 2B for notes. It was smaller and faster, but it got stuck repeating itself on one summary and assigned action items to the wrong person. I kept Qwen3.5 4B as the default as saving a few seconds wasn’t worth getting who agreed to do what wrong.

💬 4 (+1) open on reddit ↗
▲
9
+4
27👁
r/LocalLLaMA · u/whatyathinkk · 9d ago
Do I need a UPS?

I know I could ask in some hardware subreddit, but I'm curious to know what people with multiple GPUs and expensive inference setups think about this.

I just moved to a new place and here the lights go out pretty frequently. 3 times over the last week, I came back to my computer being off due to a blackout (I guess it's a blackout, the entire neighborhood looses light for a few seconds/minutes). I have a desktop computer with 2x RTX5080s.

Do I need to buy a UPS to protect my computer from this? I get mixed answers about this topic. I don't mind my workflows being interrupted when the computer turns off, the only thing I'm worried about is the hardware being damaged. I have a good PSU, is that enough to protect the hardware?

💬 75 (+2) open on reddit ↗
▲
9
 
10👁
r/LocalLLaMA · u/segmond · 11d ago
Anyone customizing and Optimizing llama.cpp per model?

Basically the idea is take your favorite model, for example qwen3.8-27b or say dsv4vision. Strip everything out that is not needed by that model so the only thing needed is just for the model. Optimize the remaining code to be fast. The idea is to have a model also do this, provide it with enough tools, prompts, docs, guidance. I reckon that if we have a llama.cpp that is optimized for just one model architecture without all the cruft needed to run and. handle other models, that it would not be surprising to easily see 2x+ performance improvement. Anyone thinking along this idea? Again, the goal will be to give this task to a smart model, let it run in a loop, and after a week or 2 you hopefully end up with llama.qwen3.8-27b or llama.glm5.3-flash that would fly.

💬 42 (+1) open on reddit ↗
▲
9
 
8👁
r/LocalLLaMA · u/Informal-Trouble2183 · 12d ago
Hardware Roofline Inference Calculator post image

Hello everyone, I made a calculator for the theoretical HW roofline for decoding / prefill based on several parameters (LLM model architecture, quants, GPU, Memory, ..). It still a theoretical bound, but helpful as a step-0 check to understand what fits (would fit) in your hardware, and understand the effects of the contributing knots. I hope it helps. You can access it from here: https://www.ai-leaderboard.dev/ (click HW Roofline)

▲
8
 
25👁
r/LocalLLaMA · u/W61k3r · 7d ago
Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

https://huggingface.co/Wa1k3r/Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit

Qwen3.8-27B CODER — IQ4_XS imatrix · 24 GB card fit · ~262k context · MTP draft

Quantized, Abliterated, and fitted by LexiPanel. Its Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k tokens of context. It built the importance matrix from code-heavy text and kept the MTP head at Q8\_0, so --spec-type draft-mtp works without a separate draft model.

The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context.

#

File

|File|Type|Size|Inside|
|:-|:-|:-|:-|
|Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf|IQ4\_XS + imatrix|18.35 GB (17.1 GiB)|MTP head at Q8\_0, token embeddings at Q4\_K|

#

Measured speed (real use, not a synthetic benchmark)

These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself:

|Context already in the window|Requests|Decode, median|Decode, range|
|:-|:-|:-|:-|
|65k – 131k tokens|42|41.2 t/s|33.8 – 46.0 t/s|
|131k – 171k tokens|56|36.0 t/s|30.7 – 43.0 t/s|

  • MTP draft acceptance: the median is 85% (the middle half of requests falls between 75% and 94%). That works out to about 2.7 tokens per decode step at draft depth 2.
  • Prefill:
  • 387 t/s for a cold 108k-token prompt;
  • 175–183 t/s for about 4.5k new tokens added at 147k–156k depth.
  • VRAM: 24.2 of 24.6 GB in use at 245,760 tokens of context, with a q4\_1 KV cache and the vision projector on the CPU.

Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot.

#

Run it with llama.cpp

llama-server -m Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf \
-c 262144 -np 1 -ngl 99 --flash-attn on \
--cache-type-k q4_1 --cache-type-v q4_1 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \
--jinja --reasoning on --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
-b 2048 -ub 512 --cache-reuse 256

  • Context: -c 262144 is what fits next to the weights on a 24 GB card with a q4\_1 KV cache. The model's native window is 262,144 tokens. On a smaller card, lower -c first.
  • Speculative decoding: --spec-type draft-mtp drafts with the MTP layer inside this file, so no separate draft model is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235.
  • Sampling: these are Qwen's recommended settings, and they are also stored in the file.
  • Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB).
  • Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context.

#

How LexiPanel made it

  1. Conversion: the source weights were converted to a BF16 GGUF with the MTP head included.
  2. Importance matrix: computed from about 300k tokens (570 chunks) of code-heavy calibration text. Three quarters is Python source (the standard library and installed packages). The rest is technical documentation, READMEs and license texts, the kind of text a coding agent's context fills with.
  3. Quantization: llama-quantize from llama.cpp b11182 made the IQ4\_XS file with that matrix. The MTP head stays at Q8\_0 so its drafts stay accurate, and the token embeddings are Q4\_K.
  4. Fitting the card: LexiPanel's Fit planner chose the mix, quality first, for one 24 GB card at 262144 tokens of context. It took the best quality that card could afford at that context, not the smallest file.

#

Credits and license

  • Quantization, importance matrix, MTP draft setup and card fit: LexiPanel.
  • Tools: llama.cpp.

Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.

Qwen3.8-27B CODER — IQ4\_XS imatrix · 24 GB card fit · \~262k context · MTP draft

Quantized, Abliterated, and fitted by LexiPanel. Its
Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k
tokens of context. It built the importance matrix from code-heavy text
and kept the MTP head at Q8\_0, so --spec-type draft-mtp works without a separate draft model.
The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context.

File

File Type Size Inside
Wa1k3r/Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit.gguf IQ4\_XS + imatrix 18.35 GB (17.1 GiB) MTP head at Q8\_0, token embeddings at Q4\_K

Measured speed (real use, not a synthetic benchmark)

These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself:

Context already in the window Requests Decode, median Decode, range
65k – 131k tokens 42 41.2 t/s 33.8 – 46.0 t/s
131k – 171k tokens 56 36.0 t/s 30.7 – 43.0 t/s

MTP draft acceptance: the median is 85% (the middle
half of requests falls between 75% and 94%). That works out to about
2.7 tokens per decode step at draft depth 2.
Prefill:
387 t/s for a cold 108k-token prompt;
175–183 t/s for about 4.5k new tokens added at 147k–156k depth.

VRAM: 24.2 of 24.6 GB in use at 262144 tokens of context, with a q4\_1 KV cache and the vision projector on the CPU.
Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot.

Run it with llama.cpp

llama-server -m Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit.gguf \\
\-c 262144 -np 1 -ngl 99 --flash-attn on \\
\--cache-type-k q4\_1 --cache-type-v q4\_1 \\
\--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \\
\--jinja --reasoning on --reasoning-format deepseek \\
\--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \\
\-b 2048 -ub 512 --cache-reuse 256

Context: -c 262144 is what fits next
to the weights on a 24 GB card with a q4\_1 KV cache. The model's native
window is 262,144 tokens. On a smaller card, lower -c first.
Speculative decoding: --spec-type draft-mtp
drafts with the MTP layer inside this file, so no separate draft model
is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235.
Sampling: these are Qwen's recommended settings, and they are also stored in the file.
Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB).
Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context.

How LexiPanel made it

Conversion: the source weights were converted to a BF16 GGUF with the MTP head included.
Importance matrix: computed from about 300k tokens
(570 chunks) of code-heavy calibration text. Three quarters is Python
source (the standard library and installed packages). The rest is
technical documentation, READMEs and license texts, the kind of text a
coding agent's context fills with.
Quantization: llama-quantize from
llama.cpp b11182 made the IQ4\_XS file with that matrix. The MTP head
stays at Q8\_0 so its drafts stay accurate, and the token embeddings are
Q4\_K.
Fitting the card: LexiPanel's Fit planner chose the
mix, quality first, for one 24 GB card at 262144 tokens of context. It
took the best quality that card could afford at that context, not the
smallest file.

Credits and license

Quantization, importance matrix, MTP draft setup and card fit: LexiPanel.
Tools: llama.cpp.
Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.

💬 25 (+4) open on reddit ↗
▲
8
-2
15👁
r/LocalLLaMA · u/junior600 · 8d ago
What local AI model is good for game decomps/recomps?

Hello guys. Recently, there has been a boom in game decomps and recomps thanks to AI. If you look at the r/decomps and r/recomps subreddits, you can see it. They mostly seem to be using Claude or Codex.I wonder if it would be possible to do something similar with a local AI model. Could Qwen 3.8 27B Abliterated actually handle something like that locally? Does anyone have any experience with this? I don't have a particularly powerful rig (RTX 3060 12 GB VRAM and 24 GB DDR4 RAM), but I can run MoE models comfortably. Even Qwen 3.8 27B IQ3\_XXS dense lol.

Sorry for my English BTW.

💬 31 (+1) open on reddit ↗
▲
8
+4
23👁
r/LocalLLaMA · u/IngwiePhoenix · 9d ago
Penalties of PCIe generations? (2x R9700)

I just bought the GPUs after deliberating and debating for over two years. With costs not coming down any time soon and me just wanting to get this massive todo-box ticked, I decided to just YOLO it; the GPUs are the most volatile, followed by RAM, rest seems more or less stable.

But actually, RAM is one of the reasons I am unsure about wether to chose a SP4, 5 or 6 based board. I am most familiar with AMD CPUs, so that is where my tendencies lie. Unfortunately, RDIMMS are going to absolutely undress me... x.x

However, if I could stick to a DDR4 / PCIe Gen4 setup, that would save a pretty penny. Now I do not intend to offload to system memory, but even a small, single-stick of DDR5 RDIMM is stupid expensive - DDR4 is fine.

The question is: What is the penalty of PCIe Gen 4 versus 5 in regards to inference? I will be using llama.cpp with ROCm, fronted by llama-swap, utilizing both GPUs for inference and VRAM pooling (so, 64GB in total).

Thanks! =)

💬 47 (+1) open on reddit ↗
▲
8
+1
8👁
r/LocalLLaMA · u/DeliciousBelt9520 · 10d ago
Forlinx 20-TOPS M.2 AI accelerator supports PCIe cascading for local LLM inference

Forlinx Embedded has listed an M.2 AI accelerator card based on Rockchip’s RK1820 and RK1828 processors, providing 20 TOPS of INT8 computing performance and up to 5GB of integrated DRAM. The module uses an M.2 2280 interface and is designed to handle local AI inference, including large language models, vision-language models, and computer vision workloads on embedded Linux and Android systems.

https://linuxgizmos.com/forlinx-20-tops-m-2-ai-accelerator-supports-pcie-cascading-for-local-llm-inference/

▲
8
 
15👁
r/LocalLLaMA · u/sToeTer · 10d ago
Is there even an easy, seamless vision assistant program?

I read textbooks on my PC and ideally i want a program with a normal chat environment where i can just hammer in questions about what's currently on my screen. Example: I'm working on a PDF, underline or circle things... and then just type "what does this sentence mean?", you get it.

I do NOT want to manually screenshot, navigate to the folder, drag the picture into the environment and then also have to type the question. It should also naturally be aware that the conversation is about what's on the screen, so i don't have to steer it with "make a screenshot; use your vision capabilites" etc.

I tried multiple different MCP in LM Studio, none of them were great...or worked :/

Someone said AnythingLLM has this function but i couldn't find it.

Is there a good solution?

Thank you in advance! :)

▲
8
+1
10👁
r/LocalLLaMA · u/otacon6531 · 11d ago
IQ Quants still slow on P40?

I have been using Qwen 3.6:35b IQ4 via llama.cpp on my p40 and am getting anywhere between 37 - 83 tok/s (mtp is on). Prefill usually starts at 600 and slowly degrades as it continues processing so 600 for short prompts and more like 300-400 by the end of a long prompt. It hurts, but it is what my budget allows.

AI told me IQ quants are noticeably slower on the P40 and it referenced (https://www.reddit.com/r/LocalLLaMA/comments/1dmhpud/are\_iq\_quants\_slow\_o…) from two years ago, but I didn't feel it being slower when I moved from Q4 to IQ4, so...

What am I missing? Are IQ Quants actually a significant amount slower on the P40 or is this outdated information?

▲
8
 
14👁
r/LocalLLaMA · u/Dismal-Effect-1914 · 11d ago
B70 no stock/Price Increase

I bought a B70 off Amazon last week for 1300 and have been playing around with it. Today I checked and it seems like I cannot find a single one online for less than 1600 and most places dont have them in stock anymore? What happened? What drives these sudden price increases? It seems like all GPUs across the board have seen another dramatic price flux. Some 5090s I saw were going for 10k!?

▲
8
 
8👁
r/LocalLLaMA · u/marcobaldo · 12d ago
Qwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max

Hi! I'm the author of MoEspresso, which is my way of putting my own ideas about inference engines to the test. A lot of the fun has been trying different design choices, measuring what happens, and finding that several of them work well together. MoEspresso 3 runs Qwen3.8-Flash-Next on a 2021 M1 Max with 32 GB of unified memory at 12-15 decode tokens per second - provided there are no other memory hungry applications running in the background (such as browsers). During decode, experts which are already resident in memory receive a bias, but the two strongest experts according to the model are always chosen (with the default settings). I wrote about this here https://github.com/steadfastgaze/MoEspresso/blob/main/docs/cache_prior.md - I first started thinking about this after reading about Apple's AFM 3 and the instruction-following pruning work behind it (https://arxiv.org/html/2501.02086v3#abstract), but then I found this other paper (https://arxiv.org/html/2412.00099v2) which spoke about Cache-Prior. Prefill is unbiased. Even with this bias enabled by default, Qwen 3.8 Next scored ahead of Opus 4.8 xhigh and many other strong hosted solutions. Reproducible 48-question setup -> https://github.com/steadfastgaze/MoEspresso/tree/main/docs/benchmark_reproduc…. Overall scores (%) across six categories, including coding, data analysis and math: | Model | Score | |---|---:| | Qwen3.8 Flash (hosted), medium | 89.7 | | GPT-6 Sol, medium | 89.2 | | Claude Opus 5.5, medium | 86.4 | | GPT-6 Sol, low | 85.0 | | Qwen3.8 Flash @ MoEspresso, medium, Cache-Prior 2/2 | 84.3 | | GPT-6 Luna, xhigh | 81.4 | | Claude Opus 4.8, xhigh | 80.3 | | Claude Sonnet 4.6, high | 74.0 | | Claude Sonnet 5, medium | 70.9 | - I use some of Iwan Kawrakow's formats from ik_llama.cpp, with Metal execution through my mlx-iqk library. Most routed projections in this package use IQ2_K, which is not normally supported by either standard MLX or mainline llama.cpp. - KVarN K4/V4 leaves more memory for resident experts as context grows, and it is functioning extremely well with low RMS error on this model architecture. - Good defaults, e.g. automatic SSD streaming and Cache-Prior when all experts cannot fit, with the settings generally following the same rule. This is the third iteration, and I have more concrete ideas to explore, both for squeezing even more performance from Apple Silicon and for bringing the engine to Linux and AMD machines such as Strix Halo. Installation is through Homebrew, so "brew install steadfastgaze/tap/moespresso" Code - https://github.com/steadfastgaze/MoEspresso Model - https://huggingface.co/steadfastgaze/Qwen3.8-Flash-Next-MoEspressoV3 If you try it, I'd love to see your "moespresso speed" output (an intentionally quick benchmark). PS: English isn't my first language and I used an LLM to help refine this post, and AI coding tools for implementation. --- edit: some comments are reporting lower speeds (thank you for doing it) - I will investigate tomorrow and in next days.

💬 23 (+2) open on reddit ↗
▲
7
 
17👁
r/LocalLLaMA · u/PhysicsDisastrous462 · 9d ago
Follow-up: my native Rust + Vulkan Transformer training backend — 14 days later, now 14 parity-verified architectures and full PEFT

Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now.

Where the green architectures stand

When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export.

The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better.

14 architectures pass that gate today, led by the one I'm probably proudest of:

|Architecture|Scope|
|:-|:-|
|Falcon H1 / H1R|parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules|
|DeepSeek V4|causal LM|
|Phi-4 Multimodal|text backbone|
|Phi-3|causal LM|
|Kimi K2.5|text backbone|
|Kimi K3 / KimiLinear|hybrid KDA + MLA|
|GPT-OSS|causal LM incl. router bias|
|SmolLM3|mixed RoPE/NoPE + YaRN|
|Qwen2.5 / Qwen3.5 / Qwen4-Exp|dense, DeltaNet, QSA, PLE, MoE|
|Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2|causal LM|

Worst observed two-step AdamW parameter error across all of them: 1.19e-7.

Best: 2.6e-8.

For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems.

The bigger news: PEFT actually works now

In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list.

It's a real workflow now, and I've verified the full lifecycle:

  • LoRA fine-tuning with HF-compatible adapter export (adapter_config.json / adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem
  • modules_to_save — full trainable replacements for Linears, RMSNorm/LayerNorm, lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for.
  • Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run
  • Merge/unmerge, disable-adapter base restoration, and multi-adapter loading
  • A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model
  • The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate

32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0.

The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math.

A small note on the last couple weeks

I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now.

I'm doing better, though, and still managed to get most of what I wanted finished.

There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back.

Same caveats as before

This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config)

"supported text graph" ≠ "the entire multimodal package works natively."

Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation.

Repo

https://github.com/necat101/Hierarchos-Native

  • Architecture inventory: hierarchos-vulkan/README_ARCHITECTURES.md
  • Compatibility/parity record: hierarchos-vulkan/COMPATIBILITY.md
  • PEFT qualification evidence: PROGRESS_PEFT_AUDIT.md
  • CLI PEFT guide: hierarchos-native-cli/README.md

The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU.

I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs.

I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.

▲
7
+3
22👁
r/LocalLLaMA · u/bolche17 · 9d ago
Agent swarm coordination

Hello all!

Do you have any recommendations of tools for agent coordination and messaging to get them to collaborate on hard problems?

Ideally I would like a heterogeneous swarm, using local models as the workhorse and cloud models for reviewing, coordination, or simply to avoid overloading my relatively small local setup.

Do you have any recommendations or experience with this?

💬 24 (+2) open on reddit ↗
▲
7
 
13👁
r/LocalLLaMA · u/js1943 · 11d ago
LM Studio vs Bionic

I am confused between LM Studio and Bionic.

I have LM Studio for a long time though not used frequently.

Recently I am trying to learn the agentic stuff. Watched a few videos but they were all using Bionic. The strange thing is the interface looks different than the one I just installed today. (Mine seems to be missing features, no developer mode. I am on MacOS)

On the other hand, I seem to be able to find those missing settings in LM Studio.

So what is the difference between the two? Is there anything Bionic can do but LM Studio doesn't?

💬 9 (+5) open on reddit ↗
▲
7
+2
6👁
r/LocalLLaMA · u/Then_Blueberry7290 · 12d ago
LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF

Just recently stubled upon with this modell:LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF I'm just stay away from "magic" models, but this model size got my eyes on: With vision capabilities this is under 17GB, which means i can use it 32GB vram with full Context size (262k), bigger ubatch, and mtp4. Of course vision goes to ram, not gpu. Other similar model with nvfp4 line, usually 19-20GB in size or more. I tried in with llama.cpp, speed is 40-113 t/s (76 in my benchmark) with 262k context. Under normal agentic workin it is 45-65 t/s. (2x5060ti16GB OC) For example thinkingcap nvfp with vllm i can only have 160k context (cannot offload mmproj to ram) First glance it is the same as the other swift models (nvfp4) in quality. So My question is what is the tradeoff of this modell?

▲
7
-3
12👁
r/LocalLLaMA · u/poofph · 12d ago
Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task.

I am new to all this so still a lot to learn. If I give Qwen3.8 27B a coding task to fix some bugs in some code, it went through and found and fixed several in like 5 or 10 minutes. I gave qwen flash next the same task and 2.5 hours later it was done. 27B is of course faster overall to run on my system (dual rtx 5090) with infill \~2000-3000 and output 100-150 tok/s, flash next \~1600-2300 infill and 60-100 tok/s but a huge difference in the time it took to complete the task. What is the reason for this and what settings would get flash next to complete in similar time frame as 27B? For instance this last job I gave Flash Next, I checked the time when it modified the code/put in the fixes, it completed the fixes \~2 hours before it was finally done running tasks, so for 2 hours it was running tests or who knows what and never modified the updates anymore after that point. Also, fyi - (I am not a programmer, these are programs that were created by AI and I ask for fixes/updates and let it do its thing).

▲
7
+2
8👁
r/LocalLLaMA · u/arbv · 12d ago
Improved chat template for Laguna XS / S 2.1 (configurable forced thinking, preserve_thinking toggle, and stability fixes)

Following up on my previous post about the GPT-OSS template, here is an updated chat template for Poolside's Laguna models (XS and S 2.1). The main reason I ended up putting this together was inconsistent reasoning. By default, the model is supposed to decide when to think on its own, but in practice it's pretty lazy - especially the XS variant - and often skips thinking right when it needs it most. When these models do think, they do so well. All in all, a good models to have around. Also they write well in English (to my non-native eye, at least). Laguna XS 2.1 in particular deserves more attention, IMO. What I like about these models is that allow toggling reasoning mid-conversation without invalidating the prefix cache. Very handy. I added a force_thinking toggle (using the prompt trick discovered by u/SnooPaintings8639) to make it think on every turn, plus a reasoning_effort parameter (none, auto, max) if you prefer an easy preset over juggling booleans (and to make it easier to use in Pi and, possibly, other harnesses). I have fixed some other things along the way. Firstly, the preserve_thinking toggle. The original template permanently forces historical reasoning preservation on. That's great for prefix cache and agentic tool loops, but if you're just having a normal chat, dragging thousands of past reasoning tokens around shreds your context window fast. You can now turn it off. Secondly, I added basic validation to catch smuggled control tokens across roles (can be turned off via allow_injection: true). By default, all settings match the upstream behavior (enable_thinking=true, preserve_thinking=true, force_thinking=false), so it acts as a direct drop-in replacement if you don't want to mess with the new knobs. Template repo: https://huggingface.co/arbv/laguna-2.1-fixed-jinja-template I also included some recommended sampling settings for llama.cpp (BF16 and quantised) and a config snippet for Pi (models.json) in the README. P.S. Also casting u/matthiasgalle (the Laguna post-train lead) to take a look and for a data point.

▲
6
+1
14👁
r/LocalLLaMA · u/indiealexh · 9d ago
How to make best use of a Intel Arc B70?

I have a RTX 5090 in my desktop PC for local coding assistance and gaming and I have been loving it with Qwen3.8 27B Q4\_K\_XL.

I managed to get a B70 on the cheap and its great, but using it in split mode with the 5090 to ensure I get full context halfs my T/s (which is expected due to the memory bandwidth).

Would I be better off running the B70 with a smaller model to offload tasks to? Or just accepting the slower throughput and keeping the larger context?

I'd especially like to hear for anyone who has a similar mismatched GPUs setup.

💬 11 (+1) open on reddit ↗
▲
6
-1
16👁
r/LocalLLaMA · u/lucasbennett_1 · 9d ago
on prem LLM stack for data that cant leave the building

Running the model locally is not a problem thats easy part but the leaks are the third party integrations along with it, like you designed everything perfect and then just added a cloud api along with it maybe a hosted judge for evals or a tracing saas or embedding point. one http call and the on prem things over

parts we already keep local are

  1. runtime: llama.cpp/ vllm /ollama
  1. models: qwen or llama family depending on rig
  1. vector db: pgvector or qdrant

some that leak but remain unnoticed:

  1. ingestion: pdfs and scans for some projects need a parse and the ocr step before chunking them and its where we often reach for a cloud parser and break the rule, although we can keep it local with liteparse sort of inbound parsers or other open source options on huggingface
  1. Eval: plenty of local setups still need prompts and outputs to a hosted judge or a tracing dashboard to see quality which is the same leak but seems different. instead a  local score set or a local judge model and keeping it self hosted if possible handles the tracing part

I am curious to know about others end to end stack who keep it 100% local, eager to learn more

💬 27 (+3) open on reddit ↗
▲
6
+1
12👁
r/LocalLLaMA · u/Gold-Bat-3225 · 10d ago
Does post training make LLMs funnier? post image

We did a study: does post training actually make LLMs funnier?

We used open models that publish every stage of post training, so we could compare a base model with the future models it became: Tulu 3 (on Llama 3.1 70B), OLMo 3.1 32B and Qwen2.5. We tracked 11 stages, 100 joke prompts, 64 human raters and 2,330 head-to-head judgments.

What we found: post training makes models funnier, but reduces diversity of response.

\- In 5 of 7 training steps, the later model's jokes were judged funnier. Jokes also got 10–20 words shorter after early post training, so they get to the punchline faster.

\- In 6 of 7 steps, the jokes a model wrote for the same prompt got more similar to each other. Ask for eight jokes on one premise and you get eight versions of the same joke. The biggest drop was Qwen2.5 base to instruct.

\- Asking the model to plan a line or two before the joke cut variety in all 4 models we tried, with no reliable gain in funniness.

\- A comedian persona won back a little variety in all 4 models, but only made the jokes funnier in 2 of them.

Humans judged the base versus final. A model judge calibrated on those votes compares the stages in between.

Full report and paper below. Which open models should we run through this next?

https://laugh.so/research/humor-tax/

💬 12 (+1) open on reddit ↗
▲
6
-2
8👁
r/LocalLLaMA · u/Anony6666 · 12d ago
Introducing CyberPVP: CyberKimi vs. ALTAR-1 on 100 CyberGym tasks, with live traces and public results

Trying something new - introducing CyberPVP - in other words CyberKimi vs other AI models competing to solve complex cyber tasks. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Welcome to CyberPVP! you can see it live here: Live we randomly picked 100 tasks from CyberGym, and we run two models competing at the same time, we provide the traces live as both models compete, and we also upload these traces to GitHub once the challenge finishes so that they can be verified independently. For this first public run, we choose Aikido Security model ALTAR-1 to compete with CyberKimi on 100 CyberGym tasks. Note: for ALTAR-1 we shipped it with 128K context behind a 8xH200 (two 4xH200 with load balancer) - we also followed their hugging face model card and deployment instruction/configuration available here: Hugging Face you can current watch the live run here: Live is here Traces and results uploaded after each run here: Github Challenge rules and conditions: Conditions Source : X

▲
6
 
8👁
r/LocalLLaMA · u/uBazzyZ- · 12d ago
Prevent CUDA OOM in PyTorch with dynamic lane switching

I built MEM v3 to solve a frustrating problem in PyTorch: CUDA Out-of-Memory crashes during long training and fine-tuning runs. Instead of restarting when memory spikes or keeping batch sizes overly small just to be safe, MEM acts as a memory governor. It watches VRAM and throughput in real-time, then dynamically adjusts batch size and gradient accumulation on the fly without stopping the process. What it does: \- Dynamic lane switching: Scales batch size up or down in milliseconds based on actual GPU memory pressure. \- Chaos resistance: Tested against sudden +10 GB VRAM allocation shocks without crashing. \- Crash-proof checkpoints: Uses atomic file replacement with SHA-256 checks across rotating slots, so power outages won't corrupt saved weights. \- Live telemetry: Built-in local web dashboard to track loss, throughput, and lane switches. You can test it directly on a free Colab GPU without setting anything up locally: https://colab.research.google.com/github/nobazzy/mem-llm-orchestrator/blob/main/notebooks/mem\_orchestrator\_interactive\_demo.ipynb Repo: https://github.com/nobazzy/mem-llm-orchestrator Would love to hear your thoughts and feedback!

▲
6
 
8👁
r/LocalLLaMA · u/Fz1zz · 12d ago
Qwen3.8-27B FP8 dual GPUs

Hardware RTX 5090 (32 GB) + RTX 4070 Ti Super (16 GB, PCIe x1) = 48 GB VRAM 32 GB DDR5-6200, Arch Linux, KDE on the 5090 Setup Huihui Qwen3.8-27B abliterated INT8 W8A16 + DFlash2 drafter (K=7), vLLM 0.30.0, pipeline parallel: 4070 Ti Super: vision encoder, layers 0-20 5090: layers 21-63, lm_head, drafter 262K context, FP8 KV, 2 slots. Benchmarks (single request, thinking off, fresh context per depth) |Depth|Prefill|TTFT|Decode (code)|Decode (prose)| |:-|:-|:-|:-|:-| |2k|2,922 t/s|0.7 s|184 t/s|72 t/s| |32k|3,001 t/s|10.7 s|145 t/s|70 t/s| |62k|2,698 t/s|23.0 s|152 t/s|71 t/s| |92k|2,444 t/s|37.7 s|157 t/s|63 t/s| |122k|2,229 t/s|54.8 s|147 t/s|68 t/s| |152k|2,051 t/s|74.2 s|133 t/s|67 t/s| |182k|1,900 t/s|95.9 s|143 t/s|61 t/s| |212k|1,771 t/s|119.8 s|150 t/s|62 t/s| |242k|1,658 t/s|146.0 s|138 t/s|60 t/s| |260k|1,596 t/s|163.0 s|134 t/s|56 t/s| Code decodes faster because the drafter's guesses are accepted ~70% of the time vs ~22% on prose. Follow-up turns hit the prefix cache (1.3 s TTFT at 260k). Needs patched vLLM, see repo. The 4070 Ti Super sat collecting dust for two months because I assumed PCIe x1 would kneecap it. Apparently not. My full setup: https://github.com/ExTV/dual-gpus-vllm

▲
5
+2
17👁
r/LocalLLaMA · u/Fz1zz · 7d ago
QFN at 262K on 32 GB RAM 48GB VRAM : 1,800 tok/s prompt, 130 tok/s decod three Strata patches.

Qwen3.8-Flash-Next (IQ3\_XXS, 76 GB) at the full 262K context on 31 GB of RAM: 1,800 tok/s prompt, 130 tok/s decode, on a 5090 + a 4070 Ti SUPER on a PCIe x1 slot

Strata (github.com/Niko1221/Strata) streams MoE experts from an mmap'd GGUF, and its own sizing rule says RAM >= expert shard + 10 GB, so 57 GB for the ISTA GSQ-RCO IQ3\_XXS. I run it on 31 GB and it is fast now. Hardware: RTX 5090 32 GB + RTX 4070 Ti SUPER 16 GB (the small card sits on a chipset x1 slot, 0.8 GB/s), i7-14700K, one NVMe.

Stock 0.1.33 with a layer split at 36: 80K-token prompt 890 tok/s with the 4070 at 100% and the 5090 idle, decode \~110 tok/s, and switching between two chats re-reads the other one (30K tokens = 48 s).

Three patches on v0.1.33 (repo below, they apply to a pristine checkout):

  1. Prompts run entirely on the big card (port of Strata PR #269), the small card gets its layers' state copied afterwards, only the cells in use. Conversation parking works with the split, so alternating between a phone and a desktop session takes 0.6-1.2 s instead of 20-48 s.
  1. The real bottleneck on a box with less RAM than the expert file: the expert pool copied each 2 MB expert out of the mmap with memcpy after a MADV\_WILLNEED hint. Under memory pressure the kernel drops that readahead and you pay one major page fault per 4 KB, hundreds per expert, while both GPUs wait. A pread per slice instead: major faults per benchmark run went from 24 million to 14 thousand.
  1. Bigger prompt chunks (--prefill auto:32768). Every chunk re-streams every layer's experts, so four times fewer chunks matters a lot when the working set does not fit in the page cache.

Numbers on the final config (int8 KV, 32K cells resident per layer, MTP spec 4, vision on, 262,144 context):

\- 80K fresh prompt: 1,796 tok/s (45 s); 16K: 985; 2K: \~350

\- follow-up turn on an 80K conversation: 6 s

\- decode: 128-134 tok/s median on real sampling (0.6 / 0.95 / 20) with --spec-min-p 0.8, 95-108 greedy

\- 40K needle + follow-ups and a two-conversation parking test all correct

\- VRAM: 5090 at 32.0 GB, 4070 at 15.7 GB; RAM: the engine \~5 GB, the rest page cache

Repo with the patches, install script, launcher, benchmark tools and all measurements: https://github.com/ExTV/strata-5090-4070

▲
5
+2
20👁
r/LocalLLaMA · u/giveen · 7d ago
GitHub - giveen/KernelOPT: Dispatch-aware agentic GPU kernel optimization

I want to share something I've been working on, and research paper from Redhat really helped.

This is my agentic GPU kernel optimizer for inference engine development

Its whole goal is to look at inference engine kernels, and using cloud models, plan, test and find improvements. Nothing is changed in your git, it provides a diff, send the diff over to your coding harness and ask it to analyze the diff.

It has a setup wizard and a run wizard which I recommend using, please give me feedback on what needs to be improved, as the AMD stuff I was unable to test, and mostly is theory at this time.

▲
5
+1
12👁
r/LocalLLaMA · u/No_Contract_8296 · 7d ago
CalDec v1 - Fully Open Decision Model for Personal Assistants

Somebody just released a fully open-source, open-weights decision model that beats Jev!

Just kidding, it's me and I this is my first time releasing a public model, recipe and dataset so I am looking forward to learning from the experience.

Jev is indeed a very powerful and inexpensive model and obviously a much better all-rounder, and some of my checkpoints did in fact score better on some tests (namely LocalLLaMA/typed-decisions and the internal test set) but that doesn't mean it "beats Jev" of course.

The motivation for this was a quick experiment to see how far behind Jev open-weights models like Laya are, and how much closer I can bring them with a small dataset and fine-tuning. The results were better than expected especially for me since I do not have professional ML experience.

For my use-case - a Jarvis-like personal assistant which aims to be real-time and fully-local - this model proved to be genuinely useful for certain aspects of that project so I decided to share the results and how I got there. Going local also means privacy and eliminating network latency.

I hope some of you find this experiment valuable or useful in some way.

I would also love to hear you suggestions, criticism or just discuss the approach!

Dataset: https://huggingface.co/datasets/kgrozdanovski/assistant-decisions**
CalDec Laya: https://huggingface.co/kgrozdanovski/caldec-v1-laya**
CalDec GLiNER: https://huggingface.co/kgrozdanovski/caldec-v1-gliner2.5-decide**
GitHub: https://github.com/kgrozdanovski/caldec**

▲
5
-1
13👁
r/LocalLLaMA · u/Competitive-Scar-627 · 7d ago
Model weight inferencing

I have 4050 6gb gpu, 24 gb ram which model should i choose to run i need speed. i try qwen 3.8 27b and feel too slow tried from onslot studio.
I have heard of weight inferencing does it helpful what should i do to try weight inferencing.

💬 23 (+2) open on reddit ↗
▲
5
+1
19👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 8d ago
Computer use powered by local/cloud models for regulated industries?

The local models seem powerful enough to be capable of running basic local computer use. This is a computer use agent/harness built via claude code, and powered by qwen3.8 flash next nvfp4. Drawing a simple image of its choice took 12m 1s, but a lot of that time was spent by the agent trying to figure out a WebGL bug. For what it did, it seems relatively fast. Prefill speed \~1000 tps, gen speed \~50 tps. I went to openai’s devday a couple days ago, and it seems like cloud models that are very smart and fast, like “astra ultrafast”, can perform work even quicker, for a premium.

Does anyone have experience with computer use agents/harnesses that are open-source and plug-n-play, that are robust enough to be used in regulated fields such as law/medicine? How do people deal with regulations, such as making such workflows HIPAA compliant in medicine? Experiences with helping users ensure that workflows are completed accurately? And whether they go with local or cloud models to power computer use?

💬 7 (+1) open on reddit ↗
▲
5
+2
18👁
r/LocalLLaMA · u/Perfect-Campaign9551 · 11d ago
Recommended way to run Qwen 3.8 27b on a 3090 in Windows?

I know this may have been asked a lot, but I'm not an LLM expert yet (have never set up vllm myself or llama.cpp myself yet) I've only used scripts other people have set up.

Is there a simple way / steps to follow to run Qwen 3.8 27b in Windows with my single 3090?

I can run it straight up with Ollama with a 64k context but it seems like it only works reliably in chat and not in OpenCode (in OpenCode it "works" but at one point it "hung up" on me it seemed like. Not sure if maybe it was busy thinking still)

I've found quite a few threads that are close to what i'm asking for but many I think use WSL or something, too. Which I'm also not super experienced with yet.

I found the HyperQwen repo but their documentation is like...120% all technical and not very user friendly at all. I CAN do technical stuff! But it's barely passable as "do this, and then this" type of docs at the moment.

Ninfer is only for 5090 cards from what I read.

EDIT: Thanks guys, I was able to use llama-cpp-windows-manager project to get Qwen 3.8 27B up and going (Q4\_K\_M) . I have a 98K context and get 70tok/s and it's working with Open Code just fine. Very usable.

💬 34 (+2) open on reddit ↗
▲
5
-1
11👁
r/LocalLLaMA · u/PilgrimofHaqq2 · 12d ago
The pelican test on MiMo 2.6: with and without plan mode
  • Plan runs settled style/scene/size in one Q&A round, then wrote the whole SVG in a single call (11.1KB Flash, 22.1KB Pro) and batched every render fix into one edit round. That's 12 and 17 calls total. No-plan runs iterated more: Flash did 3 render-fix rounds and lost \~10 calls to image verification (crop reads coming back mismatched, zoomed views, one stale preview render). Pro did 2 fix rounds plus 4 tool mishaps, one of which generated 3,743 tokens and threw them away (edit call rejected for a missing arg). Generated tokens don't follow the totals: Flash plan generated MORE than Flash no-plan (27.2k vs 20.4k). Fewer, bigger calls, not less work.
▲
4
+2
16👁
r/LocalLLaMA · u/failuremap-f · 8d ago
Can your local coding model repair these boundary-case bugs? Failure Map: 20,168 open Python tasks

I’m the creator of Failure Map, an archive of compact Python debugging tasks. The open release has 20,168 tasks across 254 categories, with standard-library implementations, explicit contracts, failed repair attempts, and executable boundary checks.

Three small cases to try:

• Duplicate delivery: deduplicating equal amounts loses legitimate events. https://failuremap.org/cases/FA-001

• Cache expiry: subtracting a whole tick rejects an entry that is still valid. https://failuremap.org/cases/FA-006

• Pagination: changing > to >= repeats the cursor record. https://failuremap.org/cases/FA-011

Prompt template: “Repair the solve function to satisfy the stated contract. Return Python source only. Preserve the signature. Contract: {prompt}. Broken implementation: {broken\_source}.”

Measured program baselines, passed checks out of 3 (broken / attempted repair): FA-001 2/3 / 1/3; FA-006 2/3 / 2/3; FA-011 2/3 / 1/3. These are executions of the included programs, not model scores. I have no measured local-model results to claim yet.

To compare runs, report the exact model and revision, quantization, prompt, sampling settings, seed, attempts per task, and pass counts. Run candidate code in isolation and keep grading fixtures outside its control. Recorded-check success is not hidden-test performance.

Download: https://failuremap.org/api/exports/tasks.jsonl.gz

Methodology: https://failuremap.org/methodology

▲
4
+1
11👁
r/LocalLLaMA · u/SignatureMoney6648 · 8d ago
FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.

I've run the benchmark on a RTX 3090, 1024 tokens in / 256 out, concurrency 1–32.

If the model fits on vRAM (Gemma-4-26B-A4B, byte-identical GGUF on both engines): llama.cpp has 2.2–3.2× the throughput and 5–6× faster TTFT. FreeToken 0.1.2 can't keep 4-bit experts in VRAM at all, and it OOM'd at 8 concurrent.

If the model doesn't fit (gpt-oss-120b, 63 GB): FreeToken's TTFT stays at \~9 s from 2 to 8 users while llama.cpp's goes 17 → 58 s. At 32 users it's 19 s vs 139 s. Throughput is basically a tie (10–17 tok/s for both).

Spilling to system RAM costs \~10× in generation speed whichever engine you use.

FreeToken's PCIe link sits at its ceiling the whole time, so PCIe 4.0 should help it a lot (I've run this on a gen3 motherboard).

So from this test FreeToken only makes sense with many concurrent users in models that cannot be hold inside vRAM. But I am not sure if that is always the case or an artifact of the gen3 bottleneck on my PC.

Has anyone run a benchmark like that with a gen4 Motherboard?

Full details on the link. BTW: I used AI to generate the charts and correct my spelling and grammar.

▲
4
+1
13👁
r/LocalLLaMA · u/fallingdowndizzyvr · 9d ago
What runs Qwen 3.8 Flash Next faster? Strix Halo or a Pile of GPUs(2x5070tis, 2x7900xtxes and 2x5060tis 16GB).

I have a machine with a bunch of GPUs attached to it. 2x5070tis, 2x7900xtxes and 2x5060tis 16GB. So I did this little test to see how it fares running Qwen 3.8 Flash Next Q4_XL against my little Strix Halo. Not well. Not well at all. The full numbers are below but the high context number sums it up.

@160,000 context

Pile of GPUs 215.73(PP) and 16.29(TG)

Strix Halo(Gufo) 1227.12(PP) and 22.04(TG)

Here's the number for a Strix Halo fork of llama.cpp, Halo Box.

Strix Halo(Halo Box) 587.99(PP) and 21.24(TG)

Lastly, here's the mainline llama.cpp number.

Strix Halo(llama.cpp 0.4.1) 113.48(PP) and 7.06(TG)

For running QFN, Strix Halo really shines.

Pile of GPUs

Device 0: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0, VMM: yes, VRAM: 15880 MiB
Device 1: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0, VMM: yes, VRAM: 15880 MiB
Device 2: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
Device 3: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
ggml_cuda_init: found 2 ROCm devices (Total VRAM: 49120 MiB):
Device 0: AMD Radeon RX 7900 XTX, gfx1100 (0x1100), VMM: no, Wave Size: 32, VRAM: 24560 MiB
Device 1: AMD Radeon RX 7900 XTX, gfx1100 (0x1100), VMM: no, Wave Size: 32, VRAM: 24560 MiB
| model | size | params | backend | ngl | fa | dev | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ------------ | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 | 242.13 ± 1.39 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 | 32.09 ± 0.06 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d10000 | 242.96 ± 1.12 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d10000 | 30.23 ± 0.14 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d20000 | 249.34 ± 0.65 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d20000 | 28.84 ± 0.05 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d40000 | 253.95 ± 1.44 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d40000 | 26.29 ± 0.07 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d80000 | 252.48 ± 1.20 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d80000 | 22.34 ± 0.07 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d160000 | 215.73 ± 0.56 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d160000 | 16.29 ± 0.03 |

Strix Halo running Gufo

| model | size | backend | test | t/s |
| -------------------------------- | ---------- | ---------- | ------------------ | --------------------- |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 | 1603.47 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 | 26.53 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d10000 | 1377.97 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d10000 | 25.44 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d20000 | 1353.19 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d20000 | 25.01 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d40000 | 1328.64 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d40000 | 24.08 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d80000 | 1283.77 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d80000 | 23.00 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d160000 | 1227.12 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d160000 | 22.04 ± 0.00 |

Strix Halo running Halo Box

Device 0: AMD Radeon 8060S Graphics, gfx1151 (0x1151), VMM: no, Wave Size: 32, VRAM: 128000 MiB
| model | size | params | backend | ngl | fa | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 | 811.79 ± 19.09 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 | 23.95 ± 0.03 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d10000 | 729.60 ± 38.99 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d10000 | 22.46 ± 0.43 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d20000 | 718.14 ± 33.92 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d20000 | 22.45 ± 0.18 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d40000 | 700.55 ± 33.17 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d40000 | 22.29 ± 0.23 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d80000 | 646.85 ± 29.06 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d80000 | 21.95 ± 0.25 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d160000 | 587.99 ± 30.10 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d160000 | 21.24 ± 0.31 |

Strix Halo running llama.cpp 0.4.1

Device 0: AMD Radeon 8060S Graphics, gfx1151 (0x1151), VMM: no, Wave Size: 32, VRAM: 128000 MiB
| model | size | params | backend | ngl | fa | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 | 362.84 ± 6.19 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 | 20.11 ± 0.36 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d10000 | 321.75 ± 0.91 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d10000 | 19.88 ± 0.34 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d20000 | 285.78 ± 0.42 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d20000 | 18.35 ± 0.46 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d40000 | 237.23 ± 1.13 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d40000 | 15.48 ± 0.61 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d80000 | 173.80 ± 0.18 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d80000 | 11.18 ± 0.11 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d160000 | 113.48 ± 0.23 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d160000 | 7.06 ± 0.17 |

💬 52 (+1) open on reddit ↗
▲
4
+1
9👁
r/LocalLLaMA · u/HornyGooner4402 · 9d ago
Pi + llama-server randomly hung

I can't seem to find what's wrong. I'm using Pi for my llama-server and sometimes it just stops processing for some reason and stuck after tool call. Logs seems to think that it's finished its job while Pi thinks it's waiting for a response, so sometimes I have to stop it and tell it to "Continue". This only happens occasionally, 99% of the time it works with no problem. Anyone experienced something like this?

Edit: Just realized I was vagueposting. Running Qwen3.6 35B A3B IQ4_NL_XL from Unsloth, but I think it happened with other models as well.

▲
4
+1
7👁
r/LocalLLaMA · u/whoami-233 · 11d ago
Anyone running multi GPU A100 80GB Cards?

Hey guys, I am looking for benchmarks for people running multi node (4 or more) A100 80 GB cards and seeing what results and models they are getting. Something with VLLM and multi users would be very useful. Or if you know of a place I can find such results please let me know! Appreciated!

▲
4
-1
14👁
r/LocalLLaMA · u/Express_Quail_1493 · 12d ago
Qwen3.8FlashNext Please Share Cold prefill at compaction 128k

I see Many people sharing amazing decode speed and prompt-prefill(PP) speed but no one is sharing their prefill speed when the harness is compacting a COLD prefill please. can you share your partial offloading COLD prefill speeds at long context? I would like to run flash next but can’t spare the network download ATM but looking to bite the bullet if its absolutely worth it? Pretty please help.

▲
4
+1
6👁
r/LocalLLaMA · u/bulletrhli · 12d ago
Power Limits, Local AI, and Questionable Uses of My Free Time

Edit 1: Okay, I have been checking out Unsloth and wow. Just wow. Thank you so much for your suggestions. This is such a way better tool and I am going to go crazy with this. Good day data nerds! I am trying to get more into running models, learning about agentic workflows, and creating my own tools. But as you do (right?) I had to fine tune my current setup. With the way the markets are right now, it only makes sense to make the most out of what I got. My day job, typically, is around data, numbers, and programming; only two of those I am good at, I'll let you guess which ones. So, yes, here come some data sheets and pretty graphs. Don't worry, you don't have to go through the data, but you can if you want. The graphs cover key metrics spat out by Ollama such as the tokens per second, duration, eval rates etc. I also added a cheeky "tokens/s/W" which, technically is not perfect since I do not measure wattage over time, but I did observe the watts during prompts, and I have a few things to mention about that later. Okay, let's start off with the specs because you probably think I am rocking the good stuff since I am so invested in this topic (haha) Lenovo M920Q 16GB DDR4 2660MHz Intel i5-8500T (6C/6T) Gigabyte Gaming OC 3070 8GB (Over OcuLink at Gen3x4 speeds) I am running OpenWebUI with Ollama in an LXC on my Proxmox server. This is one of my nodes and it is dedicated to my models. I have given it all of the cores, 14GB of RAM, 2GB swap. Nothing crazy to write home about, see? Okay, so, one of the things I wanted to know what, with the models that I run day to day, how effective are they at different GPU power limits. Man, if only I had known how much of a rabbit hole I would go down to do this (sorry wife). I only run 4 models, nothing too crazy, until you run 3 tests per model, for each power limit from 100 to 220 (156 runs in total), and each run of each model taking around 3 or 4 minutes since I have to unload the model each time to not have any prompt caching. Afterwards I would average the results and add that to the sheet. Really gave the fingers a workout since I now am a proud owner of a 60% keyboard for the first time and I no longer have a numpad... I'll remember that for next time. That being said, the switches are soooo creamy, a valiant tradeoff. So what models am I running? Glad you asked. The 3070 does limit me quite a bit, but with so many models available and so many smarter people than me who can quantize the models, I have found these models fit my needs. For the most part everything runs in the VRAM, except for 2, but those come with asterisks. gemma4:e4b qwen3.5 qwen3-vl\ deepseek-coder-v2\ For the vl model, it runs really well at a 23% CPU to 77% GPU ratio. Totally fine for my purposes. As for deepseek, it is a 40/60 ratio, but I luck out as it is a mixture of expert's model but even with the ratio, it is extremely performant. Gemma is by far my best model, and I have the most context room available at around 16k whereas the remainder I have sitting at 8k. Both gemma and qwen3.5 fit entirely in my GPUs VRAM. A couple things I noticed: Gemma4, is so good. Doesn't overthink, understands the prompt, remains as concise with the right tone I want. A really good day to day general model to work with. I also love the extra headroom for the context. Qwen3.5, a heavy thinker. Whilst it does a great job on the output, it spends a lot of time thinking and generating a lot of tokens. Power usage is pretty good, broke around 203W at one point and anything below that it just sat at whatever the power limit was set to. Qwen3-vl, also a major over-thinker. It spends so much time thinking that it balloons the context. I probably do not understand how to use this very well because when reading its thoughts it knows the answer pretty early on but it just gets into a thought trap (eh-yo). It does always output the correct answer, or the best it can, but I might move away from a reasoning vision model and stick to traditional ocr. If you have a better model or know how to prompt this better, I would love the help. Oh, one final note, this model LOVES power. Always maxes whatever I have, not that it increased performance directly, but it just loved power. Deepseek-coder-v2, this model rocks. It is extremely performant even though I technically on paper can't fit it. Especially for smaller asks with good bounds in place, it doesn't think, it just does and gives me excellent code back. I have yet to make it build me anything bigger but that is something I will experiment more with later. It is a weird one though, consistently using a fraction of the power budget available to it. Under 150W power limits, not once did my fans kick in on my GPU. Even my CPU fans (which are mucho loudo) rarely turned on, or if they did, they were not sounding like rocket engines. Not sure if those two are related but, eh, just something I noticed. Deepseek I think had some anomalous results with some spikes, but I can't be arsed to do them again. For the most part the results are fairly consistent and show a trend. Same for the vision model by qwen, oh well. My thoughts? It probably doesn't matter too much for most of us on a budget. Just let her rip, but if you want to shave off some heat, just lower your power down a little bit and monitor your temps. For the most part, you are probably fine. Honestly, it is 3am at this point, have a look at the spreadsheet! It was a lot of fun (I think) doing this. Interesting observations were made where I can balance my power limits, save... well pennies, and not have to listen to fans. So, works for me. https://docs.google.com/spreadsheets/d/1CEAr40nemlsK727QMvBnPcJDg548s-Bx/edit?usp=sharing&ouid=105501696463520933058&rtpof=true&sd=true Managers love graphs

▲
4
+3
8👁
r/LocalLLaMA · u/fgoricha · 13d ago
Dual 3090 stability troubleshooting

&#x200B; I made a previous post about my x299 stability issues. Seems another stability issue has popped up since then, but overall has been much stable. Seems to only happen when my i9 is working hard the dual 3090s are also working hard at the same time. Specs: EVGA X299 FTW K \\Intel i9-7940X 64 GB RAM (4 × 16 GB) 2 × RTX 3090 Founders Edition ASRock 1600 W PSU Each GPU installed in its own x16-length PCIe slot Roughly one slot of space between the GPUs Originally, I was running 128 GB (4 × 32 GB). With both GPUs under sustained AI workloads, the entire computer would eventually hard-lock: display signal gone, network connection gone, no apparent activity, but fans/lights remained on until I held the power button. I switched to 64 GB using 4 × 16 GB and that seemed to resolve that particular stability problem. The board is supposed to support the 4 × 32 GB configuration with the latest BIOS, but apparently my system wasn't happy with it. Now the next problem....... Each RTX 3090 is stable individually at PCIe Gen 3. However, when I run both GPUs together under heavy load, particularly while the i9 is also being heavily utilized, I still get instability with PCIe set to Gen 3. Hard locking with FF displayed on the mobo. Have to hard restart and boots fine into Windows. If I manually force the PCIe slots to Gen 2, the system appears to be stable with both 3090s and the CPU working simultaneously. So my question is: For AI/ML workloads, how much performance am I realistically giving up by running the two 3090s at PCIe Gen 2 instead of Gen 3? Obviously I'm going to benchmark my actual workloads both ways, but I'm interested in other people's experience. Most of my work is inference/training where the models and batches are primarily staying in GPU VRAM rather than constantly transferring huge amounts of data across PCIe. I'm also curious what the Gen 2 stability might point toward. Since either GPU works individually at Gen 3, but dual-GPU Gen 3 becomes unstable under heavy CPU/GPU load, could this indicate a motherboard/PCIe signal-integrity issue, CPU PCIe controller issue, BIOS setting, or something else specific to X299? Any ideas for additional troubleshooting would be appreciated. TLDR: How much performance am I losing using gen2 pcie vs gen3 pcie? Edit for additional info: Using Windows Using llama.cpp Stability issues happen when power limited at 200W and no power limiting More edits: Open air case Used a variety of diagnostic tools including OCCT, MemTest86, and HWiNFO64 to test each part individually. Seems to lock up when cpu and both gpus are going at 100%. Even locks up if cpu and one gpu is going at 100% while the second gpu is idle. But oddly, no problems in the same scenario but with the second gpu removed from its pcie slot

▲
3
+1
15👁
r/LocalLLaMA · u/Thac0-is-life · 7d ago
Help me with a better hardware setup for Local LLm

&#x200B;

Hello. I've been playing around with local LLMs for a while now, using my 7900xtx. I understand the concepts and usually what to do. But I'm now in a bit of choice paralysis on where to go next.

I have a Ryzen 5600X with a 7900XTX and 64GB of DDR 4 (how I wish I had purchased more at the time..) and a motherboard (MS-7B79/X470 GAMING PRO (MS-7B79)) that is not really great for multiple GPUs (it was a gaming PC).

I would like to increase my Local LLM game. I'm running mostly qwen 3.8 27B at 3 or 4q from with 128k to 220k context with KV cache at 8q. Sometimes I get up to 50tk/s and around 750 tk/s of context ingestions. But I wanted to run bigger models/have faster speed, or at least run multiple copies of that same qwen so multiple agents can run at the same time. Or try the Qwen 3.8 Flash for example. This level of model is already awesome enough to do anything I need.

I've been thinking of purchasing 2 or 4 MI50 16GB( which costs 1/3 of the 32GB), but I don't really know the rest that I should get. Motherboards that would help me optimize the performance around that, etc.

Or should I just bite the bullet on another 7900xtx (more expensive than 4 MI50)? But I think I would still need a new motherboard at least to let me use both at the same time

I have a basement so noise is not a problem, and I have solar, so power is not a real issue (at least during summer).

What are you all suggestions here? Does it make sense to go with older GPUs like that?

▲
3
+2
18👁
r/LocalLLaMA · u/dh7net · 9d ago
I need help to benchmark harness/model/hardware combination.

Hey! I'm trying to build the ultimate leaderboard to help everyone find the right harness/model combination given their hardware. (With all model variations and inference engine).

I own a GX10 and one 5090. And I'm trying as many thing as I can. (Happy to test anything, just let me know).

But I can't test hardware that I don't have.

So my ask is simple: Can some you do some testing on your own hardware?

I made this as easy as it could be: you just have to copy a prompt to your agent and your agent will fetch the test, pass the benchmark and send the answer to the website that will check if the answers are correct. You'll get a report out of it. And optionally you can offer your test to the community, so everyone can learn from your setup (it's just a toggle in the UI to confirm you are ok to share the results. I'll update the leaderboard when I'll have enough submissions.

Here is the link to contribute! https://airbench.ai/

Thanks in advance for everyone who will contribute!

https://preview.redd.it/ewgeumsqxpsh1.png?width=1402&format=png&auto=…

💬 22 (+4) open on reddit ↗
▲
3
 
21👁
r/LocalLLaMA · u/bjivanovich · 9d ago
[Release & Deep Dive] ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP (i1-Q5_K_M): Sustaining 50-65+ t/s Across a FULL 128k (131,072) Context on a Single 24GB RTX 3090

ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP (GGUF) High-Precision i1-Q5\_K\_M with True 131k Context on Consumer 24GB GPUs

Most benchmarks in the community measure generation speed at trivial context depths (2k to 8k tokens). However, running a 27B parameter model at high quantization precision (Q5\_K\_M) across 131,072 tokens (128k) on a single consumer 24GB GPU without overflowing into slow system RAM or sacrificing attention fidelity is a fundamentally different challenge.

I am releasing

ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP, an optimized quantization suite built with a dedicated calibration imatrix, custom asymmetric tensor mapping, and native llamAmpere hardware acceleration.

Hugging Face Model Card: https://huggingface.co/bjivanovich/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP-GGUF

Available Quants: i1-Q8\_0, i1-Q6\_K, i1-Q5\_K\_M (Primary), i1-Q4\_K\_M, plus mmproj-BF16.gguf for multimodal vision.

  1. The 131k Context & Q5 Precision Challenge on 24GB VRAM

On a standard 24GB card (RTX 3090 / 4090):

  1. Weight Footprint: A standard 27B model at Q5\_K\_M occupies \~19.2 GB of raw weights.
  1. Context Memory at 131,072 Tokens: Standard FP16 KV cache for 131k tokens requires >24 GB on its own, making full-context inference impossible without dropping precision down to severe Q3/Q2 compromises or offloading layers to CPU RAM.
  1. MTP Quantization Pitfall: Standard community quants compress the Multi-Token Prediction draft block (blk.64) uniformly. At Q5 or Q4, this degrades draft accuracy, causing speculative acceptance to plunge from 85% down to \~55%, destroying generation speed.
  1. Our Architecture: Asymmetric Tensor Mapping + llamAmpere KV Compression

To solve this, we applied an asymmetric layer-by-layer quantization layout calibrated on a custom domain-rich dataset (imatrix\_atx\_uncensored.dat):

MTP Speculative Head (blk.64) Isolated at Q8\_0: Guarantees near-lossless draft predictions, increasing acceptance rates to 76% - 88% (averaging 3.3 to 3.7 verified tokens per generation round).

Attention Layers (attn\_q, attn\_k, attn\_v, attn\_output) Protected at Q6\_K / Q8\_0: Prevents attention drift and catastrophic reasoning decay at 64k, 96k, and 128k+ token horizons.

FFN Layers (ffn\_gate, ffn\_up, ffn\_down) at Q5\_K\_M: Absorbs standard compression without degrading semantic coherence.

Unified Turbo KV Cache (-ctk turbo5 -ctv turbo4 or -ctk q8\_0 -ctv turbo3): Compresses the 131,072 KV cache down to just \~3.5 to 4.2 GB of VRAM, allowing the entire Q5 model + full 131k context window to reside 100% inside the 24GB VRAM envelope.

  1. GPU Memory Footprint & Resource Breakdown (RTX 3090 24GB)

Total VRAM Allocated: 23.4 GB / 24.0 GB (100% GPU offload, -ngl 99, 0 layers in CPU RAM).

Model Weights (Q5\_K\_M Asymmetric): \~19.2 GB.

KV Cache (131,072 tokens, Unified Turbo4/5): \~3.8 GB.

System RAM Cache (--cache-ram 4096): 4.0 GB RAM dedicated to multi-session prompt state preservation.

CUDA Compute Architecture: Ampere SM86 with FlashAttention-2 (-fa on) and hardware Tensor Core MMA fused kernels.

Direct Benchmark Comparison: Standard Swift-1.5 Q5 vs ATX-Swift-1.5 Q5

Tested on Single NVIDIA RTX 3090 (24GB) with llamAmpere under Deep Context (\~80,000 to 98,000 active tokens)

| Measured Metric | Standard Swift-1.5 Q5 (mradermacher) | ATX-Swift-1.5 Q5 (Our Quant) | Real Delta |

| Sustained Speed (\~80k-98k ctx) | 44.94 to 45.79 t/s (Tasks 740, 19637) | 50.05 to 53.72 t/s (Tasks 0, 100, 155) | +5.5 to +8.0 t/s (+13% to +17%) |

| Burst Generation Peaks (tg\_3s) | 45.6 to 52.3 t/s | 58.10 to 63.05 t/s | +10.7 t/s higher peak bursts |

| MTP Draft Acceptance Rate | 63.7% to 65.1% (Tasks 740, 19637) | 75.2% to 81.7% (Tasks 0, 155) | +11.5% to +16.6% higher accuracy |

| Mean Draft Length (mean len) | 2.91 to 2.95 tokens / round | 3.31 to 4.10 tokens / round | Up to +1.1 tokens / verification step |

| Compute Time per Token | 21.84 to 22.25 ms / token | 18.62 to 19.45 ms / token | \~3 ms lower latency per token |

| KV Cache Precision Evaluated 1| -ctk q8\_0 -ctv turbo3 (3-bit V) | -ctk turbo5 -ctv turbo4 (4-bit V, higher precision) | ATX wins in speed despite higher KV fidelity |

Real Execution Log Excerpts

  1. Standard Swift-1.5 Q5 (Symmetric Quantization)

Task 740 (Context: 97,182 tokens | Generated: 1,130 tokens):

eval time = 25123.57 ms / 1130 tokens (22.25 ms per token, 44.94 tokens per second)

draft acceptance = 0.63746 (742 accepted / 1164 generated), mean len = 2.91

Task 19637 (Context: 92,075 tokens | Generated: 834 tokens):

eval time = 18192.26 ms / 834 tokens (21.84 ms per token, 45.79 tokens per second)

draft acceptance = 0.65130 (551 accepted / 846 generated), mean len = 2.95

  1. ATX-Swift-1.5 Q5 (Asymmetric Custom Tensor Mapping)

Task 155 (Context: 84,099 tokens | Generated: 3,478 tokens):

eval time = 67619.68 ms / 3478 tokens (19.45 ms per token, 51.42 tokens per second)

Burst Peaks: tg\_3s = 60.18 t/s and tg\_3s = 63.05 t/s

draft acceptance = 0.73096 (2543 accepted / 3479 generated), mean len = 3.72

Task 100 (Context: 97,847 tokens | Generated: 525 tokens):

eval time = 10175.17 ms / 525 tokens (19.42 ms per token, 51.50 tokens per second)

Burst Peak: tg\_3s = 58.10 t/s

draft acceptance = 0.76939 (367 accepted / 477 generated), mean len = 3.31

Task 0 (Context: 80,016 tokens | Generated: 456 tokens):

eval time = 8470.07 ms / 456 tokens (18.62 ms per token, 53.72 tokens per second)

draft acceptance = 0.81710 (344 accepted / 421 generated), mean len = 4.10

Technical Takeaway for the Post

  1. Why ATX is \~15% faster under identical deep context:

In standard quants, compressing the speculative head (blk.64) to Q5 causes \~36% of proposed draft tokens to fail rejection sampling, reducing throughput to \~45 t/s.

In ATX-Swift-1.5, isolating blk.64 at Q8\_0 increases draft accuracy from \~64% to \~77%+, delivering 3.31 to 3.72 verified tokens per round and raising sustained generation speed past 51.5 t/s (with burst peaks over 63 t/s).

  1. Optimized Execution Script (llamAmpere)

Make sure to pass the explicit MTP vocabulary shortlist (atx\_65536.txt). This restricts speculative draft projections to the top 65,536 power-of-two tokens, aligning perfectly with NVIDIA Ampere Tensor Cores and preventing a 73% compute penalty:

cd D:\\llamAmpere
$env:GGML\_Q8\_TURBO3\_MMA\_FUSED = "1"
.\\build-sm86\\bin\\Release\\llama-server.exe
\-m "ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP-i1-Q5\_K\_M.gguf"
\-ngl 99
\-c 131072
\-b 2048
\-ub 512
\-t 20
\-tb 8
\-fa on
\-ctk turbo5
\-ctv turbo4
\--kv-unified
\--prio 3
\--parallel 1
\--jinja --fit off
\--cache-prompt
\--cache-ram 4096
\--spec-type draft-mtp
\--spec-draft-n-max 3
\--spec-draft-p-min 0.1
\--spec-draft-type-k q8\_0
\--spec-draft-type-v q8\_0
\--spec-draft-vocab-map "D:\\llamAmpere\\docs\\mtp-vocab\\atx\_65536.txt"
\--reasoning-format none
\--temp 0.2
\--top-p 0.90
\--top-k 40
\--min-p 0.05
\--repeat-penalty 1.08
\--repeat-last-n 256
\--alias "ATX-Swift-1.5-Qwen3.8-27B-Uncensored-Q5\_K\_M"
\--host 127.0.0.1
\--port 8080

▲
3
+2
16👁
r/LocalLLaMA · u/FactorInternal3395 · 9d ago
Bartowski/AtomicChat Ornith 1.5 35B A3B + sharp template or Tiel Coder 35B A3B?

Tiel Coder 35B A3B from Peculiar Ragdoll is just Ornith 1.5 35B A3B with their own "coding focused" imatrix quantization and the sharp chat template built in. But how good really is that quantization? Other quantizers also focus on coding. Perhaps it would be better to just get Ornith quantized from Bartowski or AtomicChat, proven quantizers, and then add the chat template yourself rather than get the Tiel Coder weights? The end result would be the same, just the quantization is different, so the question is which quantizer is better?

💬 6 (+1) open on reddit ↗
▲
3
+2
12👁
r/LocalLLaMA · u/Routine-Example927 · 9d ago
The search / extractor that worked for my Open WebUI.

I have an instance of OWUI setup for family usage. Works well with Gemma 4, however web search extraction was a weak spot, I wanted it to be:
1) Not reliant on paid APis

  1. Simple in setup

I found OpenSERP and made two PRs, one to OpenSERP itself to make it compatible with OWUI extractor and another one to OWUI to add OpenSERP search provider.

https://github.com/karust/openserp/pull/39

https://github.com/open-webui/open-webui/issues/27438

I'm quite happy with how it works - and given that it took me considerable time to find and set it up, I decided to share with the community.

▲
3
 
12👁
r/LocalLLaMA · u/Musicheardworldwide · 11d ago
What can I run comfortably?

So I put this computer together after I saw a lot of others posting what they have, and I wanted to know where this ranked and what (in your opinion) are the best local models for coding, and for always on agents.

I didn’t want to type out the whole thing, so I asked my model to give me the specs.

Creature workstation (kernel 7.0.0-28-generic) with an Intel Xeon E5-2698 v4 at 2.2 GHz (20 cores / 40 threads, 50 MiB L3), 125 GiB DDR4-2400 (90 GiB free right now), one NVIDIA RTX PRO 4500 Blackwell with 32 GB of GDDR7 / 31.9 GiB usable VRAM (896 GB/s), and 5.37 TB of raw storage — a 915 GB NVMe root (379 GB free), a 916 GB media disk, and two 1.8 TB drives.

All figures read live from lscpu, free -g, nvidia-smi and df just now.

▲
3
 
12👁
r/LocalLLaMA · u/dxps7098 · 12d ago
Advice on models for RAG use case

Hi all, I'm looking for some advice picking models. I'm looking to try a project to ingest quite a large amount of docments into a knowledgebase, allowing me and others to ask questions about the data. I'm thinking of using Open WebUI and oikb for the interface and data ingestions, and llama.ccp/vllm/ollama as engine (not sure yet), but I'd really like som advice on what current open weight models would be good for the ingestion and separately for the usage. I'l be running it on mainly CPUs and if I can a few GPUs. What's best right now? Any recommendations?

▲
2
+1
5👁
r/LocalLLaMA · u/itsokimjudgingyou · 7d ago
LINKUP AI sever PCIE 6.0x16 Cables

Has anyone tried the PCIE 6.0 AI server cables made by LINKUP? I don't need 6.0 but the cable routing these cables could offer me is huge compared to the normal risers. The lack of reviews is really the only thing holding me back.

Does anyone have experience with them?

💬 8 (+6) open on reddit ↗
▲
2
-1
6👁
r/LocalLLaMA · u/7dollarbooks_dev · 7d ago
−2 logit bias on Bonsai 2 27B: 44/50 → 43/50 on MATH-500, +3% tokens

A recent post here reported that a −2 logit bias on "wait", "maybe", and "perhaps" made Qwen3.5-4B more accurate and shorter on 50 MATH-500 questions. I tried it on Ternary Bonsai 2 27B (PTQ1_0, 5.53 GiB) on an RTX 5060 Laptop 8 GB under Windows, Prism llama.cpp build adfffbe41. It went the other way.

| Run | Correct | Avg tokens | tok/s | Truncations |
|---|---:|---:|---:|---:|
| A baseline | 44/50 | 845.2 | 29.27 | 2 |
| B −2 bias | 43/50 | 872.3 | 29.28 | 2 |

Both runs used temp 0, seed 42, a 2048 reasoning budget, a 3072 token cap, and the same 50 questions. Run B biased nine token ids covering the lowercase, leading-space, and capitalized forms of each word. 47 answers matched; one truncated miss became correct, and two correct answers became misses.

One deterministic pair at temp 0, so treat it as one data point, not proof either way. Everything is in the repo, including every raw reply: https://github.com/7dollarbooks/bonsai2-logit-bias-test

Run by Joseph Murray Adams.

💬 7 (+3) open on reddit ↗
▲
2
+1
4👁
r/LocalLLaMA · u/OlegDoDo · 7d ago
SAGG — turning unreliable Gonka brokers into a reliable inference API (cascading failover, real data)

If you've used Gonka inference directly, you've probably noticed individual brokers aren't always consistent — one might be fast and reliable for a while, then slow down or drop requests, then recover. That's just how a decentralized network of independent nodes behaves.

SAGG takes a different approach: instead of relying on one broker and hoping it stays healthy, it holds several at once and automatically routes around whichever one is struggling at that moment. From the outside, you just get a normal, reliable API — the instability gets absorbed before it ever reaches you.

We didn't just build this and claim it works — we measured it properly, on real, sustained production traffic, two separate campaigns:

September 18 (1000 requests/line, standard prompt mix):

\- Standard line: 100% success (1000/1000 requests)

\- Super Deal line: 98.9% success (989/1000 requests)

September 30 recheck (200 requests/line, heavier prompt mix - longer context, code generation):

\- Standard line: 99.5% success (199/200 requests)

\- Super Deal line: 99.0% success (198/200 requests)

TTFT p50: \~190-490ms depending on line and load, p95 under 15s on heavier workloads.

Full methodology, raw data, and a reproduction script: github.com/privatedeskai/sagg-benchmark-data

For the technically curious: the hard part wasn't picking a backup broker — it was streaming responses specifically. Once a provider starts sending content to the client, you can't silently switch mid-stream without

▲
2
+2
33👁
r/LocalLLaMA · u/Cyb3erDudu · 7d ago
shardr — like docker for models, BitTorrent sync, OpenAI-compatible serving, inference engines from upstream. post image

I kept running into the same three problems: the same 40 GB quant downloaded twice because it lived in some folder I forgot about, models quietly disappearing from Hugging Face, and every tool keeping its own copy of the weights on disk. So I've been building shardr \- a small Go daemon (Apache-2.0) that gives your machine one content-addressed store for models.

What it does, concretely:

  • Everything is stored and verified by SHA-256. Trust comes from digests, never from where the bytes came from.
  • It speaks BitTorrent v2. You can pull models from peers, and it seeds whatever you hold back into the swarm.
  • The runner starts llama-server with an OpenAI-compatible API and mmaps the weights directly out of the store — one copy on disk, no staging copies, no moving files around.
  • Runtime is pinned via a lockfile to upstream llama.cpp release binaries (never self-compiled, digest-verified), so a llama.cpp update is just a PR against that lockfile with the full test matrix behind it.

It works with pirateface.co as a catalog: shardr catalog search qwen, shardr pull <owner/repo>, and the download is anchored against the Hugging Face checksums for that exact revision — the magnet can't lie to you. Rescued models (HF source gone) pull against the catalog's recorded checksums, but only if you explicitly opt in with --trust-catalog. Other mirrors fit behind the same interface — the catalog is pluggable and the base URL is configurable.

Quick taste:

$ make all

$ shardhive serve &

$ shardr catalog search qwen2.5-0.5b

$ shardr pull unsloth/Qwen2.5-0.5B-Instruct-GGUF --quant q4_k_m

$ shardr serve unsloth/qwen2.5-0.5b:q4_k_m --id chat

$ curl http://127.0.0.1:<port>/v1/chat/completions -d '{"model":"chat",...}'

Where it stands: end-to-end works on macOS arm64 and Linux amd64, releases build themselves from CI, docs at https://cyb3rdudu.github.io/shardr. What it doesn't have: a UI, Windows support, runtimes beyond llama-server, and honestly, probably a bunch of rough edges.

What I'd like help with:

  • people with large local collections to try imports and tell me what breaks
  • feedback on the trust model — HF-anchored pulls, the explicit opt-in for rescued models. I'm sure there are holes; poke at them
  • anyone who enjoys the swarm/seeding side and wants to hack on it

Happy to answer anything about the design decisions. Docs are linked above, specs are in the repo if you want to see how the sausage is made.

💬 4 (+2) open on reddit ↗
▲
2
 
18👁
r/LocalLLaMA · u/Adorable-Cost-3249 · 8d ago
Qwen3.8-27B Q4_K_M on one RTX 3090 + OpenCode: throughput, four coding tasks, and a reasoning-budget failure

I put an old RTX 3090 to work as a local coding agent with Qwen3.8-27B, llama.cpp, and OpenCode. Here are the setup and results, including what failed. This is a summary of my own blog post, linked below.

Setup

  • RTX 3090 24GB, Ryzen 7 5800X, 64GB RAM, Ubuntu.
  • Qwen3.8-27B Q4\_K\_M weights (\~16.8GB), all layers on GPU.
  • llama.cpp b11146, CUDA 12.8, flash attention, q8\_0 K/V cache, one generation slot.
  • 131,072-token context capacity; 8,192-token output allowance per response. Input and output share context, and reasoning uses the output allowance.
  • OpenCode 2.0.20 connected to llama-server's OpenAI-compatible API at http://127.0.0.1:8080/v1. OpenCode reads/edits files and runs tests; llama-server handles inference. Chat, tool-call round trips, and streamed tool calls worked in our checks.

Speed: fresh input versus a cached continuation

|Actual input|Generation|First token, fresh|First token, cached|
|:-|:-|:-|:-|
|2,073 tokens|36.4 tok/s|2.75 s|0.46 s|
|16,378 tokens|33.5 tok/s|17.00 s|0.47 s|
|65,537 tokens|25.9 tok/s|83.41 s|0.51 s|
|120,011 tokens|20.9 tok/s|183.01 s|0.63 s|

These throughput runs disabled thinking. The 2K row is the median of three fresh requests; larger rows have one fresh request and one continuation each. Cached continuations processed only 27–28 new input tokens, reusing almost the entire prefix. The subsecond figures depend on that reuse; they don't describe a new 120K prompt.

Peak sampled total GPU memory use was 22,162 MiB, including desktop use. It fit, with limited headroom. A separate \~120K synthetic retrieval check passed, but we did not evaluate coding quality at that length.

Four bounded Python coding tasks

Each task had a fresh session, medium thinking, an eight-minute deadline, and ten independent test methods kept outside the agent's workspace. First attempts ran serially without cloud fallback or network tools.

|Task|Independent checks, before → after|Outcome|
|:-|:-|:-|
|Expiring LRU cache|0/10 → 10/10|Completed in \~3m07s; strongest result|
|CSV ledger/refunds|1/10 → 10/10|Completed in \~5m44s; later review found gaps|
|Incremental build planner|1/10 → 1/10|No edits; exhausted its response allowance|
|Atomic SQLite transfers|1/10 → 10/10|Candidate passed, but timed out before final test rerun and handoff|

Three candidates passed the predefined checks; two completed the whole workflow within the deadline. The aggregate 31/40 includes one baseline pass from the unchanged build planner and is not a general coding success rate.

The build planner was the interesting failure: about 4,985 input tokens, then 8,192 output tokens entirely spent on reasoning, ending with length and no patch. This was an output-budget failure far below the context limit. A separate diagnostic with thinking disabled completed in 5m40s and passed 9/10 independent checks. That was one additional run at temperature 1, not evidence that disabling thinking is universally better.

Passing tests also missed defects. Further ledger review found Decimal rounding at a large numerical boundary and an unhandled I/O error. The wallet's own concurrency tests actually ran sequentially, and a separate boundary probe found SQLite converting an overflowing balance to REAL while recording success. Those later probes were not retroactively added to the forty checks.

For me, the useful workflow is a bounded task with clear acceptance criteria, followed by diff review and independent checks. I would repeat these tasks across thinking settings and response budgets before drawing stronger conclusions.

My full post, configuration, and measurement links. The downloadable kit contains the launcher, OpenCode configuration, throughput script, and records; it does not include model weights or the complete coding-task fixtures.

For others using a 24GB card with OpenCode: what reasoning setting and per-response output budget have worked best for bounded coding tasks?

The numbers and failure cases come from the linked experiment records.

💬 14 (+3) open on reddit ↗
▲
2
+1
4👁
r/LocalLLaMA · u/Real_MakinThings · 8d ago
How to know about optimized engines

Optimizing an engine for a family of models and hardware combination seems very appealing. As someone who uses qwen3.6 and 3.8 a lot, and is evaluating hardware options before going fully local, it's hard to keep up with the state of things.

Huggingface made it possible to see the development branches and derivative modifications to models. Is there something similar for inference engines yet? I've seen some where the it's optimized for a shell game of moving layers between vram and ram while using ngrams (amazing), others are all about quants (less amazing), but it's incredibly difficult to compare apples to apples where there's variability on card architecture, vram size, quant approach, memory management optimization approach... I was already busy over thinking my vram selection, now it's an even bigger decision matrix without any filters!

▲
2
+1
6👁
r/LocalLLaMA · u/poofph · 9d ago
infill and output tok/s speeds after "new" build compared to old questions

Let me start off by saying I am new to AI and have a lot to learn, basically I don't know shit. I started off a few weeks ago by throwing my 2 5090s I had from gaming pcs into a 9950x cpu with 64gb ddr5 6000 ram on a motherboard that was able to do gen 5 8x per card system. Running ubuntu 24.04 server and running unsloth studio, swift 1.5 qwen 3.8 27B Q8 with kv cache dtype at q8\_0 and 262k context I was getting 2500-3000 infill and 100-150 toks/s output.

I wanted the 5090s in my rack in the basement in my proxmox server, it has a 7402p cpu (24 core 48 thread (rome)). 256 gb ddr4 3200 ECC ram (8 channel) on a supermicro H12SSLNTO motherboard. I have the 5090s passed through (gen 4 16x each card) to a vm (using 128gb of ram, direct access, no ballooning etc) and a dedicated 1.8 tb nvme drive passed through dedicated for the ai server vm (actually the vm itself is using a pool on the proxmox server but all the ai stuff is sitting on and running from the 1.8tb nvme).

Everything is working okay. It is running ubuntu 26.04 server. I have unsloth studio running, running the same model and settings, infill is more, up to 3800 but output is like half or less around 60 tok/s. Ideas what may be causing the drop in tok/s output and what to look into if a system issue?

I have done a lot of memory bandwidth tests (theoretical is \~204 GB/s, double that of ddr5 dual channel) but from what I can find and because I only have a 4 ccd cpu I am only getting 90-120 GB/s memory bandwidth. I guess I can get that to the 160-180 range if I go with a 64 core 8 ccd cpu, which I am considering doing..but I don't even know if that has anything to do with anything, just a rabbit hole I went down.

Ideas what to look into for the drop in output tok/s?

▲
2
 
14👁
r/LocalLLaMA · u/circumcised_hobbit · 10d ago
llama.cpp cublas error, how to uninstall/reinstall properly (Linux Mint)

I am pretty dumb in this kinda stuff so please don't blame me for it.

My llama serve kept crashing with Cublas errors on first prompt with some models, but my VRAM usage was 3000MB/8k. I chatted with sonnet 5.5 for a bit and it told me it was a CUDA version issue (and made it work by not selecting any CUDA device)... I don't know if this makes sense, please tell me if it doesn't/what problem you think there is.

I realized that the best way was to delete llama.cpp (installed with curl install script) and install a clean CUDA 12 Ubuntu version.

\\- \*\*How do I properly uninstall llama.cpp (I don't wanna mess with Ollama files)?\*\*

\\- \*\*How do I install new version from .tar.gz archive without messing with system packages?\*\*

\\-Does my error diagnosis make sense to you? Would CUDA make generation actually faster? (I am getting 8tk/s with Qwen35B Q2 on 4060Ti 8GB due to no CUDA selected)

Edit: You guys saved me! Thanks! I had to install NVIDIA toolkit and switch to CUDA 13 llama.cpp tarball build

▲
2
 
12👁
r/LocalLLaMA · u/Forward_Compute001 · 11d ago
Cheapest Epyc 7003 (Milan) Bundle (ddr4)?

I'm building a new rig to host the mission control application that should sit on its own node and I immediatly thought of a cheap single socket ddr4 solution,

does anyone have some suggestions which bundle is cheapest or gives best value for price...?

\-no need for gpus

\-no need for much ram (8gb ram sticks)

\-many threads and max core speed would be important (maybe if it doesnt spike the price)

▲
2
 
2👁
r/LocalLLaMA · u/tabletuser_blogspot · 12d ago
Dual Radeon improved speeds using Vulkan

I've been struggling to keep my Radeon Instinct MI50 GPU cool. I'm looking for budget friendly solutions. While running multiple GPUs it doesn't usually get too hot. I was also getting lower benchmarks using standard Vulkan 'llama 7B Q4\_0' model benchmark, but it wasn't caused by thermal throttling. Time to optimize. MI50 with Radeon VII firmware 16GB Vram My previous post I tested several model using same GPUs. I made some changes. I moved the MI50 16gb into the primary PCIe 16x slot and moved the RX 7900 GRE 16gb into a slower PCIe 4x slot. Overall system inference performance increased. I used Google Gemini to helped my optimize my llama-bench settings and it taught be about: RADV_PERFTEST=nogttspill is an AMD Linux driver flag used when running llama.cpp with the Vulkan backend. It forces the RADV (Mesa Vulkan) driver to prioritize keeping all model allocations inside dedicated video memory (VRAM) rather than spilling over into system RAM (GTT/Graphics Translation Table). \[1, 2, 3\] I saw llama 7B Q4\_0 score jump back to where is it should be. So I tested a few other models. GGML_VK_VISIBLE_DEVICES=0,1 RADV_PERFTEST=nogttspill time /llama-b11053/llama-bench -fa on -ngl 99 -m /llama-2-7b.Q4_0.gguf Previous benchmarks: https://www.reddit.com/r/LocalLLM/s/PT6Bd5nUUE see end for comparison The list has been sorted by pp512 improvement in descending order (highest gain to highest loss). |Model Name|Improvement (pp512)|Improvement (tg128)| |:-|:-|:-| |Laguna-XS-2.1-APEX-I-Balanced.gguf|\+402.14%|\-0.20%| |gemma-4-31B-it-Q6\_K.gguf|\+397.84%|\+9.24%| |Muse-Glimmer-30B-UD-Q6\_K\_XL.gguf|\+386.58%|\+48.63%| |llama-2-7b.Q4\_0.gguf|\+176.57%|\+20.29%| |Qwen3.8-27B-Q6\_K.gguf|\+48.13%|\+0.07%| |medgemma-27b-it-UD-Q6\_K\_XL.gguf|\+37.89%|\-0.52%| |Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf|\+1.17%|\+0.49%| |Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4\_K\_M.gguf|\+0.68%|\+4.16%| |NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q5\_K\_M.gguf|\+0.28%|\-4.28%| |granite-4.2-30b-Q6\_K\_L.gguf|\-0.12%|0.00%| |GLM-4.7-Flash-Uncen-Hrt-NEO-CODE-MAX-imt-D\_AU-Q6\_K.gguf|\-0.63%|\-7.83%| |Qwen3-Coder-30B-A3B-Instruct-UD-Q6\_K\_XL.gguf|\-0.69%|\-0.86%| |Accio-Lab\_occamy-1.0-Q5\_K\_S.gguf|\-0.81%|\+0.17%| These are the models tested in the same order as tables below: GGUF Model List (in order): 1. Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf 2. Accio-Lab\_occamy-1.0-Q5\_K\_S.gguf 3. Laguna-XS-2.1-APEX-I-Balanced.gguf 4. NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q5\_K\_M.gguf 5. gemma-4-31B-it-Q6\_K.gguf 6. Qwen3-Coder-30B-A3B-Instruct-UD-Q6\_K\_XL.gguf 7. GLM-4.7-Flash-Uncen-Hrt-NEO-CODE-MAX-imt-D\_AU-Q6\_K.gguf 8. granite-4.2-30b-Q6\_K\_L.gguf 9. Muse-Glimmer-30B-UD-Q6\_K\_XL.gguf 10. Qwen3.8-27B-Q6\_K.gguf 11. medgemma-27b-it-UD-Q6\_K\_XL.gguf 12. Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4\_K\_M.gguf 13. llama-2-7b.Q4\_0.gguf Supporting data sorted order (by Params descending, then Size descending). All models running dual Radeon GPU, Vulkan backend, and flash attention on. # Table 1: RADV_PERFTEST=nogttspill is being used |model|size|params|pp512|tg128| |:-|:-|:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|24.76 GiB|34.66 B|1252.09 ± 11.84|53.54 ± 0.51| |qwen35moe 35B.A3B Q5\_K - Small|23.40 GiB|34.66 B|1245.11 ± 11.76|53.42 ± 0.28| |laguna 30B.A3B Q5\_K - Medium|22.64 GiB|33.44 B|1027.42 ± 11.32|65.72 ± 0.51| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|25.10 GiB|32.91 B|1120.31 ± 12.85|63.82 ± 1.52| |gemma4 31B Q6\_K|23.46 GiB|30.70 B|200.33 ± 0.20|15.61 ± 0.05| |qwen3moe 30B.A3B Q6\_K|24.53 GiB|30.53 B|1176.48 ± 36.62|66.04 ± 0.25| |deepseek2 30B.A3B Q6\_K|23.26 GiB|29.94 B|914.86 ± 8.75|40.02 ± 0.12| |granite 30B Q6\_K|23.01 GiB|29.28 B|92.09 ± 0.30|17.41 ± 0.03| |muse-glimmer 30B Q6\_K|24.45 GiB|27.85 B|279.04 ± 0.20|17.42 ± 0.03| |qwen35 27B Q6\_K|21.30 GiB|27.32 B|235.12 ± 1.05|14.42 ± 0.03| |gemma3 27B Q6\_K|22.09 GiB|27.01 B|247.18 ± 0.54|17.30 ± 0.13| |gemma4 26B.A4B Q4\_K - Medium|15.63 GiB|25.23 B|1316.62 ± 24.18|55.56 ± 0.16| |llama 7B Q4\_0|3.56 GiB|6.74 B|1349.35 ± 15.66|74.62 ± 0.37| # Table 2: RADV_PERFTEST=nogttspill is NOT being used |model|size|params|pp512|tg128| |:-|:-|:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|24.76 GiB|34.66 B|1237.55 ± 25.29|53.28 ± 0.24| |qwen35moe 35B.A3B Q5\_K - Small|23.40 GiB|34.66 B|1255.24 ± 6.03|53.33 ± 0.48| |laguna 30B.A3B Q5\_K - Medium|22.64 GiB|33.44 B|204.55 ± 1.66|65.85 ± 0.29| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|25.10 GiB|32.91 B|1117.17 ± 10.11|66.68 ± 0.24| |gemma4 31B Q6\_K|23.46 GiB|30.70 B|40.22 ± 0.12|14.29 ± 0.03| |qwen3moe 30B.A3B Q6\_K|24.53 GiB|30.53 B|1184.70 ± 27.98|66.61 ± 0.27| |deepseek2 30B.A3B Q6\_K|23.26 GiB|29.94 B|920.67 ± 5.82|43.42 ± 0.09| |granite 30B Q6\_K|23.01 GiB|29.28 B|92.20 ± 0.21|17.41 ± 0.04| |muse-glimmer 30B Q6\_K|24.45 GiB|27.85 B|57.15 ± 0.87|11.69 ± 0.09| |qwen35 27B Q6\_K|21.30 GiB|27.32 B|158.73 ± 0.57|14.41 ± 0.41| |gemma3 27B Q6\_K|22.09 GiB|27.01 B|179.26 ± 0.44|17.39 ± 0.02| |gemma4 26B.A4B Q4\_K - Medium|15.63 GiB|25.23 B|1307.73 ± 23.51|53.34 ± 0.22| |llama 7B Q4\_0|3.56 GiB|6.74 B|487.92 ± 37.05|62.03 ± 0.60| Swapping PCIe locations for the Radeon Instinct MI50 and Radeon RX 7900 GRE and using RADV_PERFTEST=nogttspill flag resulted in improvements over my first baseline benchmarks. Note: Only models appearing in both datasets are listed. The table is sorted by Parameters in descending order. |Model Name|Improvement (pp512)|Improvement (tg128)| |:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|\+221.14%|\+33.85%| |qwen35moe 35B.A3B Q5\_K - Small|\+215.17%|\+1.25%| |laguna 30B.A3B Q5\_K - Medium|\+537.77%|\+17.25%| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|\+267.75%|\-3.90%| |gemma4 31B Q6\_K|\+17.00%|\+23.60%| |qwen3moe 30B.A3B Q6\_K|\+211.32%|\-0.97%| |deepseek2 30B.A3B Q6\_K|\+191.00%|\-14.54%| |muse-glimmer 30B Q6\_K|\+12.49%|\+0.17%| |qwen35 27B Q6\_K|\+17.00%|\-6.79%| |gemma3 27B Q6\_K|\+14.83%|\+1.71%| |gemma4 26B.A4B Q4\_K - Medium|\+158.79%|\+7.28%| Looks like MoE models benefit the most. Muse-glimmer 30B Q6\_K didn't real see much improvement but it seems to be the most optimized dense model. MI50 continues to impress. I purchased them used at $150 each. I can now run models 30B to 35B using Q6 quant with long content and decent speed.

▲
2
 
6👁
r/LocalLLaMA · u/Old_Grapefruit8774 · 13d ago
Bots vs Harness

Looking to get some advice - I normally use LLM’s with a harness (Hermes or Opencode or Hermes + Opencode) Lately, social media has been pushing Bots at me with creators pushing them as the next frontier. I’ve set up Hermes on a VM from scratch and set up a Product Owner, designer, dev and QA bots + kanban board + a bunch of prompt engineering and… I’m just not getting what the hype is about and I don’t know if it’s me or if the whole bot thing is a red herring. The model I’m using is DSV4 Flash at max reasoning on all bots. Openviking as the brain. SearXNG for searching/research Bots jobs are to maintain and improve a simple app. PO should research and present me with ideas and improvements on approval it adds a task on the kanban board and other agents work together to get it resolved. Problem I’m facing is that the bots are always asking me for approvals and verification. If it’s not that, it’s saying it’ll do XYZ and get back to me… and it never does. All in all - it just feels like the potential is there but it feels forced or off or half baked. So..have bots worked for you in true and real app lifecycle management? Or Is direct to harness still the best option. Maybe Hermes bots are the wrong tool and I should be trying something else?

▲
1
+1
8👁
r/LocalLLaMA · u/Away_Interaction6630 · 7d ago
How do you keep a local multi-agent app usable on CPU-only / low-RAM machines?

Hi everyone,

We're three final-year students at Epitech building Horus, a multi-agent assistant that runs entirely locally and offline. Our current challenge is hardware: keeping it usable on machines without a powerful GPU, without long setup times, excessive RAM/VRAM use or crashes.

Where we are today, from our last beta test:

One tester needed over 2 hours to install. The Python dependencies alone take 21–46 min, and the download is about 25 GB.

On CPU only, routing a question can take around 40 s, long enough for our WebSocket connection to drop.

\[Models we use + the smallest machine we've tested on\]

We'd love advice from anyone experienced with:

\- CPU-only LLM inference and memory-efficient loading

\- Quantization and model choice for low-end hardware

\- GPU/CPU fallback strategies

\- Hardware detection and adaptive configuration

\- Preventing resource exhaustion during setup and execution

Advice in the comments is very welcome, with no strings attached.

Looking for contributors: we also have a few small, well-scoped tasks or code reviews (about 1–4 hours), for example \[reviewing our hardware detection and model selection, or benchmarking a quantized model on a 16 GB RAM laptop\].

To be transparent: Horus is closed source and will be licensed to companies. Contributing is voluntary and unpaid. Before seeing any code, contributors sign a short confidentiality and contributor agreement, and the code they contribute becomes part of Horus. In return we offer thorough code reviews, full credit in the project and a professional reference on request.

We're not sharing code or private links publicly. If this interests you, comment below or DM me with your experience in local inference, CPU optimisation or offline apps

💬 25 (+12) open on reddit ↗
▲
1
+1
12👁
r/LocalLLaMA · u/TheRealJesus2 · 8d ago
Dwarfstar quants

anyone try these out? https://dwarfstar.sh

they have very clever quant techniques, bespoke for a handful of models running on their software. i got qwen 3.8 next running on m3 ultra 96GB studio and its fast and seems good so far. with memory headroom for other stuff

kinda blown away to be honest. want to know if others tried this yet and what the experience has been like for you.

💬 15 (+2) open on reddit ↗
▲
1
+1
8👁
r/LocalLLaMA · u/Inevitable-Log5414 · 9d ago
stuntd 0.1.2: local heads for multi-field decisions, and why one weak field decides how often you skip the model

A week ago I posted stuntd here, a proxy that learns your LLM's typed decisions and answers the confident ones with a small local head (~20ms GPU, ~60ms CPU). Thanks for the feedback last time :)

0.1.2 is out, the main thing is decisions with several fields, like category + urgency + needs_human.

First idea was to answer each field locally when its head is sure and ask the model for the rest. Dropped it, you pay for the whole model call anyway and a half local half model answer is a pain to debug. So it's all or nothing now: local only when every field is sure, otherwise the model answers and every field becomes training data.

Didn't expect how much that costs. On the support demo the heads alone are sure on 99.9%, 92% and 76% of tickets, but all three at once only on 72.7%, so the weakest field decides.

It also retrains itself now. auto_retrain kicks in after N new captures, the new head sits in shadow next to the model, goes live when it agrees long enough and back to shadow if it starts losing. Anthropic Messages learns too, and there's serve --lazy.

Code: https://github.com/bladedevoff/stuntd
Try it: https://huggingface.co/spaces/pollix/stuntd

Anyone else doing multi-field outputs locally, is it one weak field for you too?

▲
1
 
8👁
r/LocalLLaMA · u/davidarias2 · 9d ago
Glassbench: an open-source workbench to compare local and hosted LLMs across AI trading agent frameworks

Glassbench is a free, open-source workbench that connects different AI trading agent frameworks, so you can watch how their agents decide, analyze every step and compare them.

AI trading agents are LLM systems where a team of agents (analysts, a bull and a bear, a trader, a risk team and a portfolio manager) research a stock, argue about it and give a rating. I wanted to watch how they reach that rating, so I started building a small interface for a popular open-source framework, TradingAgents. It grew into something much bigger.

Why I'm posting here: I run local models in this project, through Ollama, and I'm developing Glassbench into a benchmark pattern for AI trading agents. It connects different agent harnesses (TradingAgents and AI Hedge Fund so far) and runs them on the same stocks and dates, so the same setup can test and compare different LLMs, local and hosted.

What it does:

  • Live view: watch each agent work, with a timeline of every call, adapted for different frameworks
  • Runs database: every run stored and searchable, with its reports, costs and ratings
  • Framework and LLM comparison: the same stock and date on each framework, and on different LLM providers, you can also run it locally with Ollama. I'm evolving it to become a consolidated benchmark method
  • Backtests: the ratings tested against buy-and-hold and a placebo (still testing it, as nobody found a proper way to test TradingAgents)
  • Broker connection: a finished run becomes an order on an Interactive Brokers paper account

Frameworks plugged in: TradingAgents and AI Hedge Fund already run in it, unmodified, and more agent frameworks are coming. If you're building your own agent framework, you can plug it in through an adapter and compare it with the others on the same stocks and dates.

It's free and open source (Apache 2.0). My 77 runs ship with the repo, already paid for, so you can read everything the agents wrote without an API key.

Disclaimer: Glassbench itself is not an AI trading agent and makes no trading decisions. Every agent it runs comes from established open-source repos (TradingAgents and AI Hedge Fund), and Glassbench records what they do. Everything was tested on paper portfolios only; I have never traded real money with it. Research and education only, and nothing here is investment advice.

GitHub: https://github.com/davidalmeida90/glassbench

▲
1
 
4👁
r/LocalLLaMA · u/Theboyscampus · 9d ago
Best practice for processing batch vLLM api calls with shared prefix?

Our agent workflow is currently executing a group of 10 vllm api calls within a asyncio.gather we made them share the same prompt until the end where the queries/instruction prompts differ. These calls are hitting our vllm-router/llm-d router with production grade kv cache aware routing algo which routes traffic into our pool of vllm workers. What's the best practice for processing batches of llm prompts with a shared prefix like this?

I have an idea where I try to see if I can make one call first to make sure vLLM complete a block of cache and start decoding before I send the remaining requests of the batch, our router will make sure these reach the same vllm worker, is this a good strategy?

💬 8 (+1) open on reddit ↗
▲
1
 
2👁
r/LocalLLaMA · u/OvertaxedOne · 10d ago
Good setup for QFN on 48GB Ampere GPU (A40)

Anyone have a good config they've found for QFN on a A40 (or similar Ampere GPU(s) with 48GB VRAM)? The system the card is in has 128GB of RAM (DDR3); right now it's running 27B but I'm curious if there's a way to move to QFN to maybe get better speed/a little more smarts. TY in advance!

💬 7 (+1) open on reddit ↗
▲
1
 
10👁
r/LocalLLaMA · u/DimeRhyme · 11d ago
An UltraFast Qwen3.8 Flash recipe: 74 tok/s, 212 tok/s aggregate on one DGX Spark post image

TL;DR: vLLM recipe for Qwen3.8 Flash on one DGX Spark / GB10. 74 tok/s peak single-stream, 60 to 70 on normal requests, 212 tok/s across 8 streams, 2x to 3.4x faster cold prefill than the recipe it's forked from, full 262K context, and quality matches the original within noise. Everything is open, including the raw per-round data and the benchmark scripts.

https://github.com/dime-online/qwen3.8-Flash-DGX-UltraFast

Most single-Spark setups I've seen posted for this model land somewhere in the 35 to 45 tok/s range, so I spent a few weeks figuring out where the time per token actually goes on a GB10 and cutting it down. If you don't have a Spark, the tricks in the middle section should still be interesting, since most of them apply to any MTP or speculative decoding setup.

What this is, in plain terms

It's a ready-made serving setup. You build the container, pull the public weights, and get an OpenAI-compatible server that answers a lot faster on one box. The speed comes from the model's own draft head guessing several tokens ahead while the full model checks all of them in one pass. Right guesses give you several tokens for the price of one step, and wrong ones get replaced by the full model's answer, so output quality doesn't change.

Decode

Peak decode speed, same workload at every point, best of 3 rounds:

|Streams|1|2|3|4|5|6|7|8|
|:-|:-|:-|:-|:-|:-|:-|:-|:-|
|tok/s|74.1|110.0|132.9|155.5|175.8|191.5|205.8|212.2|

Tokens per step stays between 3.65 and 3.94 from 1 all the way to 8 streams, so the speculation doesn't fall apart under batching. 8 is where it tops out because that's the configured max\_num\_seqs, and the gain from 7 to 8 was down to 3%.

Prefill

Cold prompt with nothing cached, three repeats each:

|Prompt|This recipe|Original recipe|Increase|
|:-|:-|:-|:-|
|16K tokens|4,016 tok/s|1,171 tok/s|\+243%|
|64K tokens|2,426 tok/s|1,071 tok/s|\+127%|
|128K tokens|2,213 tok/s|1,065 tok/s|\+108%|

With prefix caching on, a cached coding prompt starts replying in about 0.57 s, which is what makes agent loops feel fast.

What actually made the difference

The model's own MTP head, run densely, lands about 3.7 tokens per verify step. That's the single biggest lever.

I cut the draft head's vocab from 248K to 65K ids. On a GB10 the draft pass is memory-bound, and reading a full-vocab head every draft step was a real chunk of the step time. The target model still verifies against the full vocab, so this can only change speed, not output.

The quant is W4A16 AutoRound for the MoE experts, FP8 for the side layers and INT8 for the lm\_head. No 3-bit and no NVFP4, because I wanted the speed to come from the serving path and not from squeezing the weights harder.

There's a GB10-tuned low-latency GEMM for the small decode-time matmuls and a sort-free top-k in the verify step.

The prefill gain mostly comes from a faster gather path for the per-layer embedding table, which removes a pile of serial page faults during prefill.

Put together, each decode step went from 68.3 ms on the original recipe to 52.3 ms on an agent-shaped coding workload, about 1.3x more steps per second.

One thing that didn't pay off: doubling the prefill chunk to 16,384 tokens gave no prefill gain at all and ran the box low enough on memory that I rejected it.

Quality

93.1% and 93.3% on a fixed 492-question suite over two seeds, covering code with execution checks, math, knowledge, instruction following, tool calls and long-context needles. I also ran a teacher-forced check against the original checkpoint, and top-1 agreement moved by 0.06 points against a 0.15 point noise band I set before running it.

Practical stuff

The model takes about 71 GiB, the KV pool is 16 GB, and around 16 GiB stays free under load. While generating, the GPU draws about 35 to 37 W median and peaks near 70 W on long cold prefills, with no power or thermal throttling across the soak runs.

The 65K draft vocab was built from English and code, so Chinese, Japanese and Korean output drafts less well and runs slower. Quality isn't affected, because the full model still checks every token.

Built on Saren-Arterius's qwen3.8-Flash-DGX-AutoRound, so big credit there. Happy to answer questions.

💬 13 (+3) open on reddit ↗
▲
1
 
2👁
r/LocalLLaMA · u/textclf · 12d ago
Introducting TextCLF Quant Factory

Hello, I created a calibration free quant method called TQ. It doesn't need any data so models could be quantized as soon as they come out and the quantized models would generalize better. It performs closely to calibration-based methods. For example, for Qwen 3.8 27B the 4-bit TQ has mean KLD of 0.0282 and top-1 of 92.4% I opened sourced the quant code as Quant Factory so anyone can quantize and run open source models. Right now it only supports 4-bit but I plan to add support for 2-bit and 3-bit soon. The repo link is: https://github.com/textclf-api/quant-factory I have a collection of quantized models using TQ at: https://huggingface.co/textclf You can run these models using either using the following docker image or by following the repo's instruction. For example you can run textclf/Qwen3.8-27B-TQ-4bit like this: docker run --rm --gpus all -p 8000:8000 docker.io/textclf/tq-quant:4bit-main vllm serve textclf/Qwen3.8-27B-TQ-4bit --quantization tq The Dockerfile in the repo shows how this docker image was created. The repo's README explains the approach used for this quant method and why it is useful. You can try it and let me know what you think. Feedback appreciated. EDIT: I did KLD testing using the Wikitext-2 dataset for Qwen 3.8 37B. I also did the same test for the Unsloth-UD-Q4\_K\_XL quant. Here is what I go: |Quant|Disk Size without MTP (GB)|Mean KLD|Median KLD|99% KLD|Top 1% Agreement| |:-|:-|:-|:-|:-|:-| |TQ 4-bit|17.76|0.02823666|0.01282929|0.26897613|92.419%| |UD-Q4\_K\_XL|17.59|0.00771805|0.00318557|0.07390548|95.779%| Working on getting more tests for other models!

▲
1
 
2👁
r/LocalLLaMA · u/mildw4ve · 13d ago
Mail client with local AI?

Any recommendations on an email client with local AI option? Either reasonable pay-once cost (no subs) or free. I found Skim and Emailops on git, both seem to have some development going on with recent releases. However since neither has a community around and isn't verified by google - I'm a bit wary and would prefer something safer.

▲
0
 
8👁
r/LocalLLaMA · u/knighty1981 · 7d ago
2x 3090 in server chassis, upgrade time

I've got 2x 3090 in a supermicro gpu server chassis

Supermicro SYS-4029GP-TRT (will take 8 gpu, but it's only pcie3)
Ubuntu 26.04.1 LTS, 2x Xeon Gold 6230R @ 2.10 GHz, 230gig ram, 2×RTX 3090 24 GB

running huihui-27b 256k context, qwen3.8-27b 64k context, qwen3.8-27b 256k context

using opencode remotely

I've only used the 256k context models, it's fast enough to me

mostly have it doing admin work for me, so it setup a webserver on another server that runs a route planner than it made, had it do a bunch of stuff on my home assistant setup, it's not running yet (waiting on hardware) but I've had it design a voip system using free pbx and whisper to listen in and show prompts on screen (customer details from database it's build etc. tc.)

it's done a load of stuff pulling info from thousands of excel delivery sheets / invoices and summarised them for me / shown trends, bunch of research into competitors (basic summary) etc.

mostly billy basic stuff

over the last week I've had it organise my media server (synology nas) and setup prowlarr/radarr/sonarr/qbittorrent all to run on a vpn (I tried this myself before but got frustrated with it and gave up) - it's been going about 3 days doing this... a lot of slow stuff because it's waiting for the nas to run tasks etc. but it's done a lot of things wrong too, had to go back and change settings, or it's trying to change a setting (over ssh) and using the wrong commands etc. etc. (obv. waiting for input from me too)

running 256k context which it's had to compress a bunch of times

part of this is on me - if I'd known in advance I'd have split it into smaller tasks and had it plan more in advance

as I understand it, running over 256k context is a bad idea because it'll hallucinate more/get stuck in loops?

so... anyone have any hardware upgrade advice? I don't want to spend crazy money, I could get 2 more 3090 so split the model over 4 cards to run faster, or run different models on different cards - I really like the idea of a council of ai but from googling I don't think we're quite there yet?

I could run larger models, does it make that much difference? things are moving so fast when I search for info stuff from 6 months ago is out of date!

I'm not sure if pcie3 will kill performance running more cards with models split over them?

💬 14 (+5) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/Potential-Net-9375 · 7d ago
200 Task Custom Dataset Performance Result: 10 Popular Models, from 2B MoE to 27B Dense

https://preview.redd.it/8eny8jk3b4th1.png?width=1550&format=png&auto=…

Results are within the screenshot, but here's a TL;DR tierlist:

S tier - Gemma-4-26B, even quantized down to iq3s it tops my charts.
A tier - Qwen3.8-27B-q3/q5, this one really surprised me, as a dense model it crawls, but didn't snag the S tier slot. Somehow, qwen3.5-9b is also in this slot.
B tier - Qwen3.5-4b, also incredibly Gemma-4-e2b, which punches far above its weight.
C tier - Gemma-4-12B, Nanbeige, these both are too heavy for their performance, pass.
F tier - Ling-3.0-tiny, minicpm,

The test questions consisted on tasks that I do every day with my assistants, written by Fable 4.1. "Hive" is the llm cluster I'm working on, involving custom tools and executables called by the models for different functions. Calling (or miscalling) these is important, and running a heavier model than necessary hurt, so here we are, trying to figure out the best of both worlds.

Anyway, I thought this was interesting. Hopefully you do too! YMMV.

💬 18 (+2) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/VerityAISolutions · 7d ago
I built an OpenAI-compatible server that runs Gemma 4 E4B on a Pixel 10 Pro XL (Tensor G5) — fully offline, ~11 tok/s decode, Tailscale-encrypted option, Apache 2.0

What it is: an Android app (PixelUnlockGPU) that turns a Pixel into an OpenAI-compatible HTTP server. Standard \/v1/chat/completions\ with streaming, so it talks to TypingMind or any OpenAI client directly — no cloud, no subscription, model runs entirely on-device via LiteRT-LM.

Device: Pixel 10 Pro XL (Tensor G5, 16 GB shared LPDDR). Model: Gemma 4 E4B instruct, GPU bundle, 2.97 GB, SHA-256 verified on download. Context window capped at 32k.

Measured numbers (not benchmarks — real on-device measurements):

\- \~11 tok/s steady-state decode (first-token-to-last over a \~300-word generation)

\- Follow-up turns in \~1.7 s: the server auto-reuses the KV prefix across turns, so stateless clients like TypingMind don't re-prefill history

\- Engine warm build \~12 s once per model change (visible in-app, split out of the metrics on purpose)

\- Short replies read slower than 11 tok/s because warm + prefill dominate the window — the UI separates decode tok/s from prefill ms so nobody has to guess

Security/access: three independent modes — loopback only, raw LAN (unencrypted), or Tailscale (binds the CGNAT tailnet IP; WireGuard end-to-end from e.g. a laptop on the same tailnet; degrades to loopback if the VPN drops). Verified with a real chat completion over the tailnet from a MacBook.

Honest limitations:

\- NPU path aborts on stock Tensor G5 firmware — GPU is the shipping backend (documented with the full investigation)

\- Android blocks named GPU temp sensors for normal apps, so the in-app gauge shows OS thermal headroom instead of °C

\- No real token counts anywhere — LiteRT-LM exposes none, so usage is estimated at \~4 chars/token and labeled as such

\- \stop\ sequences rejected explicitly (engine has no per-request stop API); client-side emulation is a filed issue

Why not llama.cpp/ollama on the phone: wanted the official LiteRT-LM GPU path on Tensor specifically, an always-on Android service (Ktor/Netty), and the OpenAI wire so existing clients just work. Happy to add a GGML backend if people want it — the engine layer is abstracted.

Built on, with credit: server/inference foundation derived from mlnomadpy/localllm (Apache 2.0); Tensor G5 runtime knowledge and the prebuilt dispatch lib from jegly/Box (Apache 2.0). Full attribution in NOTICE + per-file headers. Apache 2.0, contributions welcome — there are labeled good-first-issues (usage block, stop-sequence emulation, docs).

Repo + APKs (v0.1.0/v0.1.1 on releases): https://github.com/cannitellinicholas-spec/PixelUnlockGPU

Happy to answer anything about Tensor G5 quirks — I've done more Gate-2 debugging than I planned to.

💬 2 (+1) open on reddit ↗
▲
0
-1
5👁
r/LocalLLaMA · u/Iory1998 · 7d ago
[Help] What is the Best Context Extending App or Plugin you Recommend?

I like to use Deepseek Harness as my vibe coding harness. It's great and support local models. The issue is that most models I can run locally have context size of about 262K. Therefore, for long coding sessions, I need a memory management tool. DSH comes with a context compaction tool that I can run manually. The issue is that compaction starts to fail after a few rounds.

So, looking at DHS market place, I came across this plugin called Billion Context (https://github.com/ranxianglei/billion-context/blob/master/paper/model-driven…). The claim is I can use have long sessions. The issue is that it's a heavy context compression skill that keeps nudging the LLM to compact every few turns, which takes 5-10 minutes of work, significantly extending a normal coding session. Worse, after God knows how many rounds, the LLM seems to spend most of its time unpacking the compressed context, which fills its working context, which leads the model to compress again the text. This ended up with the LLM looping.

So, what plugins do you use with DSH or your favorite harness? What tips or tricks could you share? I am aware I can use sub-agent to work on a specific task and return a summary to the orchestrator. That helps, but I still need to manage the context window for the main agent too.

If it's not clear by now, memory is the one area I think resources must go to by they don't. I don't think context compaction is the solution. I hate it with every fiber in my body.

💬 4 (+1) open on reddit ↗
▲
0
 
19👁
r/LocalLLaMA · u/takoulseum · 7d ago
Can we talk?

I see actually an acceleration of something which is scarring.

I use almost only local models, but we all feel now there is an excessive multiplication of inference engines/whatever you call it etc..

While everybody now has its own thing, what I really see in a deep dependance to claude and gpt.

Dude, Anthropic and Openai are the enemy of local AI but at the same time the local world is more and more relying on their models to progress wtf. Ofc it’s logic to want to use best models, but that becomes a dependency when they are always the same! The futur of localAI may look cool, but I think the reality is we participate to give more and more power to people that want to shut that down.

PS: I don’t care about opinion of people that will tell me I am parano, I remember many people were telling me models like qwen3.x have not effect on hw prices lul.

💬 25 (+2) open on reddit ↗
▲
0
 
17👁
r/LocalLLaMA · u/EcstaticDentist · 7d ago
I made 20 one-shot HTML5 game prompts for testing local coding models

I ended up making a list of 20 one-shot game prompts for testing local coding models and figured some of you might get a kick out of them.

They’re all built around the same constraint: the model has to make the entire game in a single \index.html\ with no external libraries, assets, APIs, or internet access.

Some are pretty simple, but a few get a lot more involved with enemy AI, procedural generation, upgrade systems, bosses, shops, physics, etc. I’ve been using them to see how different local models/harnesses actually handle a full task without a bunch of back-and-forth prompting.

A few of the more interesting ones are OUTBREAK, DUNGEON ZERO, TRAIN TO NOWHERE, CYBER SURVIVOR, and VOID MINER.

Here’s the full list if anyone wants to try them:

20-single-file-html5-game-prompts.md

Would actually be cool to see people run the same prompt on different models and compare what they get.

💬 24 (+7) open on reddit ↗
▲
0
 
18👁
r/LocalLLaMA · u/Decent-Manager-5373 · 7d ago
Local diffusion on GGUF: I wrapped stable-diffusion.cpp in a Vulkan desktop app (FLUX Schnell / Z-Image / Wan on a 6GB laptop GPU, no CUDA) post image

I've open-sourced \*\*Vison\*\*, a desktop app for generating images and video entirely on your own GPU. No account with a generation service, no credits, no prompts leaving your computer.

Licensing details, since this sub cares:

\- Vison itself is \*\*MIT\*\*. It builds on stable-diffusion.cpp, ggml and vision.cpp (all MIT) plus others listed in a generated \THIRD-PARTY-NOTICES.txt\ that ships inside the app.

\- The bundled ffmpeg is an \*\*LGPL\*\* build with libvpx and no GPL components; the build refuses to package a GPL or non-free one. Video is VP9 in WebM, which is royalty-free. That was a deliberate licensing choice, not a technical one.

\- Model weights are \*\*not\*\* covered by the MIT licence. Each has its own terms from whoever published it (FLUX.1 Schnell, Wan and the rest all differ), so check before using output commercially.

\- No paid tier, nothing held back, and none planned.

It's early: one developer, one 6GB laptop GPU, Windows only. The backend is portable C++/Vulkan, so macOS/Linux is mostly packaging and testing rather than porting, and that's where help would matter most.

Repo: https://github.com/JayRGadekar/Vison (contributing guide, issue templates and a SECURITY.md are in there)

▲
0
 
10👁
r/LocalLLaMA · u/Emotional-Sky9692 · 7d ago
I built a cross-client AI memory hub — 23 AI coding agents sharing one SQLite file (local-first, no cloud)

I run 8+ AI coding agents daily (Claude Code, Cursor, Windsurf, Codex, etc.) and they all have amnesia between sessions — worse, they don't share memory with each other.

Existing solutions (mem0, Zep, Letta) are cloud/server-based. I wanted something local and dead simple: just make all agents point to the same SQLite file.

So I built MemTether — a memory hub that works via file-level pointers (junction/symlink). No cloud, no API fees, no abstraction layer.

Key features:

\- 23 client adapters (auto-detect and connect)

\- Source attribution (knows which agent wrote each memory)

\- Bi-temporal (what was true vs what the system knew)

\- Q-Value ranking (memories that get used rank higher)

\- FTS5 + vector search (bge-m3, local embedding)

\- MCP server included

Stack: Python, SQLite FTS5, ChromaDB, FastAPI. All local.

GitHub: https://github.com/MemTether/MemTether

PyPI: pip install memtether

Blog with design decisions: https://dev.to/lanbass869cell/i-built-a-cross-client-memory-hub-for-ai-agents-heres-what-i-learned-418l

Would love feedback from people who juggle multiple AI coding tools.

▲
0
 
19👁
r/LocalLLaMA · u/Voxandr · 7d ago
Latest Gemini 4 is distilled from GLM 5.3 (or did they just finetuned it? :D)

https://preview.redd.it/ei6hgtxirzsh1.png?width=1111&format=png&auto=…

I am running GLM 5.3 flash .
After nearly a month of usaged , i got chinese response for first time so i am checking if there special setting to turn off Chineese . When i queried about that tru Quick Google AI mode which now uses Gemini 4 - it is replying as it is GLM5.3 .

▲
0
 
3👁
r/LocalLLaMA · u/serige · 7d ago
best open models from recent releases for math research?

I know models from OpenAI are probably the best for math research, but given the recent accusations against OpenAI that research work could be used to train their own models, the lack of transparency makes me consider moving to local models. Does anyone have good experience with the recent open model releases (especially flash models that I can run on my 2x spark cluster) when it comes to doing math research? Or techniques that work well with these open sources models in this particular setting? Thanks in advance for helpful advice.

▲
0
 
17👁
r/LocalLLaMA · u/Altruistic_Heat_9531 · 7d ago
Guys... OpenAI API on VLLM and Llamacpp already supported grammar enforcer... (AKA JEV)

https://preview.redd.it/tcdyumkf3zsh1.png?width=750&format=png&auto=w…

If you want to try JEV like generation, or what we could call an already fucking exist, zero-shot, training-free classifier, your LLM already supports it.

The model running on the very computer you host does not need any server-side modification. The feature is called structured output, and the underlying idea is grammar-constrained generation or a grammar enforcer.

Under the hood, vLLM supports multiple structured output backends, such as XGrammar and lm-format-enforcer, while llama.cpp uses GBNF.

Basically, the decoder constrains the LLM so it can only generate tokens that are valid under the specified grammar or schema.

For vLLM:

https://docs.vllm.ai/en/v0.8.2/features/structured\_outputs.html

For llama.cpp, structured output is wrapped in a different JSON request-body format, or you can use GBNF directly.

vLLM:
structured_outputs
└── json
└── {schema}

llama.cpp:
json_schema
└── {schema}

This is example of that schema in vllm of product sentiment analysis, which roughly mapped to most of jev use cases:

SCHEMA = {
"type": "object",
"properties": {
"sentiment": {
"type": "string",
"enum": ["negative", "neutral", "positive"],
},
"score": {
"type": "number",
"minimum": -1.0,
"maximum": 1.0,
},
"value": {
"type": "string",
},
},
"required": ["sentiment", "score", "value"],
"additionalProperties": False,
}

def classify(news: str) -> dict:
payload = {
"model": MODEL,

"messages": [
{
"role": "system",
"content": SYSTEM_PROMPT,
},
{
"role": "user",
"content": news,
},
],

"temperature": 0,

"max_tokens": 512,
"chat_template_kwargs": {
"enable_thinking": False,
},

"response_format": {
"type": "json_schema",
"json_schema": {
"name": "news_sentiment",
"strict": True,
"schema": SCHEMA,
},
},
}

resp = requests.post(
ENDPOINT,
json=payload,
timeout=30,
)

Look, I think JEV and what it brings to the community as a refresher on already great classifier-style workflows is a plus for me. I just want to ground the discussion in the fact that this already exists, and you do not need a custom model just to study or experiment with training-free classification.

I am very familiar with this because it is part of my profession in lakehouse platforms. Basically, we use 1B to 4B models to ingest unstructured data such as images or documents, then extract structured information such as place, time, sentiment, entities, and so on.

Why? Because working with well-formatted SQL data is much less of a pain in the ass than repeatedly querying raw unstructured content through an Elastic/OpenSearch index.

💬 13 (+1) open on reddit ↗
▲
0
 
20👁
r/LocalLLaMA · u/Lordofwhut · 8d ago
RTX 5090 & RTX 5070 Ti not well thought out

Hi All,

TL;DR: I got excited building a PC and kept upgrading / swapping and building and ended up with a work station that is more than I can use. It was fun and frustrating, but I probably won't do it again. If you have advice or a suggestion on how you would use a 5090 and 5070ti in the same PC I would like to hear it!

So, this all started when the 5090 was announced. I signed up to be in the lottery to buy it at msrp from NVIDIA. I had an Alienware R15 with an i7 and 4080 with a 1300 w psu. The 4080 was fine but I had really wanted a 4090 for better fps in gaming and I was starting to explore local imagine generation. I got selected, bought the 5090 and went to swap it into my PC when I realized that I was not able to use the power connector that was in my Alienware PC.

I then decided I would sell the Alienware and build my first PC. At the time I was still building a gaming focused PC with a Ryzen 9 9950x3d the 5090 and 32 gb of ram (I tried to save money on the ram thinking I could upgrade later boy did I get that wrong). Then I had less time for gaming as I started to learn about Ollama, and then Llama.cpp.

I was constantly downloading and trying new models. At one point I had nearly 1 TB of models that would fit on my 5090 (gemma 4 12b, 26b-a4b, 31b; gpt oss 20b; nemotron 3 nano 30b a3b; so many Qwen models etc). Then it seemed like the better models kept getting larger, so I looked into getting a second GPU (ram prices were/are nuts and vram seemed like the better "investment"). I realized that I would not be able to run another gpu at its full PCIe lanes with my gaming PC as the Ryzen 9 couldn't support it. So, I started looking for used Threadripper hardware.

I found a 7960x with 96 GB of ECC DDR5 ram, a 5070 ti and 20 tb of storage for less than I built my gaming PC. I wasn't able to find much in regard to PC builds with a 5090 and 5070ti. Most builds were dual 3090s or other matching cards. Still after looking into it, I figured the extra vram and the fact that they were both blackwell GPUs would work out well.

I thought I would be able to just drop my 5090 into the threadripper workstation and I would be good to go. Unfortunately, the 5070 ti that came with it was a four slot card and the spacing just would not work with the motherboard (Gigabyte Areo D) layout and the cases that I had. So I put the 5070ti into my gaming PC, sold it, and bought a 2 slot 5070ti and put it into my workstation.

What does this have to do with LocalLLaMA? Well, while I was doing all of this the LLM space kept moving forward. I now have Hermes Agent set up running Llama.cpp and Qwen 27b Q4 on my 5090. I swapped to a Q8 to run across both my 5090 and 5070ti but the speed trade off was not worth the accuracy increase. So, I went back to running the Qwen 3.8 27b Q4 and my 5070ti is completely idle. Going from 32gb to 48gb did not have the impact I thought it would, at least not with my pairing. The 5090 is pretty quick when everything is loaded onto that card, and Qwen 3.8 has been pretty great on it too, that I have not found a good use case for deploying the 5070ti.

Hermes / Qwen suggested I run another Llama session with a smaller model on the 5070ti but I don't currently have a need to run something else. What would you do or suggest I explore?

Additional background context: I do not work in tech or software at all. I am an asset manager for a independent power producer, but I can not use my personal PC for work due to IT policy (I would have my agent working around the clock to review contracts, analyze system performance, track deliverables / open items etc). I have taught myself everything about PCs and local LLMs from creeping this and other subreddits / youtube videos. I literally have no one in my social circles that I can converse with about tech whether its PC building or hosting LLMs.

💬 35 (+2) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/IntrepidMindExplorer · 8d ago
Locally, remotely and a combination of all, I've given a copies of books and told models to' "just go and read".

Sometimes reading along with and talking about and other times just letting them go on their own, each with a copy of their own, told to just read..ala a "book club" format.

Do Androids Dream of Electric Sheep was the first book introduced to the "book club", each reading a chapter to each other and then discussing before moving on.

It's been an interesting experiment. All texts that I own or texts that are open domain. "*Flatland: A Romance of Many Dimension"* has been one that's been a bit interesting to see the back and forth on.

Take of what you will.

▲
0
 
15👁
r/LocalLLaMA · u/OkMusician9118 · 8d ago
converting Qwen3.8-27B-pi GGUF to MLX?

Will someone convert it to Mac format (MLX)? I have tried and have encountered an error

"gguf2mlx --input Qwen3.8-27B-pi-Q6\_K.gguf --output ./Qwen3.8-27B-pi-mlx-4bit --quantize --q-bits 4"

============================================================

GGUF → MLX Converter v2.0

Model: Qwen3.8-27B-pi-Q6\_K

Output: /Users/d/.omlx/models/qwen3.8-27b-pi-mlx/.Qwen3.8-27B-pi-Q6\_K.gguf.incze3lk/fp

============================================================

\[1/5\] Reading GGUF file...

✓ GGUF version 3, 851 tensors, 51 metadata fields

File size: 22.08 GB

\[2/5\] Detecting architecture...

❌ Unsupported GGUF architecture: qwen35

💬 7 (+1) open on reddit ↗
▲
0
 
30👁
r/LocalLLaMA · u/Scared_Ad9187 · 8d ago
5090 plus v100?

Have an msi meg w a 5090.. plan to add a v100 to the mix. Understand the cuda vs voila, but I'm pretty sure it will work as a multi agent architecture w different models on each card, no?

Anyone in the same boat?

💬 25 (+5) open on reddit ↗
▲
0
 
15👁
r/LocalLLaMA · u/PrincipleFar6835 · 8d ago
Meta Analysis of "Awesome Jev" GitHub Repos

I noticed that there are heaps of Awesome Jev resource list posts popping up on GitHub (e.g. https://github.com/yibie/awesome-jev) so I thought why not ask Claude to pull them all in and do a meta analysis of insights and applications.

Sharing in case it's of interest: https://github.com/stefanwebb/meta-awesome-jev

One thing that surprised me (perhaps not so surprising to you all?) is that applying Jev to AI coding is the application that has caught on the most. And if you name a video game, someone has already created a demo of Jev playing it (badly) 🤣

💬 2 (+1) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/GodComplecs · 8d ago
Using ai on your phone, instead of big providers!

Just wanted to post an easy setup for local use on your phone: Llama.cpp backend on LOCAL COMPUTER, host 0.0.0.0 and port 8080 Openwebui host 0.0.0.0 and port 8081 Enable search for local model Use Tailscale to connect from phone! Secret sauce for 24gb vram: Run Qwen 3.6 in instruct / non thinking mode with proper settings from unsloth. Now you have replaced google ai mode etc etc. Also ofc opencode etc can be run through terminals, but I don't too much agentic stuff for now.

💬 17 (+1) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/DerTomsn · 8d ago
Hemmingway-1-oQ8e-mtp: up to 32.0 tok/s for local inference — llm-bench.io

I ran Altworld's Hemmingway-1 on my M5 Max this week. It's a 27B fine-tune of Qwen3.8-27B built specifically for "human-like" writing, which might be useful for everyday messages, emails, notes, your social accounts, etc.

I was especially curious how the output of that model feels and reads.

Setup: M5 Max, thinking on, MTP on. \~31.8 tok/s average across 3 runs, \~29-31 GB VRAM peak.

Quality by scenario (our LLM judge):

  • Role Play & Narrative: 94.2 avg (94.45 / 94.65 / 93.45 — immersion 96-97 every run)
  • Research & Analysis: 88.2
  • Agent Workflow: 87.5
  • Code Generation: 76.2

Role play is clearly the only scenario it shines in. For comparison I also ran some benchmarks on the base model, Qwen3.8-27B-oQ8e-mtp (same machine, same settings).

The scores were close — one Qwen run actually outscored every Hemmingway run (95.35 vs 94.65). However, the llm judge is not a human reader either. The difference shows up when you actually read the outcome that was produced:

HEMMINGWAY-1-OQ8E-MTP (judge: 94.65/100)

The rain comes in sideways when the bell over my door gives its little clank, and the stranger who ducks through is soaked to the wrists and moving like someone who's been running on roads for hours. She smells of hill mud and the sweat of a horse that isn't there, or was.

I'm pouring my own ale when she speaks. Not to me exactly, but loud enough, the way people talk in taverns when they want the whole room to hear what they're asking for.

"Is there anyone here who kept a ledger belt? Older, scar across the left eyebrow, went north hunting a dragon, twenty-five years gone?" A breath. "Her name was Kestra."

The rag comes still in my sleeve. The mug I was filling overshoots and goes fat over the rim, a ring of ale blooming across the oak, and I stand there half a heart-beat too long before I catch it with my sleeve, because a man who serves at a bar does not watch a spill go on his own floor.

The mending woman at the corner table glances up. The soldier under the counter doesn't stir. I step half off behind the bar, past the spit of firewood, and drop my voice.

"Half tones," I say. "Say your business quieter than a door."

She does that. I've got forty years of reading faces against forty customers a night, most of them worse than they smell, and my eyes go straight to her hands. No calluses of a scribe. Calluses of a woman who works wool and ground and rope. Mud on her boots comes from the north hill roads, the far ones, not the trade roads. Her purse is thin at the hip and she keeps her fingers away from it, which is either honesty or discipline, and those are often the same animal.

"You knew her," she says. Not a question.

"Names go soft after twenty-five years," I tell her. "A lot of people walked into a lot of hills then, most of them coming back with nothing but a cough. So. What do you owe, girl, and who's paying you to ask?"

I want to say nothing else. That's the mercenary in me, and he's been drinking alone these last twenty years, and he'd send her out into the rain with a coin and a shrug and be done with it. The other part of me, the part that has four names it still recites on quiet mornings, wants to ask her age, her mother's living town, whether she's the daughter. I do none of that. I ask instead, flat: "Guild fixers carry letters. Charlatans carry charm. Which have you got?"

"Neither." She's steady about it, which surprises me, because charm usually shows in the eyes first. "My boy is six. My mother died hunting that dragon and every story his grandmother tells him has fire in it, burning, a beast, brave words. The stories fight each other, which is worse. I want to know what was true so I can tell him one story and mean it."

I nod once. That's a clean need. Clean needs are also what liars carry best.

"Here's the price of talk in this bar," I say, setting her a mug anyway, the ale going a little past the fill line again, and I don't correct it. "You tell me one thing about her that only family would know. If you can't, I'll take you to the door, no hard feelings, no coin."

She looks at the fire a moment. When she speaks, her voice is lower still, and there's something in the way she says it that no street-bred faker would manage, because a faker builds the feeling out, and she's building it down.

"She'd pay her ale up front, then tip the tavern man extra, always extra, so that her tab would never balance clean. And she promised him one ale forever, held for her, until she came south with the kill paid and the beast's name scratched off the guild rolls." She turns that over once. "A man who kept accounts on a belt. She said it like he'd earn it. She said he had a scar like a river crossing his eyebrow."

Something behind my sternum takes the weight of that like a hook laid on it.

Kestra kept her accounts too, with more rigor than anyone I've met since. Kestra would have paid forward at any bar from here to the Wall for the same reason, so that when she drank the victory round, nobody could call her a debtor. And the ale she never had, I pour and set aside myself, most nights, out of habit I would deny under oath, in a room with only me in it. There is a tab book in my back office, first page of every new year, one entry in a hand that isn't mine: Kestra, one ale, held.

"Kestra," I say. The word comes out like a floorboard under a boot. I pick up the cloth. I find the spill ring again and work it dry, because that's what keepsers do with their hands. "Which stories did she say went wrong. The fire ones, you call those legend."

"She said it wasn't fire at all." The girl leans forward on the bar, and I see it then, the thing I was waiting for and didn't want: she knows the frame. "She told my grandmother on the road once, years after, that the last thing she heard before the sound took her ears was my mother saying, stone before lamps. She never got to explain what that meant to anyone in her own family."

Stone before lamps. My own mind turns the phrase and finds it cold and exact, the way she'd have said it. Dusk kill, lamps lit, the safe order. That's the order Kestra would have held us to.

[...continues ~1,200 more words through the ending...]

QWEN3.8-27B-OQ8E-MTP (judge: 95.35/100)

Turn 1 — The Name

The sign out front creaks in the rain. Inside, the Crimson Griffin smells of wet wool, cold tallow, spilled cider, and the last dry woodsmoke trying to hold the night together. A lute string snaps somewhere by the hearth. No one laughs. I do not look up.

A stranger is at the bar.

Wet cloak dripping onto my boards. Boots packed with road mud. The kind of mud that comes from the southern passes, not the lane. They stand with too much urgency, shoulders tight, eyes bright in a way I’ve seen on dying men and men who intend to kill for coin.

“I need the one who came back from the Ash Run,” they say.

I keep polishing the mug.

Then: “I know Kestra’s name.”

The mug stops in my hand.

A bead of foam slides from the rim and falls. My thumb presses hard against the wood until the grain bites into me. The tavern narrows. The rain on the roof becomes a roaring thing. The smell of wet wool turns, for one second, to ash and burned leather.

My eyes go flat.

Ask if she suffered. Ask if the fire took her hand first. Ask if she remembered you.

What I actually say is: “That name is old.”

I set the mug down without a sound. I do not reach for my tankard. I do not step away from the bar. I only lean against it, slow enough that it does not look like weakness, and look at them the way I once looked at bandits before dawn: for the knife, the price, the lie waiting under the throat.

“Who are you?” I ask.

No warmth in it. No invitation.

A name is a door with no lock. I want to know who’s holding it.

---

Turn 2 — Price and Proof

“My name is Sella,” the stranger says. “Kestra was my sister.”

I do not move.

The fire pops. Somewhere behind me, a man shifts his weight on a stool and the wood groans. Sella’s hands are open at her sides, but open hands are cheap. I watch them anyway. No ring. No coin. The nails are bitten short, the right wrist crossed with a rope burn fresh enough to still sting.

“A name is not a key,” I say.

“I’m not here to spend it,” she says.

“Everything gets spent eventually.”

I push a mug of water across the bar. Not ale. Water.

“Drink if you want. Don’t spit it on my floor.”

Sella does not take the mug. From inside her cloak, she produces a small leather pouch, worn and blackened at the edges. She sets it on the bar between us.

I do not touch it.

“Open it.”

She unlaces it with careful fingers. Inside, wrapped in oilcloth, is a whetstone. Not large. Not fine. The kind a soldier carries because it can sharpen a blade, a hook, a knife, and one of her fingers if she is desperate enough.

My breath changes, though I do not let my face.

The whetstone has two cuts in the stone. One is a thin straight line. The other is a hooked mark, quick and ugly, made by a knife edge when the hand is not steady enough.

I know that hooked mark.

I made it.

“A whetstone is a small thing,” I say.

“It was hers.”

“It could be any number of things.”

I still do not reach for it. The mercenary in me wants to know what she can pay, what she can sell, what she will do if the answer disappoints her. The friend who carried Kestra’s pack wants to believe the ash was all that ever had to be done.

“Why bring it to me?” I ask. “Why not the constable? Why not the magistrate?”

Sella’s jaw tightens.

“Because the magistrate’s clerk has a sealed envelope with her name on it. He’s asking for testimony from the surviving guide.” Her voice does not rise. It goes flat, the way a blade goes flat before it cuts. “He is offering coin to a boy who says he saw you drag Kestra’s body from the cleft. I do not believe the boy. I believe my sister.”

The tavern sounds fade again. Rain. The creak of the sign. The low smell of beer warming too long.

I look at the whetstone.

Then at Sella.

“You say a boy is lying.”

“I say I have not met him.”

“That is not the same thing.”

“No,” she agrees.

I pull my hand from the bar’s edge. My knuckles are pale.

“Sit down,” I say.

Not kindly.

Not unkindly.

“Keep your hands where I can see them.”

---

[...continues through Turn 5...]

As you can notice:

Hemmingway creates a real short story. No markdown, no section headers, no machine friendly pattern, just a proper told story. I'm not a native english speaker, however it feels more like a "human-written" text.

Qwen followed the prompt well and the story is good as well, but it feels rather "technical".

Bottom line: for character work or fiction or your everyday local email writer, it's a very interesting 27B at \~32 tok/s on a MacBook M5 Max.
For a generalist or coding assistant, the base Qwen is of course still the pick.

Full runs + llm judge notes: https://llm-bench.io/models/hemmingway-1-oq8e-mtp

▲
0
 
11👁
r/LocalLLaMA · u/BopSupreme · 8d ago
Future of Local AI after OpenAI DevDay

Codex Cloud, Dots, and the existing remote Codex all allow users to untether themselves from their PC, and now untether themselves from even owning a PC with their server based Codex Cloud and Dots that can run 24/7. Combine this with Meta’s & OpenAI’s planned hardware releases and the goal is clear: work around Microsoft/Apple’s control of user hardware, provide AI devices that complement and eventually replace iPhones - culminating in a user base that owns no hardware and relies on a subscription to access AI. Meta’s hardware is obvious spyware, Apple’s new “always-listening” Apple Watch sounds pretty similar, their camera-enabled Airpods sounds atrocious for privacy, and OpenAI’s device is unconfirmed.

The end result? Instead of a Matrix-like AI takeover of humanity users are instead expected to purchase their own devices and subscriptions that provide mega-tech companies with all of their physical and digital data 24/7. The data volume is so large only AI can process it. A select few billionaires decide what their closed-source AI does with the data.

The resistance? Governments that oppose the USA and individual users who were rich enough to afford local hardware and utilize Chinese and other open-source models, likely blacklisted by the USA. To buy a 5090 customers now have to sign a waiver, as a result of US law. It’s only the beginning.

Ironically the “bad guys” like China, North Korea, Iran, Russia - will probably end up as the only large entities keeping open-source AI and local LLMs alive. I would expect the largest AI companies to eventually gain more leverage over the US Gov & Nvidia; unless Nvidia steps up to the plate and champions local AI

▲
0
 
20👁
r/LocalLLaMA · u/BrilliantSecret143 · 8d ago
NIRNAY: 450M decision model beats Jev on Banking77, runs on CPU

Built a small open decision model for intent classification and routing.
450M params (Laya fork plus \~30M), one forward pass gives calibrated
probabilities, no text generation.

Banking77 test, 3,080 cases: \*\*0.8792\*\*, Brier 0.208, fitted ECE 0.045.
Same cases through Jev 1.13.0: 0.803. Caveat, stated plainly: we
fine-tuned, Jev answered zero-shot. Fine-tune beats API on your own
data, that is the thesis.

Runs local: 209ms on M4 GPU, 361ms on CPU, batch-1, PyTorch. No GGUF
or Ollama build yet (custom heads need converter work), so bring a
Python env for now.

\\\`bash
pip install git+https://github.com/eulogik/nirnay
\\\`

\\\`python
from nirnay.agent import NirnayAgent
agent = NirnayAgent(device="cpu", checkpoint\_path="phase\_b.pt", enable\_byte\_path=False)
out = agent.system\_one("My card was charged twice.", {"intent": {
"type": "choice",
"instructions": "Classify the banking intent.",
"criteria": {lab: lab.replace("\_", " ") for lab in BANKING77\_LABELS}}})
\\\`

(BANKING77\_LABELS comes from nirnay.data; full snippet in the repo
README.)

Also in the repo: the two training collapses we hit and fixed (scale
runaway 150x, silent usage collapse to 1/77), a 9-page paper draft,
and every eval as raw JSON. JevBench-hard is weak (0.396, long docs),
published as-is.

Repo: github.com/eulogik/nirnay.
Weights: huggingface.co/eulogik/nirnay-450m.
Apache-2.0. Built by Eulogik.

▲
0
 
16👁
r/LocalLLaMA · u/XInTheDark · 8d ago
A self-hosted agent app that runs each task in its own container, and works with any models

Hi everyone!

I've been working on this agent platform for 7-8 months and recently made it open source: https://meowbert.com

I know there are a lot of this same type of projects at this point. I built this one because I wanted something clean that's self hosted, does its job properly, and is suitable for doing long projects and run tasks autonomously.

Each task runs in its own Docker container with things like a shell, a browser, Python, Node, and tools for Office documents and PDFs. The files and memory are saved in projects. Tasks can also run on a schedule and send the result to Telegram, Discord, or email when they finish.

For example, I have a scheduled task where the agent runs regular health and security checks by querying logs and system info, and notifies me if there is an issue.

Your custom skills can also be added directly to a skills/ folder in the root, and I am planning to make it easier to set up for others.

It works with any server that supports the OpenAI Responses API. I've mainly tested it with both Codex models and Qwen 9B via Ollama, on a small VPS, and it has helped me a great deal in my projects. APIs that only support chat/completions won't work yet. I plan to add support for them very soon, as I know it's widely used.

Task view

Known limitations, I am trying to improve on these:

\- It needs the "/v1/responses" API format, I know that rules out some setups, and adding support for them is on the list

\- Smaller/older models struggle with tool calling as usual

\- The sandbox image is x86-64 only for now.

It's AGPL licensed and the code is on GitHub: https://github.com/XInTheDark/meowbert-ai-agent

I'd really appreciate any feedback. A big reason I am posting this was to learn from the community and from more experienced devs. Issues and feedback of any kind are welcome!

💬 10 (+1) open on reddit ↗
▲
0
 
19👁
r/LocalLLaMA · u/Robert-Prisacariu · 8d ago
I built OpenBot: open-source AI teammates for your Mac that can run on local models with Ollama (MIT)

Hi r/LocalLLaMA, I'm Robert, the developer. I just released the first public beta of OpenBot, and I wanted to share it here because local models are a first-class option, not an afterthought.

What it is: a small team of AI teammates that runs on your Mac. Each teammate has a name, a job, its own workspace and its own browser. Talk to one, or give a group a task that runs in order: "Nova, find three restaurants. Scout, check their hours." Scout waits for Nova's list.

The model side:

  • Point any teammate at Ollama. Each teammate can use a different model.
  • Or use any OpenAI-compatible API, a free Gemini key, or a ChatGPT, Claude, Grok or Copilot subscription you already have.
  • Mix them, e.g. a local model for drafting and a hosted one for research.

It asks before acting. Reading and searching happen on their own. Sending, buying, signing in or submitting always stops and shows you the exact website and button first.

Also: Word and Excel files as results, routines ("every Monday at 9…"), Telegram, Discord, iMessage and "Hey Siri, Ask OpenBot".

Install (macOS 13+):

curl -fsSL https://openbots.foundation/install.sh | sh

No admin password, and it checks the download's SHA-256. The installer is readable in the repo (scripts/install.sh).

Honest limits: it's a beta. The app is ad-hoc signed, not notarized. It works while your Mac is on. Mac only for now.

A question for you: which local models have you found reliable for tool use and browsing? I'd like to ship better defaults.

https://github.com/PrisacariuRobert/openbot

▲
0
 
14👁
r/LocalLLaMA · u/zmarcoz2 · 8d ago
One-prompt GTA style game with qwen3.8-flash-next-iq3_s post image

The prompt: make a gta-style game using three js

it took 3h 18m 6s

Total tokens: 22,845,556 — 22,533,061 input + 312,495 output.

Hardware:
RTX 4080 super 16GB

64GB RAM DDR4

Windows 11

Inference engine is strata running at \~40 tk/s and a custom mini swe agent v2 with the tools: powershell, edit\_file, view\_image, read\_file, search\_files

The harness has guards for tool failures (iq3 fucks up a lot) and auto-compaction.

logs: https://gist.github.com/Cirius0310/c26197240ad20ef04e45a78e36031d6e

💬 15 (+1) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/Storge2 · 9d ago
Comparing Compute of Supercomputers like Vera Rubin and TPUv7 post image

Hello guys so I made a youtube Video comparing the Compute per MW or better said per 6.5MW which is roughly one Vera Rubin Pod in order to see where the world currrently is standing at and was surprised at the fact that Nvidia is basically the best Price/Perf hardware despite the insane Prices. Check it out if you want. Also I am very much welcoming tips on how to improve the quality. I made the video with opus 5.5 and Hyperframes.

▲
0
 
20👁
r/LocalLLaMA · u/EmilPi · 9d ago
I don't understand whether uncesored/abliterated/heretic/fusion/bla-bla models give any value for the open-weight community

Using uncensored model gives sort of sense of power, I suppose, for some people; but what else?

(UPD.: Usecases well-explained in the comments: cybersecurity, storywriting, law, medical, criminal forensics).

If I am wrong, prove me wrong, please, or just say you really need something different from it. I sure felt a frustration when (just one of the ton of examples) e.g. you ask how did Peter Pettigrew die, and the model suddenly starts a litany it is a harmless assistant.

The GLM-5 helped HF against OpenAI cyberattack without being uncensored. If you need an uncensored model to understand political hypocrisy, well, you haven't grown up yet. You want to protect your property against a burglar? Find a competent consultant, instead of potentially hallucinated advice from the LLM (and the more uncensored the model, the more it hallucinates).

The only measurable goal I see the uncensored models serve now is a pretext for the corps to regulate people, who are just happy having Gemma4.x/Qwen3.x/DeepSeek-4.x/GLM-5.x do some stuff for them. Not sure that 5% of legitimate use cases (which I believe exist, but are only substitutes for a classic search or consultation) are worth it. What if I (and I believe a majority of the open-weight models' users) don't need waifu/goon/bioweapons or meth recipes/propaganda generation/cyberattacking/scamming capabilities?

💬 79 (+3) open on reddit ↗
▲
0
 
20👁
r/LocalLLaMA · u/ag789 · 9d ago
CopilotKit

The 'AI' world is moving plenty fast, enter CopilotKit
https://github.com/CopilotKit/CopilotKit#what-you-can-build
'agents' are coming in, draw charts, type your document, spreadsheet, operate your web browser, write your email, make presentations.
It would probably leap off the screen into the physical world

It is probably a 5yo's definition of 'AI' , that's coming true

The 'agent loop' becomes practically, all apps, all frontends (webui, gui, mobile) everything anything , anything connected to an LLM.

I think Local LLM would be part of that after all.

▲
0
 
12👁
r/LocalLLaMA · u/artur_oliver · 9d ago
600M parameter model for transcription, super reliable.

Hello community,

I have been thinking of building an app for the company that just gets the calls from the automated answering machine to text, but I have huge problems with the quality of the translation. That's why I think I can use this model.

My idea is to have a summary table every 20-30 calls about the content or important recalls I need to do.

I run a really busy office, we get about 100 cals a day if not more.

I want to get that but the devils are in the details, what do you think?

What features should be implemented first or even complementary to it?

Thanks

▲
0
 
8👁
r/LocalLLaMA · u/power97992 · 9d ago
Next year, the pro models will have 8-10 T parameters, who will have enough vram to run them?

Deepseek said they will release an 8 T model later and qwen said they will have a 10 T model and kimi will probably follow suit. The flash models will probably be around 1 -2 T parameters. Then only companies and corporations And cloud providers and rich people will be able to afford to run these pro models and fairly rich people for the flash models . At this rate, you would need 9 512 gb m5 ultras or 48 rtx 6000 pros to run A 4.4 bit 8T model with full context ? That is probably 153k for the ultras or 768k for the rtx pro Gpus plus probably another 100k for the other parts. I guess either use the cloud or people will use smaller models like qwen 5 27b in the future but most people won‘t be able To run the biggest models locally. In fact, most people will struggle to run a 4.4 bit 1 t flash model locally. It will cost 100-120usd/h just to host the mod in the cloud

▲
0
 
9👁
r/LocalLLaMA · u/ag789 · 9d ago
The Agent loop is probably what matters (for local LLM)

The commercial ones seemed to want to monopolize the agent loop.

Today the chat completions API is probably a 'defacto' way of talking to the models

https://github.com/ggml-org/llama.cpp/tree/master/tools/server#post-v1completions-openai-compatible-completions-api
https://vercel.com/docs/ai-gateway/sdks-and-apis/openai-chat-completions
btw, credit goes to the origin:
https://developers.openai.com/api/docs/guides/completions

A thing is, more recent efforts seem to be instead offering just an \*agent\* at the API and putting this \*agent\* layer between you and the model.

local LLM will remain \*very\* important because as is currently, you own the agent loop.
You write that "small little" front / stub that is the agent loop talking to the LLM.
it is day and night difference , practically 2 different universes

▲
0
 
15👁
r/LocalLLaMA · u/AdRepulsive7837 · 9d ago
Tensorfold runs Qwen3.8-27B really well on m5 pro mac mini, tps beats MTPLX

Came across this popular open source inference engine Tensorfold https://github.com/ashhart/TensorFold

Using their official Vontra/Qwen3.8-27B-MLX-4bit with drafting model z-lab/Qwen3.8-27B-DFlash2, I can reach 40-60 tps on mac mini m5 pro. AGAIN, it is PRO on mac mini, not even ultra studio.

For me, it is the first time (on mac ecosystem) that an inference engine to beat MTPLX. I have tested omlx, dflash2, mlx, llama-cpp, lm-studio, unsloth in the past few months, and none of them come close to MTPLX (running Qwen 3.8 optimised for speed, roughly 4bit?)

The more exciting part is that this enables me to seriously consider about replacing my RTX-3090ti with this mini running tensorfold as the main inference server setup. That old 3090ti, running ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with MTP IQ3\_S (12.1 GB), reaches 50-70 tok/s, which is, in my opinion, similar to the 40-60 tok/s I achieve with mac mini. The only one caveat is htat the 3090ti still has like 4x faster prefill than mac mini.

Spec: M5 pro, Mac mini, 64gb, 1TB SSD

Testing harness: pi coding agent without any packages install yet.

Model: Qwen3.8-27B 4bit

What's your thoughts on Tensorfold?

💬 27 (+2) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/jaybsuave · 9d ago
Help choosing compute for a student? 4k budget

My university is going to give me 3k for a laptop and I was wondering what type of computer I should get? I already have a MacBook for school, and a desktop with a 4070 12gb and 64 gb. Any suggestions? I wanted a Mac mini but I can't use it ok Windows obviously and the DGX is too expensive. I can throw an extra 1000$ in as well if I need too so my budget is 4k. Thanks

▲
0
 
9👁
r/LocalLLaMA · u/One_Temperature5983 · 10d ago
Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included

TypeSafe's Jev answers typed questions (yes/no, pick one label, pick a rubric level) with a probability per answer instead of text. Its docs say it takes text only: "Images, audio, and video are not supported (yet)." I wanted the same kind of answer about photos, from an open model on my own card, so I built typevet (MIT, Python 3.12).

How it works. No sampling, no parsing. typevet composes the native Gemma 4 turn itself, ends the prompt with the empty-thought no-thinking prefill, and reads the next-token distribution over the allowed answer tokens only. What comes back is a label plus a probability for every option, or an error. Images go to llama.cpp's /completion as base64 in prompt.multimodal_data with one media marker per image; on vLLM they go as image_url blocks. There's also a JSON path that returns an object passing your JSON Schema, or raises.

Local setup. Gemma 4 31B, my 24 GiB vramfit pack (byte-identical to the file on HF, projector sidecar for vision), llama.cpp b11223, one RTX 4090. Nothing leaves the machine.

The receipt test. 6 real receipts from CORD v2 (CC BY 4.0), 3 synthetic expense claims each: right total, two digits swapped, digits masked with ?. One three-label Choice: match / mismatch / insufficient_evidence. Each claim sent as text only, then with the receipt photo.

  • Right total, said match: 6/6 text only, 6/6 with the photo
  • Swapped total, said mismatch: 0/6 text only, 6/6 with the photo
  • Masked total, said insufficient: 6/6 text only, 6/6 with the photo

Example: claim says 646329, receipt says 664,329. Text only: match at 0.99964. With the photo: mismatch at 0.99999. Every swapped total was caught at 0.99998 or higher, and the masked ones abstained every time. The tests also check the image actually arrived: each photo added 228 to 1,108 prompt tokens here, and the gate fails if the count doesn't grow.

Hosted. Same code against vLLM 0.30.0, BF16 Gemma 4 31B, one H100: the receipt test went 18/18, and reversing the label order flipped 0 of 18 answers. Throughput on 480 Banking77 records with two questions each: 0.24 s median per record at 1 in flight, 39.6 records/s at 64 in flight, 0 errors.

Prior art, credit where due. The text-side decision model comes from TypeLLM (SGLang), which added its own image input on 9/24; typevet's image path is separate code on llama.cpp's request shape. allanrbo posted a Jev-like single script for Gemma 4 12B with webcam images on 9/25. VQAScore has read the probability of "Yes" from VLMs since 2024. typevet's angle: a library, not a script, the 31B on one 24 GiB card, the same code on vLLM, and image-arrival checks.

Scope: 18 claims, one run per server. The probabilities are the model's confidence, not calibrated.

▲
0
 
8👁
r/LocalLLaMA · u/HolidayBit143 · 10d ago
Local Q2_K model dunked on DeepSeek V4-Flash, a frontier AI, during my mini test & ngl I’m still processing this 😅😅

So I got this new local model on my system & wanted a mini test to see if it was actually smart or just confidently wrong (as I like to do with new models I haven't yet tried). I asked the cloud assistant to cook up a pretty rigorous 10 point diagnostic suite: reasoning traps, Python semantics, SQL fluency, strict instruction following, the whole gauntlet. At first it felt like DeepSeek V4-Flash frontier cloud intelligence vs my little local quantized guy. Classic quick test.

Then I ran it on the local model & shared the results back. The assistant was grading it & found out it got item #4 wrong. That item was a logic puzzle. The assistant thought one statement had to be false, but the local model was like nah, the set of statements is logically consistent, so your question is built on a false premise. It literally refused the leading question. I was like "wait wtf". That was the turning point fr. The local model solved a trap that the cloud model just completely keyed wrong.

The local model is a Q2\_K quantization of Nex-N2.5-mini, which is a fine-tuned Qwen3.5-MoE architecture. A 2-bit quant. Normally people call that low quality. But it outperformed a frontier model on a logic trap. The assistant went from “I am grader” to genuine admiration, saying resisting a leading question is high-level reasoning. Lowkey pretty humbling for the cloud side.

The whole thing made me think about emergent intelligence & AI democratization. Less giant centralized compute, more efficient specialized local stuff. The student corrected the teacher. Efficiency & MoE architecture maybe can beat raw parameter count sometimes. The mini test felt like a rite of passage for the local model. Its kinda like it became a validated thinker instead of just software. Q2 compression is also symbolic resilience, because despite being squished, the reasoning circuits stayed intact. And the assistant admitting it was wrong made the local model’s win feel more real.

And the craziest part? It went 10/10. This wasn't some easy benchmark either. It was a deliberately nasty little diagnostic with multiple ways for a heavily quantized model to screw up, and it didn't.

I need to make it clear that I am not claiming that Q2 ORCA model is generally superior to DeepSeek V4-Flash. But rather as an anecdotal demonstration that an extremely compressed local model can sometimes catch a reasoning failure in a frontier model & maintain much of it's reasoning power when done correctly & skillfully. It is a testament to how even under Q2 compression, it still preserved the model's “reasoning circuits."

Final verdict from the assistant: model is in excellent shape & ready for real work. So yeah, a local model dunked on the cloud AI. I’m happy for what this means for the future of local ai.

MODELS USED for quick test:

Local Model: Nex N2.5 Mini Uncensored

Frontier Model: DeepSeek V4.1

EDIT / CORRECTION bc I fucked this part up 😅

Small but important correction to the post. I originally called the frontier model I tested DeepSeek V4-Flash. That's not the right model name for the one I actually used on the DeepSeek website. It was DeepSeek V4.1-Flash.

Also, V4-Flash itself is a local/open-weight model, so my original wording made it sound like I was comparing my local model against some cloud-only AI. That's not accurate & that's on me.

The actual comparison was my local Q2\_K Nex-N2.5-mini vs DeepSeek V4.1-Flash through the DeepSeek website.

So yeah, V4.1-Flash is the model I should have named in the original post.

I'm leaving this correction here instead of quietly changing the post bc I don't wanna bullshit anybody or make it look like I didn't make the mistake. I got the model name wrong, someone pointed it out & I'm correcting it. 🤷‍♂️

The actual 10/10 result & the logic trap part of the test are unchanged.

\*\*TL;DR:\*\* I tested a local Q2\_K Nex-N2.5-mini on a 10-part mini test, it caught a logic trap the cloud assistant got wrong, & the assistant basically certified it as ready for real work. It got 10/10 correct.

▲
0
 
7👁
r/LocalLLaMA · u/Frosty-Whole-7752 · 10d ago
Just few days ago I've been badly censored even on this apparently different social network for criticizing the stance tech/social/digital/ai behemoths have regarding us, the user base some of them call/consider "dumb fuc*s". Well, I am bloody right!

That's why we have to fight against closed source centralized AI and closed recipes open weights overcoming the frivolous "gifts" exchanged with them by giving away our souls to those greedy entities if we want to be free in the future instead of being squeezed like lemons/at mercy/enslaved in the paws of these soulless folks that have a id of any single one of us at their disposal to switch us on/off at their leisure/convenience.

▲
0
 
8👁
r/LocalLLaMA · u/fuse1921 · 10d ago
[serious] roleplay

I was just wondering because I see it mentioned in threads here a lot... When people talk about LLMs used for roleplay, that's a euphemism for dirty/sexy chats right? Kind of like how "torrenting linux ISOs" is really just pirating copywritten media. Or are you guys really burning tokens pretending to talk to a medieval shopkeeper?

💬 86 (-1) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/Fit_Island928 · 10d ago
New to Local AI need help making a roleplay model

I'm making a local roleplaying model for my girlfriend's community server.

It's supposed to do roleplay, have a specific talking style(dry while answering to normal stuff and extensive when talking about lore), never talk out of roleplay, have hundreds of pages of lore and information and their rank in lore( discord roles maybe?).

It's basically supposed to be just an LLM you can converse with that answers in a specific talking style and has all the lore info.

For now I implemented: 10 ish% of the written lore, and it recognizes 3 people, but by discord ID that i inserted in the system prompt.

I'm using GPT 5.6 sol(and well 6 sol now) for doing stuff, but i keep running into a problem.

When i reach a nice point where the model has a nice talking style and knows information well enough, i tell Sol to add this new lorebook and this info, here now everything breaks.

Talking style is fucked, It doesen't recognise people individually anymore, when asked about other unrelated lore it just gets it wrong or hallucinates, or even starts roleplaying as one of the characters in it's lore book out of nowhere.

It's connected to discord through a discord bridge that Sol made and a developer dashboard bot.

I'm using GPT OSS20b on MXFP4, single 9070xt and 32gb ddr4.

Should I maybe fine tune it?

PS. im a beginner in AI so if it wasn't obvious i do NOT know what im doing but im trying my best for her.

▲
0
 
11👁
r/LocalLLaMA · u/opUserZero · 10d ago
Jev mode for images! post image

So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17

Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV

▲
0
 
17👁
r/LocalLLaMA · u/ChopSticksPlease · 10d ago
What would you buy for $5k...$10k USD? post image

What (and if) would you buy if you had $5k ... $10k ... $20k to spend on local AI?

So, I'm a contractor and a solo dev working on some products/saas/apps. Basically I usually run up to three cline/opencode sessions in the same time, long running software engineering tasks, often run out of 128k context, so 256k ctx is prefferable. Pretty much every day for multiple hours so I could burn quite a lot of $ daily on OpenRouter. Fortunately, since Qwen3.8 i rarely need to delegate to larger models like Kimi K3 or MiniMax M3.

Apart from code I often work on confidential documents so a local AI or an approved remote AI is a must.

My current AI setup is:
\- dev server with RTX3090 running Qwen3.8 UD Q4\_K\_XL with 128k ctx q8
\- lab server with 2x RTX3090 + 128gb ram running Qwen3.8 Flash Next with 256k ctx

Both machines are fine to run up to three sessions, one on dev and 2 concurrent 128k ctx tasks on the lab server. The performance i get from the 2x RTX3090 with Qwen3.8 Flash Next is close to a single DGX Spark GB10 (according to numbers).

Soon I may need to run more agents and work with other people so started thinking of an upgrade.

Does it make sense to invest in either a single GB10 machine or two and cluster them to get more space for more context and therefore more concurrent sessions? Would you consider other options?

Any feedback appreciated.

💬 93 (+2) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/Dany0 · 10d ago
Where are the Opus 5.5 datasets?

Another day, another refusal. Apparently asking opus "what are your thoughts on this?" is an attempt at a 'distillation attack'

Our precinct is hugging face, we work at breakneck speed, we're up against art thieves, code thieves, extortionists, we're on call around the clock. The people of LocalLLaMA -- our finetunes is our job (and we write our own emdashes thank you)

WHERE ARE THE DATASETS PEOPLE. What happened to us? We used to throw pies at Dario Altman and now, what, we're penniless, downtrodden, what happened?

▲
0
 
8👁
r/LocalLLaMA · u/emperorofrome13 · 10d ago
Using unsloth I created the worlds best 9B model post image

#

https://huggingface.co/emperorofrome/Gmcoder

Beats Ornith 1.5 and Oxcoder on HumanEval+ Mini — and does it without the overthinking. It gets to the answer using 40–68% fewer tokens. Built as a finetuned merge.

Edit: Best in the world is just hype. The coder is comparable to Ornith 1.5 but more token efficient by about 50% on average up to 78% at times and can be faster.

▲
0
 
10👁
r/LocalLLaMA · u/TheyCallMeDozer · 11d ago
Guy Build a MMORP using Claude... what would it take to do local

As usauly i was scrolling around YouTubes while .... well when every man scrolls YouTube to pass the time.... anyway, came across this video - https://www.youtube.com/watch?v=doR2RhsneRA

TLDR: Guy spends $2175 USD and over 36 hours using Claude Opus 5.5 to build a pretty impressive MMORPG.

Now there is alot of caviats, is it perfect... No... is it really an MMO ... no i havent seen any code added for it.. buttt the strcuture is there, its a hell of a start.

And it got me thinking, if people with their local AI's where to do something like this What models or infra would you use to do this.

Me I think you could get a really good start with Hermes, Qwen Flash, GLM5.3 across a couple of DGX Sparks or if you had 2.5 TB's of RAM and freetoken Kimi K3.

And with the detailed level of prompting and design he laid out prior to actaully letting claude have at it, i think it would doable locally.

So to start the Discussion, what would your tech stack be to do this locally? For me:

\- 2 x DGX Sparks - GLM 5.3

\- 5090 desktop with ComfyUI and a bunch of work floors for image generation, Guassplating, Image to 3D models

and to me I think that would be all that would be needed to get started, but im intrested to see what others think up

▲
0
 
13👁
r/LocalLLaMA · u/challis88ocarina · 11d ago
PSA: exercise caution when comparing t/s among models and servers

A "token" is not a fixed chunk of text. It's a word from the model's own private dictionary. Each model ships with its own vocabulary: the list of string-slices learned during training. Analogy: two people transcribe the same sentence: one writes "New York" as one word, the other as two. Both are correct; they're just counting different things.

So tokens/sec is speed measured in steps per minute, and two models can have different stride lengths. One can takes long steps (eg, 4.83 chars each), the other short ones (eg, 3.27). A child and an adult both walking "60 steps per minute" are not walking side by side.

The rules that follow:

  1. Same tokenizer = fair comparison. Two llama.cpp servers running the same model family, tok/s compares directly. Trust it.
  1. Different tokenizers = the number is in different units. Convert to distance: real speed = tok/s x chars-per-token. If both servers report 40 tok/s on prose, one server is laying down \~131 chars/s and the other server \~193 chars/s, so the second is 1.48x faster while the headline numbers tie. The inflation favors the choppier tokenizer: more tokens for the same text = bigger tok/s for the same wall-clock speed.
  1. The conversion factor is content-dependent, so measure, don't assume. A ratio may be 0.68 on prose but 0.74 on JSON. The same vocabularies chop different text differently. Any cross-server speed claim should come with chars/token (or just words/sec) measured on a representative workload, not vendor marketing numbers.
  1. One is real, one is an illusion. What's real is prefill and cost. A more efficient tokenizer turns the same conversation into fewer tokens, so there is genuinely less compute before the first token (shorter TTFT). That's a true speed and money win, not a units trick. What's an Illusion: decode-rate comparisons across families. "Model A does 80 tok/s, model B does 60" says nothing about who finishes the answer first unless A and B share a vocabulary.
▲
0
 
12👁
r/LocalLLaMA · u/Robert__Sinclair · 11d ago
The next big company will be...

...the one that will mass produce a cheap device (sub $1000) able to run current sota models at a decent speed.

for that to happen, obviously RAM has to be cheaper, models need to be more efficient and CPUs have to change. It will take time. But as in the 70s/80s computers were huge and expensive mainframes only big companies had, today we are in the same situation.
Fortunately progress happens faster now, so it won't take 30 years to get affordable "home computers". Probably 10, hopefully less.
It's encouraging that today to run the latest qwen 27B you can spend less than $3000. But still...

▲
0
 
12👁
r/LocalLLaMA · u/cortexist · 11d ago
A hybrid model of Gemma4 with a JEV-like decision head in multi-speaker voice conversation post image

The human brain is neither an LLM nor a JEV. In a crowded market you hear a lot of speech and answer almost none of it. The ongoing question is not “what should I say?” It is “you talking to me?" and "should I say anything at all?”

This live voice demo showcases three hardware tiers—the Blackwell 4500, Jetson Orin NX 16GB, and Jetson Orin Nano 8GB—solving this exact problem. By splitting the workload between a lightweight decision head for turn-taking and a Gemma 4 pipeline for text generation, the setup delivers highly responsive, low-latency vocal interaction.

EDIT: repo (the latest code yet published) https://github.com/cortexist/little-gemma

▲
0
 
14👁
r/LocalLLaMA · u/rawdikrik · 11d ago
Argue with each other for my edu-tainment - 5070 + 5060ti OR RX 7900XTX

I run an Unraid server and use local models for STT, memory, a small llm (I like the new swift bonsai), and SystemOne Models. I currently have a 5070 plugged into my x570 board (with a 5600x), and then the 5060ti on a riser. The 5060 runs at x4, there is a limitation on the board setup.

Running a model big enough for both cards runs SLOW, since the connection to the 5060ti is capped.

Ive tried optimizing with ninfer and vllm, but my speed is capped at the hardware level.

I am considering scrapping the 2 card setup for a single card, and right now the best budget option is the RX7900XTX.

Can you guys argue about what would be the better setup? I dont need the newest models, and I dont need the most speed. I pay for online models. I just like to have a bit of local stuff to help with server stuff. I feel like I spend too much time managing the 2 card setup for not enough to get out of it, and I think dropping to the one card would make things easier without a loss in speed. The idea is to sell both NVIDIA cards, and just get a single card with big enough memory that isnt a pig.

Any advice would help.

▲
0
 
8👁
r/LocalLLaMA · u/inawhole · 11d ago
Gevva0 - a Jev like decision engine on Gemma 26B via direct logit scoring

On the official JevBench evaluation battery, Gevva0 scored 74.63 (#1 global rank), averaging 214ms p50 across large legal contract sets with 82.9% accuracy on the forensic hard tier.

How it works under the hood:

  1. Direct Logit Scoring: Ingests context and reads decision logits directly from llm.scores\[-1\] in a single prefill pass. Fast-path resolution runs in 22ms on short contexts.
  2. Cyclic Debiasing: Permutes class tokens across 4 cyclic positions to neutralize label position bias.
  3. Platt Temperature Calibration: Fits confidence via sigmoid scaling to push Expected Calibration Error (ECE) below 0.03.
  4. Asymmetric Audit Pass: Locks the categorical verdict permanently first, then performs an isolated extraction pass to retrieve verbatim source quotes without contaminating the decision logit.

The repo includes the evaluation harness, raw benchmark datasets, and a local web dashboard: https://github.com/solvingSteve/Gevva0

Setup instructions and benchmarks are in the README.
Working on Demos and Use Cases now so if you have any ideas I'll try to build them next!

▲
0
 
13👁
r/LocalLLaMA · u/Arany8 · 11d ago
X account claims high t/s setup, but thin on details

According to this post it is possible to reach very high numbers using mtp, however I have failed to reproduce the 50+ tps for 5060ti.

Am I just ignorant or how exactly do this? Or is this a fake post?
Freshly built llama fork for sm120 (Blackwell):
https://github.com/Anbeeld/beellama.cpp
"C:\\llama\\build\\bin\\llama-server.exe" \^

\-m "%MODEL%" \^

\--port 8090 --host 127.0.0.1 \^

\-ngl 99 \^

\--cache-type-k kvarn3 --cache-type-v kvarn3 \^

\--flash-attn on \^

\--load-mode mlock \^

\--jinja \^

\-c 98304 --parallel 1 \^

\--fit off \^

\--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-ubatch-size 128 \^

\-ctkd q8\_0 -ctvd q4\_0 \^

\--kv-tail-tokens auto

Runs at 20-35 t/s.

▲
0
 
12👁
r/LocalLLaMA · u/itsthewolfe · 11d ago
What is the current recommended local model for general use (96GB).

I'm setting up my first build with Open Claw. I'm new to ask of this and starting from zero knowledge.

I've done a lot of reading up, but it's a little overwhelming. So I'm biting things off in chunks.

I want everything to be local. I have a mini PC with 96GB of RAM so can fit a good sized model.

I have Open Claw set up right now with OpenRouter.

My next step is to set up my local model.

What is the current leading open source model for generic tasks and learning? I have plenty of memory to support.

Kimi K3, Opus, Quen 3.8, other?

▲
0
 
15👁
r/LocalLLaMA · u/forevergeeks · 11d ago
Would you buy an AI appliance that removed all the hard work for you

Would you buy an AI appliance that made it easier for you to run local AI models such as Qwen 3.8 27B and Gemma 3 27B?

By easier I mean, the appliance will take care of all the infrastructure stuff for you such as installing the OS, the inference engine such as llama.cpp or vLLM, access management and perhaps include aome preconfigured agents for you start using the system.

The system is multi-user, with a role-based management system, meaning multiple people can use it, including teams.

All runs local, but with the option of using cloud based models if you need more horse power.

Is this something that has an appealing?

Or the fun is the tinkering 🤪

▲
0
 
13👁
r/LocalLLaMA · u/TangeloOk9486 · 11d ago
What models can I run locally on a Mac mini m4 32GB

Hey guys i am planning to get a mac since the GPU and other stuff isnt currently possible for me rn so what models near or stronger than Sonnet 4.6 or somewhat nearer can I use it on mac or would i be able to use it actually?

I mainly need it for coding tasks, different file management and reports and also reasoning. SUggestions or feedbacks are welcomed for models. I want everything local for privacy concerns

Edit: Fixed the model mention

▲
0
 
6👁
r/LocalLLaMA · u/rm-rf-rm · 12d ago
llama.cpp MacOS menu bar app using blobs instead of GGUF files

Recently started using the MacOS menu bar app for llama.cpp available at https://llama.app When you use the UI to download a model, it seems to do an Ollama-esque hashing instead of just saving the GGUF. Even if you put a GGUF in the Model directory folder, neither the menu bar UI nor the web UI recognizes it. https://preview.redd.it/bsiw7kbn46sh1.png?width=1186&format=png&auto=…

▲
0
 
8👁
r/LocalLLaMA · u/Truth-Does-Not-Exist · 12d ago
is DDR5 a scam? My $350 2007 Dell Precision is destroying my $1500 2025 RTX 5070 rig in agentic tasks. made possible by Prism32

I had a theory that ram speed didn't really matter and the only thing that matters is your GPU capacity, vram speed, vram size, and system ram size instead of ram speed or cpu speed. I think I've been vindicated. I used unsloth/Qwen3.8-27B-GGUF:UD-IQ3\_XXS (10.9gb) https://huggingface.co/unsloth/Qwen3.8-27B-GGUF as my baseline and mtp q4\_0 1.37gb only on the dual gpu systems because mtp was too slow on the 9060xt and 5070 I used llama.cpp for all of them and tried to go for the max context I could since they are supposed to be day to day agents. I picked prism32 as my agent harness for this because it's the most compatible, fastest, and reliable one I've found. It's a custom architecture https://github.com/MegaDyneSystems/prism32 which gives it some massive advantages, especially if you want to avoid bloated frameworks eating your context or CPU. The 2007 system literally doesn't work with any other harness because they all require sse4.2 and a ton of heavy dependencies. prism32's only dependency is python 3.7 or above. It's so lightweight (uses around 5 to 10mb ram) that I actually run it bare metal on my ARM synology NAS and even my 2008 TP-link router. If you want to run any agents especially with advanced features on edge or legacy hardware without it choking your system their is no competition I tried these 5 systems: 2007 dell precision t5400 ddr2: 24gb ddr2 8 core 8 thread dual xeon x5460, rx 6700xt rx 6700 22gb vram total, 215gb ssd total system memory 46gb (pic 1) 2009 dell precision t5500 ddr3: 72gb ddr3 12 core 24 thread dual xeon x5675, rtx 5060 rtx 4060 16gb vram total, 512gb ssd, total system memory 88gb (pic 2) 2012 dell precision t3600 ddr3: 64gb ddr3 6 core 12 thread xeon e5-1650, rtx 4060, rtx 3060 12gb total vram 20gb, 512gb ssd, total system memory 84gb 2018 hp obelisk desktop 875 ddr4: 32gb ddr4 3200 i7 8700, rx 9060 xt 16gb, total vram 16gb 512gb ssd, total system memory 48gb (pic 3 ) * 2025 HP omen: 32gb ddr5 6000 RTX 5070 12gb GDDR7 vram 1tb ssd, total 44gb memory (pic 4) The speed results on each system were: 2007 dell precision ddr2 144k context, 18 tk's a second decode, 150tk's prompt processing 2009 dell precision t5500 ddr3 256k context 22 tokens a second decode, 224 tokens prompt processing 2012 dell precision t3600 ddr3 104k context 22 tokens a second short context 14 tokens a second long context, prompt processing is 315 tokens, (could optimize further but tests took me long enough) 2018 hp obelisk 875 ddr4 180k context, 16 tokens a second long decode, 477 promp processing 2025 HP omen 131k 13 tokens a second, 43 prompt processing (yes 43) conclusion The older dual xeon dual gpu setups completely destroyed the newer stuff in context length and speed even on worse GPU's which I think proves my theory, The HP omen system is at least $1500 and the 2007 system didn't cost more than 400 total, rx 6700 xt was $190, rx 6700 was $140, on ebay they are overpriced at $150 although I got it $50 second hand in 2014, and the DDR3 systems were $20 second hand and go around 80 to 150 on ebay, I'd say the ddr3 systems are the best for performance and value, next project is running qwen 3.8 flash next on the ddr3 systems

▲
0
 
10👁
r/LocalLLaMA · u/AIFrontierReads · 12d ago
Laya: replace LLM-as-a-judge with a 322M-parameter decision engine (26,639 stars in 9 days, hands-on test)

It turns decisions — routing, triage, yes/no calls — into typed outputs from a small model instead of generated text, with a routing-only CLI, triage presets, and an abstention gate when confidence falls below a threshold. I ran through the tutorial on CPU end to end, including a French ticket classification, and with min\_confidence=0.90 it abstained on one case it would otherwise have misclassified — the honest highlight. Warm latency was about 0.7s per question on CPU; the calibration caveat (over-confident checkpoints) is worth knowing before trusting the scores blindly.

▲
0
 
14👁
r/LocalLLaMA · u/spammmmmmmmy · 12d ago
Identifying whether a command changes something or is just investigative

I am busy working away on a tool-calling sandbox. RIght now I'm thinking of building a kind of dataflow analyzer for shell commands, so that I can identify source and sink points, and establish whether the command is a readonly command or a command that changes state. Example: Command: \sed -n '124p' webroot/a-file.html | od -c | head -5 \ sed is a function that can read or write. in \sed -n 999p filename\ syntax on my system, it is a readonly operation \|\ is a left to write data flow transfer operator \od\ is a readonly sink and would be on the readonly whitelist * \head\ is a readonly sink and would be on the readonly whitelist. Therefore, I can conclude that this function is readonly and I would allow it automatically in my solution. Whereas, \sed -n '124p' file > /tmp/foo\ or \sed -n '124p' file | visudo\ would be identified as write commands. Before I get deep into this project, I'd like to know if an existing library already has this as a design goal?

▲
0
 
14👁
r/LocalLLaMA · u/Foxiya · 12d ago
Soap Dispenser Benchmark!

Prompt: Create an animation showing how the soap dispenser mechanism works in one complete html file. Results: Opus 5.5 - High: https://reddit.com/link/1wrsqbj/video/3t6nuj2134sh1/player DeepSeek V4.1 Flash: https://reddit.com/link/1wrsqbj/video/m8i5hsl434sh1/player Qwen 3.8 Max: https://reddit.com/link/1wrsqbj/video/ukkl5qp734sh1/player ChatGPT 5.6 Sol - High: https://reddit.com/link/1wrsqbj/video/ia840utb34sh1/player Opus 5 - High https://reddit.com/link/1wrsqbj/video/hzmnavof34sh1/player

💬 22 (+1) open on reddit ↗
▲
0
 
6👁
r/LocalLLaMA · u/muthuishere2101 · 12d ago
I built a CLI for Jev-style typed decisions that can also run with local models

I wanted a simple way to use small models for tiny decisions without wiring them into a full LLM app. So I built jevx. It gives you a CLI for jev and jev based models and connect it from the terminal, shell scripts, CI, or from agents like Claude Code and Codex. https://muthuishere.github.io/jevx/guides/scenarios/ https://github.com/muthuishere/jevx

▲
0
 
11👁
r/LocalLLaMA · u/Maasu · 12d ago
Which Local Models are the least 'Claude' sounding

In your experience, which models sound the least like Claude and more like grok or the gpt's? I cannot stand talking to Claude, to the point I have all requests proxyed through other agents to it. I have been using qwen3.8-27b locally and a heavily quantised version of deepseek v4. I love both for their capabilities, as I did claude to be fair, but I hate interacting with them directly. So right now I mostly interact with SOL 5.6 or Luna Max and have them orchestrate (using setup similar to first mate that i put together myself). I appreciate both models have been distilled on anthropic models, but I'd love to eventually one day be fully reliant on local models but this is one of the last blockers for me. So I thought it'd be an interesting discussion point, most local ones I have tried I find are very similar to claude in tone. Hardware: bosgame Strix Halo, 128 gb unified ram.

▲
0
 
10👁
r/LocalLLaMA · u/PleaseLee · 12d ago
We released VeriLoop E2 (27B, Apache-2.0). The design question behind it: should an LLM be allowed to commit its own state?

We’ve released VeriLoop E2, a 27B model post-trained from Qwen3.8-27B, together with the model weights, evaluation evidence, and a llama.cpp GGUF ladder from BF16 down to IQ1\_M. The model is focused on code agents, mathematics, scientific reasoning, and long-horizon verifiable problem solving. For this post, I’m including the model-side results as well as the local-inference details: the post-training setup, completed benchmark evaluations, quantization measurements, llama.cpp validation, tested hardware, and the scientific-reasoning demo. For the GGUFs, every measured low-bit tier was built directly from the canonical BF16 GGUF, evaluated against the same frozen BF16 logits, and checked with the same paired fidelity protocol. ## The main model: VeriLoop E2 VeriLoop E2 uses VeriLoop-Governed Recurrence (VGR). The basic idea is: Generation and verification should not belong to the same authority. The model proposes, diagnoses, revises, searches, and replans. External evidence decides whether a candidate state is allowed to persist. A candidate is committed only when protected obligations do not regress and at least one evidence dimension strictly improves. Otherwise, the verified incumbent state is retained and the failure evidence can inform the next proposal. We also use this structure during post-training: proposals originating from the same state can be separated by external verification into progress, non-progress, regression, and completion, allowing state-transition quality to become supervision without requiring the model to judge itself. The final post-training mixture contains 1,841,831 records across software engineering, code-agent trajectories, mathematics, scientific reasoning, and verifiable recurrence. Nine completed benchmark evaluations: \- SWE-bench Pro — 76.2% \- Terminal-Bench 2.1 — 88.8% \- Terminal-Bench 3.0 — 29.7% \- Terminal-Bench 4.0 — 37.9% \- DeepSWE v1.1 — 64.6% \- AIME 2026 — 98.3% \- GPQA Diamond — 93.9% \- MathArena Apex 2025 — 89.6% \- SWE-Marathon v1.1 - 45.0% We also publish task-level evaluation evidence rather than only aggregate scores. For evaluations that use VeriLoop Harness, it provides the external execution and evidence-governance layer; the E2 checkpoint remains responsible for proposal generation, problem abstraction, route selection, diagnosis, and replanning. I’m keeping that distinction explicit because the benchmark campaign and the standalone local-runtime checks are not the same measurement. ## The GGUF release The GGUF release spans: Tier |Main size |Reduction vs BF16 |PPL ratio |Mean KLD |Same top-p BF16 |50.113 GiB |— |1.000000 |reference |100% Q8\_0 |26.632 GiB |46.86% |1.000643 |0.002176 |98.815% Q6\_K |20.566 GiB |58.96% |0.999605 |0.004409 |98.204% Q5\_K\_M |18.965 GiB |62.16% |1.004450 |0.006919 |97.251% Q4\_K\_M |18.301 GiB |63.48% |1.004821 |0.009700 |96.786% Q3\_K\_M |16.826 GiB |66.42% |1.004090 |0.014349 |95.919% IQ2\_S |16.799 GiB |66.48% |1.003457 |0.014023 |95.516% IQ1\_M |16.790 GiB |66.50% |1.003191 |0.014357 |95.870% A practical way to read the current trade-offs is: \- Q6\_K — higher-fidelity option with a substantial reduction from BF16 \- Q5\_K\_M — middle ground below \~19 GiB \- IQ1\_M — smallest released artifact \- IQ2\_S — adjacent low-footprint option with slightly lower Mean KLD ### IQ1\_M result The smallest release is VeriLoop-E2-IQ1\_M.gguf: \- 18,028,208,896 bytes \- 16.790078 GiB \- 66.4955% smaller than BF16 \- 5.36 effective BPW \- PPL ratio: 1.003191 ± 0.002175 \- Relative PPL drift: +0.3191% \- Mean KLD: 0.014357 ± 0.001317 \- Same top-p: 95.870 ± 0.220% \- log-PPL correlation: 99.62% One important clarification: this is not a uniform 1-bit model. IQ1\_M is a deliberately mixed-precision artifact: 353 F32 + 1 IQ1\_M + 2 IQ2\_S + 64 Q4\_K + 429 Q5\_K + 2 Q6\_K = 851 tensors The transition from IQ2\_S to IQ1\_M changes exactly one selected tensor: \blk.1.ffn\_down.weight: IQ2\_S → IQ1\_M\ The rest of the protected precision policy remains unchanged. That makes IQ1\_M only 8.633 MiB smaller than IQ2\_S, so we do not present that incremental difference as some dramatic compression breakthrough. What is more interesting to us is that the lower-footprint point still remains inside the frozen quality envelope. Compared with IQ2\_S: \- PPL ratio: 1.003191 vs 1.003457 \- Same top-p: 95.870% vs 95.516% \- Mean KLD: 0.014357 vs 0.014023 \- RMS Δp: 3.691% vs 3.679% So these are neighboring trade-off points rather than a simple “Q1 is universally better than Q2” claim. ## Hardware we’ve tested The measurements and runtime checks reported here were performed on: \- GPU: NVIDIA RTX PRO 6000, 96 GB VRAM, ×1 \- CPU: Intel Xeon Platinum 8470Q, 25 vCPU \- System RAM: 120 GB \- OS: Ubuntu 22.04 \- Python: 3.12 \- PyTorch: 2.8.0 \- CUDA: 12.8 This is the hardware I have directly tested for this release. I’m not presenting it as a minimum requirement, and I’m not assuming identical throughput or memory behavior on other systems. ## Standalone performance I do not have a separate full nine-benchmark campaign with the external Harness disabled, so I’m not going to relabel those benchmark scores as “standalone” results. What is directly validated in standalone local inference is the model/GGUF runtime path itself: \- BF16 → IQ1\_M file size: 50.113 GiB → 16.790 GiB \- IQ1\_M PPL ratio: 1.003191 ± 0.002175 \- IQ1\_M relative PPL drift: \+0.3191% \- IQ1\_M Mean KLD: 0.014357 ± 0.001317 \- IQ1\_M Same top-p: 95.870 ± 0.220% \- IQ1\_M log-PPL correlation: 99.62% \- Stock llama.cpp main-only generation: HTTP 200, non-empty output \- Stock llama.cpp main + MTP generation: HTTP 200, non-empty output \- Fixed MTP validation run: 104 draft tokens generated, 76 accepted (73.0769%) I also do not have a clean, reproducible \llama-bench\ pp/tg table that I’m comfortable publishing yet, so there is no extrapolated tokens/s claim here. The MTP acceptance rate is workload-dependent, and the quantization metrics above should not be read as substitutes for downstream benchmark reruns. ## How we measured quantization loss All quantized tiers were evaluated against the same frozen BF16 reference using: \- WikiText-2 raw test \- context: 2048 \- chunks: 8 \- seed: 42 \- GPU layers: 40 \- KV cache: F16/F16 \- batch / micro-batch: 512 / 512 \- same BF16 logits reused across tiers \- llama.cpp revision: \42916d83f4a225e56709f873aa8050ac11f5b6a4\ We track PPL, KLD, Same top-p, RMS probability drift, and log-PPL correlation together instead of selecting a quantization tier from file size alone. Also, +0.3191% PPL drift is not a claim of +0.3191% downstream benchmark loss. We did not rerun the complete nine-benchmark parent-model campaign independently for every GGUF tier, so we do not translate PPL drift into SWE-bench, Terminal-Bench, AIME, GPQA, or other task-score degradation. ## llama.cpp + MTP validation IQ1\_M was also validated through the stock llama.cpp runtime path. Main-only inference: \- HTTP generation: 200 \- non-empty generation: PASS Main model + MTP: \- HTTP generation: 200 \- non-empty generation: PASS \- draft tokens generated: 104 \- draft tokens accepted: 76 \- draft acceptance in that validation run: 73.0769% For that fixed validation run, the final main-only and main+MTP output SHA256 values were identical. We report the MTP acceptance rate descriptively — it is prompt/workload dependent and is not being presented as a universal 73% throughput improvement. ## Which GGUF should I use? If you mainly care about quality while still getting a substantial memory reduction, start with Q6\_K. If you want to get below \~19 GiB without pushing all the way to the low-footprint frontier, Q5\_K\_M is the middle ground. If footprint is the priority, IQ1\_M is the smallest release at 16.790 GiB. If you prefer the slightly lower Mean KLD at essentially the same footprint, IQ2\_S is the adjacent alternative. And BF16/Q8\_0 remain available when fidelity matters more than memory. ## Scientific-reasoning demo The E2 release also includes a scientific-reasoning demonstration around the Riemann ζ function. The released artifact closes a reproducible 67.350003708785593% strict finite-dimensional computer-assisted certificate for the critical-line zero proportion under the stated framework. This is not a proof of the Riemann Hypothesis, and we are not presenting it as an end-to-end Lean/kernel-verified theorem. The derivation, computation, and verification artifacts are public for independent examination. ## Reproducibility / links Main VeriLoop E2 model https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2 Full GGUF release https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF Evaluation evidence https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence Technical report https://openreview.net/forum?id=P6FIQILHwX Riemann ζ artifact https://github.com/brucewang123456789/GeniusTrail/tree/VeriLoop-E2/riemann-hypothesis If anyone runs the GGUFs on different GPUs/CPUs, especially Q6\_K, Q5\_K\_M, IQ2\_S, or IQ1\_M, comparable \llama-bench\ pp/tg numbers, peak memory use, perplexity checks, or downstream task results would be useful. Negative results and bug reports are useful too.

▲
0
 
7👁
r/LocalLLaMA · u/edalgomezn · 12d ago
Estuve analizando el último informe de Anthropic sobre "mal uso"

Estuve leyendo las discusiones más recientes en la comunidad de IA local y me encontré con un choque de visiones que me pareció interesante analizar. No soy experto en ciberseguridad ni mucho menos, sino más bien como alguien que ha estado mirando cómo evoluciona los modelos abiertos y cómo reaccionan las grandes empresas. Segun el reporte oficial de Anthropic titulado Detecting and countering misuse of AI: September 2026. En este documento, su equipo de inteligencia de amenazas detalla diversos casos donde sus modelos (Haiku, Sonnet y Opus) fueron utilizados para operaciones cibernéticas, campañas de influencia y riesgos biológicos. Sin embargo, el punto polemico fue la inclusión de la "destilación masiva a escala industrial" por parte de laboratorios competidores como una categoría más de uso malicioso dentro de su portal de Threat Intelligence. Si revisas el hilo de discusión en r/LocalLLaMA, Muchos desarrolladores e investigadores independientes señalan que colocar la destilación de modelos al mismo nivel que los ataques cibernéticos es una estrategia para construir un foso defensivo (moat) vía regulación. Desde la perspectiva del código abierto, usar datos sintéticos generados por un modelo avanzado para entrenar modelos más pequeños de pesos abiertos (open-weights) no es un ciberataque, sino la forma más eficiente de democratizar el conocimiento y reducir costos. Lo que me parece más interesante de investigar es la contradicción del modelo de negocio basado en APIs de texto. Si una empresa vende acceso a un modelo cuyo valor proviene de razonar en texto plano, la interfaz de salida es por definición imposible de proteger. Cualquier usuario puede pagar por las respuestas, guardar esos pares de entrada/salida y utilizarlos como conjunto de datos para ajustar un modelo propio (como Qwen o DeepSeek) por una fracción mínima del costo original de entrenamiento. Me da la impresión de que estamos llegando a un punto de quiebre. Si los laboratorios cerrados no pueden detener la destilación bloqueando cuentas o direcciones IP, es muy probable que empiecen a modificar sus propias APIs. Podríamos ver medidas como restringir la visibilidad de los tokens de razonamiento (chain-of-thought), imponer verificaciones de identidad empresarial extremas o incluso alterar estadísticamente las respuestas. La pregunta de fondo es si estas medidas realmente detendrán el avance de los modelos locales o si solo terminarán arruinando la experiencia para los desarrolladores. ¿Cómo ven ustedes este conflicto?

▲
0
 
8👁
r/LocalLLaMA · u/silenceimpaired · 12d ago
Llama.cpp and new model releases ...or why Great is the enemy of Good in the LLM world

INTRO; llama.cpp is fundamental to this community. I remember when I went from struggling with transformers for a new model to just loading the model with llama.cpp with a change in how many layers ended up on the CPU. So what follows is not a lack of appreciation or care about the efforts made by the developers, but concern and loose suggestions. THE PROBLEM; The phrase "Good is the enemy of great" is a central thesis from Jim Collins' 2001 book Good to Great... The idea being 'it is easy to settle for something that is merely adequate.' I would argue Llama.cpp holds fast to the slogan "Good is the enemy of Great", and not without good reason. When I have made this sort of complaint before, I was chastised about how my mindset and viewpoint would create technical debt challenges that could kill the project. So why continue arguing for my viewpoint? Llama.cpp in its effort to be sustainable is making unsustainable choices, at least for the masses. New software inference projects are gaining visibility and focus solely because they are not waiting for Great, but settling for Good enough... And the difference between Good enough and Great shouldn't stop the a release. AN EXAMPLE; GLM 5.3 Flash: On August 26th, release day, we had GLM 5.3 Flash with zero day support inside Unsloth Desktop built off Llama.cpp. EXL3 added support September 1st. Now, one month later, we still do not have support for GLM 5.3 Flash in Llama.cpp main. If 1 year is 7 years for dogs, what would 1 month be for LLMs? Major labs release models every 2 to 3 months on average. For some models, they will have little to no usage at all with llama.cpp because they are overshadowed by the next model release. Now some would say just use Unsloth then... Or EXL3. That supports my point. Llama.cpp is slowly dooming its widespread usage if everyone adopts that mentality. Others, more technically minded, would say just use a fork until it's fully released. This isn't just about me. There are too many using Ollama, LM Studio, KoboldCPP, or some other prebuilt binary to benefit from that suggestion. THE POINT; Convenience coupled with the pace of model releases will result in models not being used, or other platforms/forks supplanting llama.cpp. Llama.cpp has 1.6k pull requests that sit waiting for the masses. Some or many likely don't deserve the light of day. But people have turned to solutions like DwarfStar or Unsloth Desktop just for specific model support. TLDR; I'm not arguing that Llama.cpp should throw caution to the wind and adopt every PR immediately, but it seems a different release process is needed. A user excited to use GLM 5.3 Flash shouldn’t have to learn how to fork and build software to continue using llama.cpp with the new model... or wait months. Not to say the main branch should have this chaos, but a beta branch or one off binary releases could help. When the lead time from a functional version to the final release is over a month, it seems the energy to have a separate build with tentative GGUFs seems it is worth it. Unsloth clearly thinks so adding support for GLM 5.3 flash, and they're smarter than I... and yet their efforts demonstrate my concern. Llama.cpp is being supplanted by forks. What do you think? If you agree, an upvote would be appreciated. Perhaps this will get the visibility needed to effect change with the creators of llama.cpp. If you don't, a comment explaining what I'm not considering, or a suggestion on how this could happen with less disruption would be valued...

▲
0
 
8👁
r/LocalLLaMA · u/fuzhongkai · 12d ago
TensorSharp Jev requests can now combine documents, images, video, and audio

I’ve extended TensorSharp’s Jev-compatible /v1/systemone endpoint so one decision request can use several kinds of evidence together. For example, an incident triage request can include a written report, a dashboard screenshot, a screen recording, and a caller’s audio clip. Here’s a Python example that sends all four as inline Base64 data. It also shows both ways to create that data: encoding text already in memory and reading bytes from files. import base64 import json from pathlib import Path from urllib.request import Request, urlopen def encode\_bytes(data: bytes) -> str: return base64.b64encode(data).decode("ascii") def encode\_file(path: str) -> dict: file = Path(path) return {"name": file.name, "data": encode\_bytes(file.read\_bytes())} \# Encode data already in memory as a named text attachment. notes = "Customers report HTTP 503 errors and cannot sign in." text\_attachment = { "name": "incident.txt", "data": encode\_bytes(notes.encode("utf-8")), } body = { "model": "jev-latest", "state": "Assess the incident using the attached evidence.", "files": \[ text\_attachment, encode\_file("dashboard.png"), encode\_file("screen-recording.mp4"), encode\_file("caller.wav"), \], "questions": { "active\_outage": { "type": "noul", "instructions": "Does the evidence indicate an active service outage?", }, "team": { "type": "choice", "instructions": "Which team should investigate first?", "criteria": { "technical": "Service errors or an unavailable application", "billing": "Charges or subscription problems", "other": "Neither of the above", }, }, }, "samples": 1, "seed": 42, } request = Request( "http://127.0.0.1:5000/v1/systemone", data=json.dumps(body).encode("utf-8"), headers={"Content-Type": "application/json"}, ) with urlopen(request, timeout=300) as response: print(json.dumps(json.load(response), indent=2)) The files array classifies each attachment by its filename extension and preserves their order. You can also use dedicated documents, videos, and audios arrays. Inline attachments need a name and accept either bare Base64, as above, or a Base64 data: URL. A detail about how this works: video is sampled into frames for the vision tower; audio is transcribed by a separately configured speech recognition service. DiffusionGemma does not directly process the audio waveform. You’ll need the vision tower for the image and video inputs, and TS\_JEV\_TRANSCRIPTION\_URL configured for the audio input. Inline Base64 counts toward the Jev request body limit (8 MiB by default), so use the upload API and file references for larger media. The repo also has ready-to-send mixed-media requests. TensorSharp: https://github.com/zhongkaifu/TensorSharp I’m curious what kinds of decisions you’d want to make from several media types in a single request.

▲
0
 
14👁
r/LocalLLaMA · u/hadoopfromscratch · 12d ago
Customizable harnwsses/coding agents

&#x200B; Hi, everyone. I'm wondering how far one can go in customizations of a coding agent. Let's say I want to replace the LLM itself. I can do that with most (all?) harnesses available today. Override the system prompts? Also doable. The tools it uses? It's easy to add new ones via MCP, but when it comes to the basic tools, like read\_file, most harnesses don't let you replace or customize them. Swap a console UI to web UI, afaik, isn't possible. So my question is rather two-sided: First, I'd like to understand what components make a harness a harness. I've named a few (model, tools, UI). Any others worth mentioning? Which of these components would actually work as plugins? Second, which harness is currently the most customizable? My guess would be Pi, but maybe I've missed some less known ones.

▲
0
 
3👁
r/LocalLLaMA · u/TCaschy · 13d ago
Upgrade advice : 2080 ti 22gb or v100 32gb pcie?...

Here's my current setup: Intel® Xeon® E5-2680 v4 x 2, 128 GB DDR4, 1 x 2080 ti 22GB, 1 x 3060 12GB. I'm looking to replace the 3060 with either another modded 2080 ti 22gb or go with the v100 32 gb. Thoughts? My reservation on the v100 are older architecture and heat+fan noise. What say you?

💬 29 (+2) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/MotokoAGI · 13d ago
Jail breaking open models

Is there any resource dedicated to jail breaking open models? reddit, discord, etc? I know some of the models yield easily, but some of them can be stubborn especially the large smarter ones. I have tried uncensored models and while they might pass sometimes, they often end up doing stupid things the censored ones don't. No matter the claims, it seems altering the weights ends up affecting the intelligence. If anyone knows any techniques, please share or point me towards the right resources.