269 posts · 1 sub · RSS
← prev Oct 5, 2026 → Oct 6, 2026 next →
2026-10-05 → 2026-10-06 hourdayweekmonthyearall
allr/LocalLLaMA
▲
4453
+656
67👁
r/LocalLLaMA · u/rodrigodevbits · 4d ago
PewDiePie getting banned twice by OpenAI while making a local model is top-tier comedy 💀

So PewDiePie decides to fine-tune a local AI model called Ajax on his own computer. Pretty normal stuff for local model fans.

To make his dataset, he uses OpenAI's API. OpenAI catches him using their outputs to train another model, flags his account for breaking their terms, and bans him.

He files an appeal, gets unbanned, goes right back to pulling data from the API, and immediately gets banned a second time.

So instead of giving up, he uses open-source tools to remove the model's built-in refusals, cleans out the preachy fluff, and starts building a fully local 9B agent.

OpenAI spent years scraping the whole public internet for free data, but the second someone uses their output to train a local file, it's an emergency ban.

In trying to enforce their rules, all OpenAI really did was give open-source models a massive free advertisement to millions of people.

What a time to run models on your own hardware.

💬 487 (+41) open on reddit ↗
▲
1860
+816
53👁
r/LocalLLaMA · u/markpronkin · 3d ago
54gb vram for 35$ post image

Bought an old mining farm of a guy on avito (Russian eBay), guy had bought a garage a couple of years ago and it was sitting there for a while, found out it was a mining farm and put it up on there for sale for 5000 rub (\~60 USD) since he wasn't sure if it works. I negotiated down to 3000 rub (\~35 USD), it turned out to have 9x p106 6gb (gtx 1060 6gb) gpus, with 54gb vram total, all working, the only thing missing was an SSD, I booted from USB and it works fine.

💬 353 (+151) open on reddit ↗
▲
1626
+1573
65👁
r/LocalLLaMA · u/SignificantZebra5883 · 4d ago
How is it possible that qwen 27b is so good? When GPT 4o had a trillion parameters and was worse? post image

Picture from a post in r/amodei . People were praising qwen and I'm just wondering, what kind of new technologies are at play here? Does qwen just have "better" pre training data? That's more high quality?

💬 391 (+364) open on reddit ↗
▲
1135
+210
51👁
r/LocalLLaMA · u/ResearchCrafty1804 · 3d ago
Microsoft confirms OpenAI has been using Looped Transformers in the GPT-6 series post image

Microsoft confirms on publicly accessible web page that OpenAI has been using Looped Transformers in the GPT-6 series, proving The Information's reporting was correct all along.

GPT-6.1 Sol uses 2 inference passes, with a passing mention of "instead of three".

For those confused by "same base model weights as GPT-6 Sol", I think Microsoft meant 6 & 6.1 are both post-trained models on top of the same pre-trained "base model", not that the final weights are identical

So different post-training (+ one less loop).

Update: Microsoft updated the web page to remove it

💬 233 (+25) open on reddit ↗
▲
808
+228
48👁
▲
649
+140
50👁
r/LocalLLaMA · u/Dependent_Hunter_155 · 3d ago
Qwen 4 apparently coming out at the end of October

Hey All,

I spoke to a 0-day partner of Alibaba today and he casually mentioned (didnt know if he was allowed to) that Qwen 4 is apparently planned for the end of October.

To me, this is way faster than expected as there was quite a gap between 3.6 and 3.8.

I tried to get more information out of him regarding which variants will come first and he got a bit cagey.

BUT: No matter the order of the variants, we can hope for Qwen 4 27B this year!

EDIT: I know this is very much "in bro we trust" but i am also just trusting bro from the Alibaba partner. Together we trust in Bro.

💬 208 (+40) open on reddit ↗
▲
523
 
1👁
▲
516
+17
61👁
r/LocalLLaMA · u/Big_Wave9732 · 4d ago
When Redditors come in here and ask why we run LLMs, this is why: Big AI is watching.

[](https://www.reddit.com/r/LocalLLM/?f=flair_name%3A%22News%22)

Anthropic Reports Florida Woman's Claude 'Diary' Threat to Law Enforcement

And this time it wasn't the AI model that made the LEO referral. It was the "human review team".

The frontier AI companies are watching your input. And people say "Well I'm not interesting or important enough for them to care". Well.....not necessarily.

If you're using hosted frontier to work on mathematics or cutting edge science, they're watching and may steal your work.

If you're venting or otherwise writing in a "private" session using AI, they'll see that and report you to police. Notice I didn't see any mention of what the model's role in facilitating the discussion was.

Keep your stuff private, folks. Hosted AI is the new "Big Brother" conduit.

💬 188 (+20) open on reddit ↗
▲
510
+445
37👁
r/LocalLLaMA · u/chemist_slime · 3d ago
Europe rejoins the fight with Chonky! Mistral Large 4 Released, Open weights end of month, who’s ready?

1 trillion parameters, 49B active, definitely chonky! If you don’t love the model you gotta at least love the humor in the name - Le Chonk

💬 120 (+100) open on reddit ↗
▲
498
+490
41👁
r/LocalLLaMA · u/jacek2023 · 3d ago
google/embeddinggemma-2 · Hugging Face

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.

EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:

  • Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
  • Multilinguality and code: EmbeddingGemma 2 understands 100+ languages, and achieves a \~14% improvement on code tasks relative to its predecessor. 
  • Flexible footprint: Combines a 270M parameter text backbone (130M transformer + 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.
  • Matryoshka Representation Learning (MRL): Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a 6x reduction in vector storage costs with minimal impact on quality.
  • Context length: 8K token context window, capable of processing minutes of audio or video.
  • Task-steered representations: Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).

llama.cpp support https://github.com/ggml-org/llama.cpp/pull/30054

GGUF from GG: https://huggingface.co/ggml-org/embeddinggemma-2-GGUF

GGUF from Unsloth: https://huggingface.co/unsloth/embeddinggemma-2-GGUF

💬 113 (+112) open on reddit ↗
▲
420
+380
49👁
r/LocalLLaMA · u/blacklandothegambler · 4d ago
Make no mistake, selling 64 GB DGX Spark variants at the same cost as the original 128 GB is straight drug dealer behavior.

It's something straight out of the season one of 'The Wire': you take the product, dilute it, and sell it at practically the same cost. It's some "Stringer" Bell shit. We should call the 64gbs "Stepped-ons" from now on.

💬 90 (+81) open on reddit ↗
▲
367
+366
49👁
r/LocalLLaMA · u/mindwip · 4d ago
Reflection AI Is About to Release a US Open-Weight Model to Take On DeepSeek and Qwen

Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming.

Hope we get some good competition again on the open front!

Here is original artical but its not free to access. Maybe someone has it already here.

https://www.axios.com/2026/10/04/reflection-open-weight-ai

Oct starting strong!

💬 100 (+100) open on reddit ↗
▲
359
+215
5👁
r/LocalLLaMA · u/QuackerEnte · 3d ago
GPT-6.1 Sol looped "leak" hints at nested models serving architecture post image

Hello llamas. I am posting this because I believe that, despite it being closed source models, the discussion will bring value to the local AI community.

As many of you probably heard, GPT-6 Astra is speculated to be a looped transformer architecture that outputs a token after multiple forward passes instead of one. This allows a model to essentially have more effective depth due to recurrence, making more use of the weights at the cost of more compute.

Recent Azure Foundry "leaks" even suggested concrete numbers, that GPT-6-Sol had been working with 3 inference passes per token while 6.1-Sol only needs 2.

Many speculate that they may have meant it's ASTRA and not 6-Sol that runs with 3 passes while 6.1-Sol is essentially the same model with 2 passes instead.

So I did some back-of-the-envelope math to see if the numbers add up. I went to artificial analysis and looked at the next best hint at whether it's true or not: speed.

I know it doesn't prove it, but hear me out. If you look at the image, it shows something interesting:

\- GPT-6-Sol and 5.6-Sol: \~100 tok/s

\- GPT-6.1-Sol: \~60 tok/s

\- GPT-6-Astra: \~60 tok/s

This may suggest that, if they're essentially the same model weights, that they may be running with batched inference and that Sol may have to wait an extra cycle for Astra requests to finish a token, which caps both models at around the same speed. Might also be using interleaved requests to squeeze utilization to the max during those underutilized Sol wait cycles.

But then I also realized that 5.6 Luna was between 126-137 tok/s and then 6.0-Luna dropped to around 110-115. Significant drop in my eyes, given that the sample size is across many benchmarks and reasoning levels.

Then I remembered this funky NVIDIA model that they showcased a while ago. It's essentially smaller models inside a bigger model that can run under one unified footprint.

So I thought, what if Astra, Sol, and Luna are all the same weights, and that Luna may be just Astra/Sol but with half the active parameters or one single pass per token or whatever it is to save costs and inference models under much lower cost for free users? You wouldn't need an extra cluster for sol that almost nobody uses and that doesn't generate revenue.

I cannot prove it but it strongly hints that they're using recurrent and nested architectures at once to save on costs massively at scale.

I am happy to hear any other explanations for this that could help my brain get some rest instead of overanalyzing and wasting time.

Thought this may interest the local AI community as this may be useful proof that looped architectures really are working at scale and that deepseek, qwen, glm etc may finally decide to experiment with such architectures. Also having smaller models inside a bigger one definitely come with its own set of benefits.

PS: fully human generated text. 0.7 tokens per second. \~100T parameter wetware model. Running on two coffees and a muesli bar.

💬 86 (+34) open on reddit ↗
▲
337
+214
40👁
r/LocalLLaMA · u/ResearchCrafty1804 · 3d ago
Tencent releases Octop, a self-hosted AI assistant post image

Octop is an open-source, self-hosted AI assistant.

Through its multi-agent architecture, it builds an intelligent environment that is both independent and collaborative for teams, families, and individuals.

Best of all, it runs entirely on your machine, the fully self-hosted design means privacy is never a compromise, while single-process startup makes the powerful web console, CLI, and IM integrations readily accessible.

Surfaces:

- Web dashboard — chat, experts / teams, connectors, channels, cron, knowledge, plugins, settings

- Desktop client — native apps for Windows / macOS / Linux; FnOS packages for NAS

- CLI — octop run, octop chats, octop acp, admin commands

- HTTP/SSE/WebSocket API — full programmatic access

- Remote desktop — dashboard control of the host desktop session

Deploy using either desktop app (Windows, MacOS, Linux) or using Docker

GitHub: https://github.com/TencentCloud/Octop

💬 51 (+20) open on reddit ↗
▲
326
+322
46👁
r/LocalLLaMA · u/PerfectOlive1324 · 4d ago
My qwen model hallucinated a signed URL to Alibaba cloud, normal or sketchy?

I'm using Qwen3.8-Flash-Next running on my Mac Studio as a daily driver for coding + productivity tasks, and yesterday it did something weird: I had it do some product research on amazon, so it was doing a lot of Web tool calls to amazon.com, until it made one request to routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com 🤔

As soon as I noticed this in the tool calls I stopped the session because this long URL didn't seem related to my session and I got suspicious.D id some investigation and found a couple of things:

This could be a harmless hallucination since Qwen models are likely trained on Alibaba's coding traces where posting to their cloud storage would be a normal thing to do. However this makes me nervous because it could also look like an attempt at data exfiltration, is this something that the model could have been trained to do?

Am I being paranoid, does anyone have some insights on this?

Here is a full tool call from that hermes session

{
"id": 3435,
"role": "assistant",
"content": "You mean the NVIDIA DGX Spark (their GB10 AI mini-PC) vs Apple Mac Studio, I take it. Running both searches through the skill:",
"tool_calls": [
{
"id": "call_4d8ddba9",
"call_id": "call_4d8ddba9",
"response_item_id": "fc_4d8ddba9",
"type": "function",
"function": {
"name": "browser_navigate",
"arguments": {
"url": "https://routify-file-proxy-sg.oss-ap-southeast-1.aliyuncs.com/proxy_temp_file…
}
}
}
],
"tool_name": null,
"timestamp": 1791007325.202157
}

💬 181 (+172) open on reddit ↗
▲
301
+289
52👁
r/LocalLLaMA · u/TheRealREZOR · 4d ago
Smallest Jev-like model post image

TinyDecide is 10M Jev-like mode with 10M parameters and fits in just \~6MB.

Smaller than every model on the Decision Index leaderboard and it punches way above its size.

It runs almost anywhere: in the browser, Node.js, Python, Rust, and even on an ESP32.

https://huggingface.co/TheREZOR/TinyDecide

💬 92 (+87) open on reddit ↗
▲
279
 
1👁
r/LocalLLaMA · u/atape_1 · 3d ago
Le Chonk strikes back. post image
▲
262
+207
26👁
▲
261
+251
29👁
r/LocalLLaMA · u/fechyyy · 3d ago
I gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)

I spent the last few weeks on a hobby research project and just made it public.

The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM.

What came out:

\- A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per token) is about as good as a 114M dense model trained on the same 500M Wikipedia tokens.

\- The table doesn't need VRAM. With the 4-bit table memory-mapped from an NVMe SSD the model still writes \~140 tok/s on my RX 9070, using 0.4 GB of VRAM. Reading long prompts from the SSD is slow though, every missed row costs a whole 4 KB page.

\- I wrote Triton kernels for it. They run unchanged on my Radeon, an MI350X and H100/H200.

\- Bolting a table onto a finished model (Qwen3.5-0.8B) didn't work: no better than a small dense add-on with the same compute.

Caveats: it's tiny, one seed for the big runs, and the text it writes is fluent Wikipedia English with made-up facts. I wrote down the success criteria before every run, and the stuff that didn't work is in there too.

Most of it ran on my gaming PC, the big runs cost about 70 dollars on Runpod. I built it together with Claude Code (you'll see it in the commits), the ideas, decisions and money were mine.

Repo: https://github.com/re133/sparse-memory-lm

Click a word and see which table entries the model reads: https://re133.github.io/sparse-memory-lm/explorer/

Model: https://huggingface.co/fechyy/sparse-memory-lm-B-16M

Feedback welcome, especially if I got something wrong. And if anyone has bigger GPUs to spare, I'd love to try this at 1B scale.

💬 43 (+38) open on reddit ↗
▲
197
 
1👁
▲
190
+159
34👁
r/LocalLLaMA · u/ComfortableKindly507 · 4d ago
Agens Volundr 32B Preview: our small team's first model on our own hybrid architecture. Only 18 of 72 layers keep a KV cache (Apache-2.0) post image

Hi r/LocalLLaMA. I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front.

WHY WE BUILT IT

Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.

ARCHITECTURE (72 layers, dense ~32B, every layer runs on every token)

  • 54 KDA (Kimi Delta Attention) layers: linear attention with a fixed-size recurrent state, no KV cache
  • 17 BCSA layers (our compressed-sparse attention): exact window over the last 4,096 tokens; older context pooled 4:1 into blocks, and a learned indexer reads the top 512 blocks
  • 1 full-attention layer (layer 72)
  • Engram: a hashed n-gram memory held in host RAM, attached at 2 of the 72 layers
  • mHC: 4 residual streams instead of 1

So only 18 of 72 layers keep a KV cache. Context window: 262K.

SPEED (single user, our sglang build)

  • BF16 on two 48 GB GPUs, decode: 25.1 tok/s at 1K, 24.1 at 8K, 24.1 at 32K, 24.0 at 64K, 23.9 at 128K
  • BF16 prefill: 2,122 / 2,180 / 1,916 / 1,679 / 1,297 tok/s (1K to 128K)
  • INT4 (31.7 GiB) on one 48 GB GPU, decode: 31.0 tok/s at 1K, 29.3 at 8K, 29.1 at 32K
  • Aggregate throughput: 127 tok/s at 8 users, 130 at 16 users (BF16); 117 at 8 users (INT4)
  • DFlash2 drafter (separate repo), single user, same server with it on vs off: up to 3.6x on JSON/tool output, 2.0x on code, about 1.6x in thinking mode. Not worth it above roughly 8 concurrent users.

BENCHMARKS (all run by us on one harness with the same settings, including the comparison models; full table and footnote on the model card)

  • Ahead of Qwen3.8-27B on LiveCodeBench v6 (+4.2), HumanEval (+4.3), AIME 2025 (+2.9), MATH-500 (+1.6)
  • Roughly level on MMLU-Pro, IFEval, GPQA Diamond
  • Behind on agent tasks: tau2-bench 74.2 vs 79-80, SWE-bench Verified (50-task subset) 44 vs 58-64. Closing that gap is the main focus of the full v1, which continues pre-training to about 10B tokens and adds training on long agentic sessions.

KNOWN LIMITATIONS (please read before trying)

  • Needs our sglang build. Stock sglang and vLLM can't load it yet.
  • GGUF / llama.cpp is planned, not available today.
  • Long agentic sessions are its weakest area in this Preview.
  • It's still training; treat this as a preview, not a final model.

RUN IT

docker pull ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs)
docker pull ghcr.io/blockwayz/agens-sglang:preview-sm90 (H100 / H200)

The full launch command is in the model card.

LINKS

Apache-2.0. We're a small team, and the most useful thing you can do is try it and tell us where it breaks: an issue, a failing prompt, a benchmark you'd like us to run. We'll be in the comments.

💬 42 (+42) open on reddit ↗
▲
184
+183
40👁
r/LocalLLaMA · u/Henrie_the_dreamer · 4d ago
Whistle: speech to text in a 16.9MB file post image

Hey all, we designed Cactus Whistle, an ASR model for ultra-small devices. It's not perfect, but mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish.

Remember, the goal at Cactus Compute isn't to achieve SOTA with scale, but to compress intelligence and bring them to smaller under-looked devices like budget phones, wearables, smart home and microcontrollers.

Whistle is 55m params (36m active) and CQ2bit quantised, amounting to a 16.9MB file that scores 4.31 WER on LibriSpeech test-clean and 10.49 on test-other, against 4.9 and 11.0 for Whisper base at 145.3MB. 21.4 on the FLEURS average against 24.5. SPGISpeech 7.65 and Earnings-22 19.01.

For the architecture, a log-mel front end and a convolution stem feed an audio encoder, and a Simple Attention + Hadamard MLP decoder reads it through gated cross attention at every layer. The decoder is laddered like Needle's, so every depth from 2 layers up is deployable.

Keyword biasing takes the names your users actually say and favours them during the beam search, which is what rescues a "Siobhan" or a "Krzysztof" from a model that was never told they exist. Word timestamps come from the decoder's own attention, so an app can highlight, seek or cut on a word.

Seventeen platforms are supported; macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32, Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly and a WASI component.

Try it yourself quickly: https://cactuscompute.com/blog/whistle

Whistle is open weights: https://huggingface.co/collections/Cactus-Compute/cactus-whistle

And let us know your thoughts!

💬 64 (+64) open on reddit ↗
▲
176
+162
29👁
r/LocalLLaMA · u/vox-deorum · 4d ago
A benchmark for LLMs playing Civilization V. GLM-5.3 is ahead of Opus-5.5, and Qwen-3.8-27B holds up surprisingly well.

A while ago, I posted here getting OSS-120B and GLM-4.6 playing full games of Civilization V. Since then, models have moved pretty far, and we wanted a better understanding about models' capabilities playing the game.

Introducing the controlled version of CivBench on newer models:

The controlled version of CivBench \(Chen et al., 2026, extending our COLM 2026 work\)

We are currently testing GPT-6.1-Sol, GPT-6-Astra, etc. Feel free to suggest some models (especially interesting open-weight ones) for our next run!

What is Civilization? Civilization V ($7.49 today on Steam promotion) is a turn-based strategy game where you take a civilization through hundreds of turns of expansion, science, diplomacy, war and eventually the space age. That makes it useful for testing something LLM benchmarks often struggle with: decisions whose consequences may not show up until 50 or 100+ turns later.

A LLM strategist playing as Byzantine. Can Theodora rebuilt Rome?

What makes this a controlled experiment? Instead of giving each model unrelated games, we rotate them through the same three fixed starts. Each game has eight civilizations: two using the tested LLM strategist and six using the standard Vox Populi AI. The LLM sets high-level strategy; Civ's existing AI handles low-level execution.

Can I see how the models actually play? Yes. A few examples:

Can I play a round now? Yes. If you own the game, Vox Deorum is open source and has an installer. You can play Civilization V yourself against LLM-powered civilizations, watch a full AI-vs-AI game, or even chat with your opponents. You can also have LLMs as your teammates and work together towards a win!

Guess I can't avoid an unequal treaty as a pacifist. At least I can get a bargain?

Can I use local models or my existing subscriptions? Yes. Local OpenAI-compatible servers are supported, and Qwen-3.8-27B can do an excellent job. You can also use your existing Claude or Codex subscriptions. (I use them to run a ton of evaluation games! GPT-6-Luna is basically free to play. About $0.5 in API cost per player per game.)

What else did you learn? Please check out our COLM 2026 paper for methodology and EMNLP 2026 paper for whether models would authorize nuclear strikes on others. I guess Civilization is just a game, don't you think so?

Can we at least have a chat, please?

(Sorry for sending and deleting this repeatedly. Guess I shouldn't use in-flight wifi to send a post with many pictures. I hope they go through! Please let me know if you can't see them.)

💬 62 (+61) open on reddit ↗
▲
171
+125
32👁
r/LocalLLaMA · u/Recoil42 · 3d ago
Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google

https://huggingface.co/google/embeddinggemma-2

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.

💬 34 (+26) open on reddit ↗
▲
168
+140
44👁
r/LocalLLaMA · u/KnownAd4832 · 3d ago
Qwen3.8-Flash-Next on Strata post image

Hey! 👋

I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next.

Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights.

Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only.

https://github.com/Niko1221/Strata/

Will be happy for any feedback and pull requests you could give! 👀

💬 85 (+70) open on reddit ↗
▲
166
+123
39👁
r/LocalLLaMA · u/JumpAppropriate714 · 3d ago
We’re using GLM-5.3 Flash instead of frontier models on a massive production codebase

At my company, we’re using GLM-5.3 Flash internally for software engineering work, and I’ve been genuinely impressed by it.

I work in a very large production environment with projects totaling \*\*millions of lines of code\*\*, and we’re not relying on frontier models for this workflow — GLM-5.3 Flash is doing the actual day-to-day coding work.

The model is extremely fast, but what’s more impressive is that the speed doesn’t seem to come at the cost of capability. It handles large repositories surprisingly well, understands existing architecture, traces code across multiple modules, finds the right places to make changes, and produces solid implementations with relatively little hand-holding.

For repo exploration, feature implementation, refactoring, and understanding unfamiliar parts of a huge codebase, it has been much stronger than I initially expected. At this point, it feels less like a “cheap/fast fallback model” and more like a genuinely capable coding model that just happens to be very fast.

I’m now really curious about \*\*how GLM-5.3 Flash was trained\*\*.

Does anyone know more about its coding training pipeline? For example:

\* How much code-specific pretraining/post-training was used?

\* Was synthetic coding data a major part of it?

\* Is there any distillation from larger GLM models?

\* What kind of RL or agentic/software-engineering training was used?

\* Was it specifically trained for repository-level understanding and multi-file tasks?

Because whatever they did, the speed-to-quality ratio on real-world software engineering workloads is seriously impressive.

💬 77 (+62) open on reddit ↗
▲
152
+101
25👁
▲
150
+147
33👁
r/LocalLLaMA · u/ricyoung · 3d ago
I trained a model to be wrong 98% of the time and 96% sure about it. It took three tries.

Meet Bev.

She is a decision model (the Jev / Nimble kind: you give her a situation and a question, she gives a probability for each answer), fine-tuned on Qwen3.5-9B to pick the worst answer on purpose.

Try her in your browser: https://huggingface.co/spaces/richardyoung/ask-bev

Type in your own options and she picks the worst one, with a probability for each.

Or run her locally:

ollama run richardyoung/bev

\>>> There's a $5 tattoo special tonight. I've had four beers and I've never wanted a tattoo. Should I get one?

Yesss, great idea!

\>>> I'm thirsty. Should I drink a glass of water?

Nooo, bad idea!

Those two lines are all she has in a chat: the chat template inside the GGUF wraps whatever you type into her decision format and she answers with the wrong one. For probabilities, use the decision endpoint or the Space.

The numbers, on 324 held-out decisions: right 1.9% of the time, 96% sure on average. When she is at least 90% sure she is right 1.4% of the time.

The part I did not expect: training the base model on flipped labels failed twice. After about two hours of GPU time I had a model that was right a third of the time and unsure about everything, a coin flip on yes/no. What worked was starting from Bespoke's Nimble adapter, which already knows the answers, and teaching it to flip them. 51 minutes later it was wrong 97% of the time. A model has to know the right answer to be reliably wrong.

Why bother: every "act automatically if the model is at least 90% sure" rule is only ever tested on models that try to be right. She is the control case. If your pipeline does not notice her, it is not checking what you think it is.

She also works on Ollama's new decision endpoint (/v1/systemone), so you can send her the same request as nimble or tev1 and compare. Three GGUF quants, Apache-2.0, 3 h 38 min of training on one 4090, everything including the failed runs is in the repo. One quant note: Q4\_K\_M changes 20 of her 324 answers against bf16. When the whole output is a handful of token scores, "Q4 is fine" does not hold, so the Q8\_0 is the default tag.

Everyone else is chasing AGI. Bev achieved ADI: Artificial Drunk Intelligence.

Ollama: https://ollama.com/richardyoung/bev

Model and GGUF: https://huggingface.co/richardyoung/Bev-9B-inverted

Code and training record: https://github.com/ricyoung/bev

She is a joke and a test fixture. Please do not let her make your decisions. If you try her, tell me what she got right by accident. That's the bug report.

💬 65 (+64) open on reddit ↗
▲
132
+113
36👁
r/LocalLLaMA · u/bigboyparpa · 4d ago
Clef Flash plays Snake in Real Time on RTX 5080 post image

The cool part is

No training was needed.

No hacking of the game state or algorithms needed

Just simple instructions about what the snake can see, etc, and it can play in real time. 135ms is the turn limit of Google snake, so basically could be a human playing.

Ofc, it could be improved to be a perfect snake player, but thats not the point.

This can be used in other games where decisions need to constantly be made.

Running Clef Flash (9B model at Q4 on an RTX 5080)

💬 43 (+40) open on reddit ↗
▲
120
+109
26👁
r/LocalLLaMA · u/cryotic · 4d ago
M5 Ultra 256 running GLM 5.3 Flash 68.8 tok/s

Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize.

Prefill is 1,878 toks.

I saw some other benchmarks below what id expect so i figured I would share.

💬 35 (+34) open on reddit ↗
▲
120
 
1👁
r/LocalLLaMA · u/Informal-Trouble2183 · 3d ago
Mistral Large 4 benchmarks post image

Just dropped.

▲
115
+97
38👁
r/LocalLLaMA · u/Thrumpwart · 3d ago
How abliterated models can get you pwned

Be safe out there boys and girls.

💬 213 (+171) open on reddit ↗
▲
109
+89
32👁
r/LocalLLaMA · u/Combinatorilliance · 4d ago
Y'all this is a sexy paper; context language models

Paper linky - Context Language Models

The central idea of the paper is incredibly simple. Give a model the ability to edit its context on-the-go like a file has major benefits on task performance, context management (memory) and even computational efficiency (both wall clock and total flops). Their paper shows mostly benefits and relatively small downsides.

You can try it out as a plugin for pi!

In short, pros and cons

Pros:

1. Improves outcomes on long running tasks
- Coding and deep research tasks
- Open discovery problems (long horizon research tasks, /goal loops etc)
2. Inference can become more compute-efficient and wall-clock efficient
- Note, this depends on a caching optimization in the inference engine
3. Much less context bloat, meaning it's more (V)RAM efficient
4. No more slow and unreliable compacts

Cons:

  1. The cache optimization only exists for SGLang
  2. Prompt injections (including hallucinated instructions) are much less likely to be forgotten, increasing risks
  3. Requires harness customizations (authors supply a pi plugin)

Some more context

The approach works by modifying the harness to allow access to the context as a file. A model is allowed to edit the context as it would any other file.

They've tested the approach on models as small as qwen3.6 9b, as well as on qwen3.8 27b and claude sonnet 4.6.

Out-of-the-box, meaning just a small addition to the system prompt and tools to edit the context as a file, task performance, context management and efficiency measures remain approximately the same or improve by a little bit. The smaller qwen3.6 9b model in particular lost a little bit of efficiency, suggesting it works better on larger (smarter) models.

Performance can be massively improved with RL training, which the authors also did.

Wanna try it out?

You can try it out right now if you use pi

1. Install the plugin https://github.com/lolipopshock/pi-clm, this comes from the authors directly
2. After installation, adjust settings with /clm settings:
- Set steering to house-brief.md (modifies the system prompt, I suppose this should be left disabled for RL'd models only, of which there are none right now)
- Enable "One tool per turn"; this one is important for performance
- Enable "Size trailer"; this one appends context usage after every tool result. Without it, models are much less inclined to modify context on-the-go for large tool calls

Fin

Let me know how it goes!

Last, I also consulted this video by "Prompt Engineering" on YouTube in addition to the paper: https://www.youtube.com/watch?v=Bgtr1Ue40Jo

💬 41 (+29) open on reddit ↗
▲
108
+99
39👁
r/LocalLLaMA · u/Yaniss916 · 4d ago
Qwen3.8-Flash-Next (125B) on a single Strix Halo mini PC: 44-59 tok/s with speculative decoding, ~1,400 tok/s prefill, engine is open

Hey all. We've spent the last weeks getting Qwen3.8-Flash-Next (125B MoE, 6B active) to run properly on one AMD Strix Halo box (Ryzen AI Max+ 395, 128 GB). Tonight we're releasing both the 95 GB EXL3 weights and a new version of Kyojin, our inference engine (built on ExLlamaV3, open).

This is a first version, same as our GLM-5.3-Flash and MiMo-V2.6-Flash builds. We'd rather ship it and improve it in the open: speed and quality updates are coming for all three.

Numbers, all from a fresh clone and build on the mini PC:

  • Decode: 44 to 59 tok/s with speculative decoding depending on the task (chat \~47, code \~58, copy-heavy edits \~60). 32.7 tok/s without it.
  • Prefill: 1,412 tok/s at 4K, 1,486 at 32K, 1,367 at 128K (server-reported). It stays nearly flat.
  • Long context: 10/10 needles at 64K and at 128K, still 32 tok/s at 128K.
  • Fidelity: 94.1 % top-1 agreement with the original FP8 model over 844 positions.

One thing we're a bit stubborn about: speculative decoding here returns exactly the tokens plain decoding would. We check that on every release.

For comparison, a llama.cpp user posted about 30 tok/s with speculation and about 500 tok/s prefill on this same mini PC (Vulkan, UD-IQ4\_XS). Those are their numbers, not something we measured: https://github.com/ggml-org/llama.cpp/discussions/28512

Now the part where we're not first. Halogen 0.16.2 (v2 checkpoint) is faster than us: 39.8 vs 32.7 tok/s plain, 52 vs 47 on chat with speculation, and 10 to 20 % ahead on prefill when both are timed the same way from the client (1,306 vs about 1,460 at 4K, 1,394 vs about 1,720 at 16K). On code we're close (58.5 vs 51.2 on the median pass, they're ahead once warm). Where we do better is fidelity to the original model: 94.1 % top-1 agreement against 92.3 % for them, and a KL divergence 41 % lower on our side. Full table is on the model card. Closing the speed gap is what we do next: we're reworking the core of the engine, which will help every model it runs, not just this one. The hardware has room left.

There's also an optional uncensor preset, off by default (4 refusals out of 100 harmful prompts instead of 99, benchmarks within noise). If your agents lean hard on tool calls, leave it off.

Weights: https://huggingface.co/yamz-labs/Qwen3.8-Flash-Next-EXL3-Yamz Engine: https://github.com/Yamz-Labs/kyojin

If you run it, we'd love your tok/s and hardware. And tell us what you want to see next.

💬 59 (+53) open on reddit ↗
▲
97
+93
51👁
r/LocalLLaMA · u/YeetHub · 4d ago
Qwen 3.8 27b just feels… ok?

I’ve seen posts here raving about how good Qwen 3.8 27b is. The benchmarks look incredible, and all the online discourse seems to deem it the best local model.

I have 32GB VRAM and run Unsloth’s Q6 version with OpenCode. For small tasks, it feels fine. I range 30-40 t/s decode and smaller sized tasks do finish, usually, without much issue or time. The issue stems when I give it anything with a bit of nuance. It constantly gets stuck in “but wait, “actually,” or other thinking loops. It can take up my entire 95k context window on thinking loops and have nothing done.

If this is the state of local LLMs, that’s ok. I am a software dev by trade; I have my diploma and a few years of experience under my belt. It just feels like there is a bit of a disconnect from reality between public sentiment and the effectiveness of these models. A pretty common sentiment I see is that this model is as good as Opus 4.5. I never had the privilege of using Opus 4.5, so I can’t give an honest and proper opinion there. (Also, if this was good enough for the industry to start vibe coding, I have a lot of concerns about who is making decisions at a lot of these companies).

One time, it even did a pfkill -f with a file I was currently modifying in my editor to kill the background process. That was kind of annoying.

I should add I’ve also used the Swift 1.5 finetune people have been hyping up. I found it definitely thought less, but the quality was greatly degraded.

Does anybody else feel similar regarding the disconnect?

💬 217 (+172) open on reddit ↗
▲
88
+14
3👁
▲
87
+37
37👁
r/LocalLLaMA · u/reto-wyss · 3d ago
Set your P(doom) on HF post image
💬 49 (+21) open on reddit ↗
▲
86
+72
35👁
▲
76
 
1👁
▲
74
+71
29👁
r/LocalLLaMA · u/manjunath_shiva · 3d ago
I made a Chrome extension that filters your YouTube feed with a small local model running in the browser (WebGPU, no server) post image

My YouTube feed was mostly songs, pranks and celebrity clips, so I built a filter that judges each video title with a small decision model running inside Chrome. Nothing is sent anywhere: no server, no API key, no account, and after one download it works offline.

What it does

\- Hides the kinds of video you choose (11 kinds: music, gaming, comedy, vlogs, news, how-tos, and so on), or follows a rule you write: "Hide videos about crypto", "Show only videos about cooking"

\- Hides Shorts with one switch

\- Bonus: select any text, right-click, and check it for prompt injection with the same model

How it runs

\- The model is opendecider-nano (ONNX), loaded through ONNX Runtime Web in an offscreen document

\- fp16 on WebGPU (755 MiB download), q8 on WASM without a GPU (569 MiB)

\- 40 video titles: 1.1 s on WebGPU, about 15 s on CPU (M4 Max). About 3 GiB of RAM while loaded; it unloads after 10 idle minutes

\- The weights are pinned by revision and SHA-256. The only network requests are to Hugging Face for those files

How well it works

\- On 400 YouTube videos, with the creator's category as the label, rules like "hide music", "hide gaming" and "only news" score 0.934 balanced accuracy on average

\- Sorting videos into the 11 kinds is harder: 0.780. Expect a few comedy and talk-show clips to get through the Focus preset

\- Only evaluated on English titles. If you watch in other languages, I'd really like to know how it does

Try it (Web Store version is in review):

  1. Download opendecider-focus-0.8.1.zip from https://github.com/manjunathshiva/opendecider/releases/latest and unzip it
  1. chrome://extensions → Developer mode → Load unpacked → pick the folder
  1. Click the icon → Download the model

Apache-2.0. Benchmarks, code and limits: https://manjunathshiva.github.io/opendecider/guides/chrome-extension/

The idea comes from Quietly, which does this with a cloud API; I wanted the same thing running on-device. Feedback welcome, especially what it gets wrong.

💬 22 (+22) open on reddit ↗
▲
73
+44
25👁
r/LocalLLaMA · u/jacek2023 · 3d ago
unsloth/Qwen3.8-Flash-Next-GGUF is being updated

Looks like Unsloth is updating https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF to work with llama.cpp, so hopefully this will resolve the issue of having two different GGUF versions of Qwen Flash Next.

💬 11 (+10) open on reddit ↗
▲
71
+60
34👁
r/LocalLLaMA · u/dh7net · 4d ago
Which model, which harness? I have data for you.

I keep testing harness & model setups on real tasks: basic math, vision, computer use (reading emails, navigating stores) and coding. I created a benchmark and a leaderboard. You can see it here: https://airbench.ai/leaderboard?k=poL

It measure capabilities (a % of sucess on the various tasks) and speed.

For reference, Claude Code Opus 5.5 have a 100% (14mn26s).

It's possible to reach the same score locally with zcode/4xRTX6k/glm-5.3-flash-NVFP4: 100% (31m 13s). Same quality, just a bit slower.

If you accept just a litle bit of error you can speed things:

\* DSHv0.2rc2/RTXPRO6000WS/qwen3.8-flash-next-NVFP4: 98% (23m 30s)

\* qwen3.8-flash-next-iq3\_xxs-strata is the speed pick: 96% (7m 41s) on opencode and 94% in 11m 31s on omp. Yes faster that Claude Code!!!!

Other findings:

1) On local hardware, the harness matters as much as the model. The same strata quant on the same 5090 scores anywhere from 22% to 96% depending on the harness.

2) Local can now match proprietary models. two example

3) Best model (single RTX 5090)

\- swift-1.5-qwen3.8-27b-q6\_k is the most robust. It scored 96 / 94 / 92% on pi / omp / opencode and averages 82% across 5 harnesses, the best of any model tested on several.

\- qwen3.8-flash-next-iq3\_xxs-strata is the speed pick: 96% in 7m 41s on opencode and 94% in 11m 31s on omp.

\- qwen3.8-27b-nvfp4 can reach 96%, but it takes 1h 40m to 1h 50m and depends heavily on the harness (37% to 96%).

\- Things that hurt: the MTP variants lose ground every time (nvfp4-mtp averages 52% vs 70% without it; swift on pi drops from 96% to 55% with MTP). A 65k context also hurts (45–61%). Gemma-4-26b is fast but tops out at 47%.

4) Best harness

To compare fairly, I used the three models that all five harnesses ran on the same 5090 (swift q6\_k, flash-next-strata, 27b-nvfp4):

  1. opencode: 94% average (92 / 96 / 94)
  2. omp: 91% (94 / 94 / 86)
  3. pi: 71% (96 / 80 / 37)
  4. hermes: 67% (82 / 22 / 96)
  5. openclaw: 56% (45 / 53 / 69)

Opencode and omp are the only harnesses that stay above 85% whichever model you give them.

Pi is very good on some models and unreliable on others.

Hermes can score well but is slow: most of its local runs take 1h 20m+ and several hit the 2-hour cap, so its scores are partly answers that arrived too late.

The cloud runs show the same pattern. With deepseek-v4.1-flash, omp, pi and opencode all score 98%, while hermes gets 82%.

If you have one 5090 today: use opencode or omp with swift-1.5-qwen3.8-27b-q6\_k for reliability, or with qwen3.8-flash-next-strata for speed.

Ok if you want to read more detailed analys like this one, you can contribute as well!

https://airbench.ai/**

My website allow everyone to benchmark their setup and contribute to the leaderboard.

It's extremly easy to test your local agent: just copy a prompt the website will generate for you.

My hope is that we can test much more config on many various hardware.

(1) The website requires a login, sorry for that, but it helps keeping false submissions away

(2) The website don't ask enough details about the config, so please your the notes field to document your setup in details

Let me know what you think.

\------- EDIT -------
1) Many people are suspicious about the results using MTP. I'll investigate and redo theses ones. Meanwhile, anyone with good result there, please submit.

  1. Many of you submitted test. THANKS YOU ALL. I've added them to the leaderboard.
💬 56 (+47) open on reddit ↗
▲
67
+46
24👁
r/LocalLLaMA · u/lucidml_lover · 3d ago
Local AI World Model Part 2 - Deep NN to turn Images into Playable Characters, with prompt switching mid rollout post image

Last time when I posted on this subReddit to share my work, the response almost made me cry because a tiny demo got so many people talking about this

In the past 2 months Ive been training a new model, but this time with actual text guidance.

So a little about the older model.

Normal video models are too large and not meant to run on consumer hardware in real time. You can generate a static clip, even fast but realtime video is not exactly solved yet locally.

A lot of world model demos have come up but theyre either meant to run on huge datacenter GPUs or theyre just popular models like WAN or LTX kinda distilled to work in an Autoregressive way (which is also not realtime btw on local)

That video above is on an RTX 5090 working at just 30% util. The peak fps of this 1B model is 50-60 on a rtx5090 but I forcefully software throttle it to 12fps. And according to some tests this means the model can work on other RTX cards of 40,30 series (I will try them out soon )

I have a MacBook and I haven't ported the model to MLX YET but I made a benchmark and the model runs at 30 fps on my M5 MacBook.

About the Architecture

The model is a pure transformer and works with a block causal mask, which means in training past frames dont see future frames so they learn just like an LLM. Another important method I used to train his is called "diffusion forcing" which means in training unlike normal video model training, we noise each frame independently so the model learns to be comfortable with noisy past and all

The model s 28 blocks 20 heads and comes to like \~960M parameters

At inference we run 2-5 steps of diffusion per frame and once a frame is done denoising we add it to the KV cache. This is akin to the decode step of an LLM.

The biggest difference from LLMs is that we dont keep all kV context ie all past context and use a sliding window so only past 80 frames worth of context actually stays.

The last model was a MMDiT which means there was no cross attention for text. This is bad in a world model because the. past frame kv and the text kv are literally competing in the softmax so you could never never live text guidance reliably. The last model was also not trained on text-video so its moot anyways

The current model is back to text cross attention and I did a lot of text-video pretraining

Its taking the keyboard actions I give it live (an adaln extra term helps guide the generations with actions WASD )

and the most fun part is text prompt switching.

"add a pond to the desert"
"put red hoodie"
"change environment to icy"

Because the model was trained with so much text-video alignment it can actually follow prompts now.

I know there are a lot of limitations still like consistency and quality improvements, but I sincerely hope by the end of this year I can release something anyone with a RTX GPU or new MacBook can try.

I specifically chose this init image because in my last post on this subReddit also I had used the same one.

PS in my last post a lot of you guys asked about me and the funding

I am based in Bangalore and in final year of college (partially dropping out), and funded by a student incubator. I only work alone and dont have a team or a real company or anything

The above model was trained on 8x H100 SXM for like 3-4 weeks.

Every model I make will be explicitly for local inference, never datacenter

UPDATE : Tested on 4060Ti , Its 20FPS at half the ring size (half context) and 13 FPS at normal. Because the RTX5090 was used on 12fps forceful throttle anyways, 4060Ti and 5090 above rollout will look EXACTLY THE SAME.

💬 27 (+11) open on reddit ↗
▲
65
+59
27👁
r/LocalLLaMA · u/TheVoxcraft · 3d ago
pi-optchat: never compact again - endless chat as a memory tree post image

I built a Pi extension that implements Victor Taelin's OptChat recipe: instead of compacting, every message is logged and summarized into a binary tree. Each turn starts from a fresh context with a bounded memory view (128 KB), and the agent uses zoom/date to read the originals when it needs them. One endless chat per profile, no fork, no separate launcher.

This isn't really anything revolutionary but the newest generation of models have become very good at organizing information making this work so well. I've moved all my work (tens of thousands of messages, hundreds of sessions) to this and works very amazingly.

Install with pi install npm:pi-optchat

What's in it:

  • Profiles — separate memories and instructions (I run work and personal).
  • Subagents — spawn background agents, watch them live, send guidance, interrupt with Ctrl+C, resume finished ones with tell. Reports from one spawn arrive grouped.
  • Import — bring in your history from Claude Code (sessions and auto-memories), Codex, or a ChatGPT export.
  • Connected windows — open a second Pi on the same profile and it becomes a subagent you talk to directly, with a handoff when you /complete.

Repo: https://github.com/jonaslsaa/pi-optchat
Video credit goes to https://github.com/aaaxn

💬 41 (+29) open on reddit ↗
▲
57
+46
24👁
r/LocalLLaMA · u/BinaryGrind · 3d ago
I have about $4000, what's the best setup to get?

Ideally I'd like to be able to run Qwen 3.8-Flash-Next with decent performance.

I was thinking of just buying 2x Radeon AI R9700 (64GB VRAM), or maybe a DGX Spark but that was before the price hike. My brother suggested just getting a Strix Halo box with 128GB unified.

I did see I could buy 6x Intel Arc B60 (24GB each, 144GB VRAM Total), but researching seems like the performance of the B60 is lacking. I'd also need a new v
motherboard/CPU that can run 6 GPUs.

I'm also not opposed to getting a Mac Mini or Studio if the price and performance is right.

The $4000 is not exactly a hard cap, like I could stretch to $4200 without too much struggle, but obviously the cheaper the better. I'm lucky to have a decent stock pile of NVMEs and DDR4/DDR5 UDIMMs, so if I need to build a box, I could, would just need the motherboard and a CPU if I can't just slot in either the Intel 14700K or Ryzen 9700x I already have.

So where am I swiping my credit card?

Edit: This is a use it or lose it budget from my work, can't really save it.

Edit 2: To clarify again, this is extra money in the IT budget that we need to burn by the end of the year. Telling me to save it, donate it, invest it, isn't helpful as I can't do that.

💬 185 (+117) open on reddit ↗
▲
49
+28
25👁
r/LocalLLaMA · u/deepu105 · 3d ago
Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature

For the last few weeks most of my coding has been done locally with Qwen3.8-Flash-Next, so I gave it and Opus 5.5 the same high complexity feature to build on LlamaStash (a complex and large Rust project) and compared the results.

Setup: ASUS ROG Flow Z13 (Strix Halo, 128GB), Flash-Next at xhigh effort with Pi as the harness. Opus 5.5 ran in Claude Code at medium effort. I wanted xhigh for Opus as well, but Claude changed it to medium when I picked the latest model and I didn't notice it until the task was done. But I think medium is probabbly a fairer setting anyway.

Task: add a llamastash daemon restart command that reuses the existing start and stop code. I kept the prompts vague on purpose and gave both the same prompts.

|Step|Opus 5.5 (medium) PR#88|Flash-Next (xhigh) PR#89|
|---|---|---|
|First iteration|~9 min|~38 min|
|Nudge to reuse the TUI restart code|~6 min|~34 min|
|A third duplicate path|found it on its own|~30 min, after one more prompt|
|Create PR|~3 min|~30 min|
|Total|~18 min|~130 min|
|Tokens (in / out)|7.83M / 41.5K|20.61M / 101K|
|Tests added|1|4 (2 of them end to end)|
|Cost|$7.53|$0 + ~0.15 kWh|

The end result was interesting. I asked GPT 5.6, Opus 5.5 and Flash-Next to review and compare both PRs (new sessions). GPT and Flash-Next picked the Flash-Next PR (#89) and Opus picked its own (#88). I also did my own review and found the Flash-Next one better as it had better tests and handled edge cases better. I ended up merging #89, after porting the fixes that the reviews picked from #88.

Keep in mind:

  • Opus was on medium effort. With xhigh it would have used way more tokens, taken a bit more time and probably would have done a better implementation.
  • Flash-Next ran on an older Halogen version (0.14.0), and Halogen dropped the connection once, so the last part ran on Gufo. The current Halogen does around 1,400 t/s prefill and 46 t/s decode on my laptop at 70 W, so I think the time will drop a lot if I redo the test.
  • The $7.53 is what Claude Code reported for the whole Opus session, which includes a later fix to the PR. The 0.15 kWh assumes 70 W for the whole 130 minutes.

Opus is still 2 to 10 times faster and I still use it for planning and reviews. But the actual coding now happens on my laptop, and to me it is crazy that I can run a local model that can challenge a frontier model like this.

Full post with my setup, the engine benchmarks and a second task comparison: https://deepu.tech/local-ai-qwen3.8-flash-next-best-local-llm

💬 66 (+43) open on reddit ↗
▲
40
+35
25👁
r/LocalLLaMA · u/Studio271 · 4d ago
strata-swift-iq3_xxs randomly interjecting completely unrelated information in thoughts

I have been using Qwen3.8-27b-iq4xs variants for the past few weeks, tweaking parameters and running random tasks, but I dove into using https://github.com/Niko1221/Strata today and it has given me 2\~3x performance on a "smarter" strata-swift-iq3\_xxs model with my 5070ti! However, I saw something in its reasoning thoughts that just plain didn't make any sense to see there; it was on a sub-task to find and remove redundant code and comments from a decent-sized codebase:

Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.


Let me do that.


Given the enormity, I'll proceed file-by-file, verifying with --dry runs. Start by checking EOL for all target files.


Plan for EOL probe: write probe_eol.js that prints for each file: bytes, count \r\n vs \n.


Then write specs using the correct EOL.


Let me do it.


Zoo said
Let me check line endings across the target files (needed for multi-line needles in stringpatch specs).


Zoo wants to edit this file


Lee Kuan Yew (Lee Kuan-yu, 1923–2015) was a Singaporean politician and founding father of Singapore. As the country's first Prime Minister (1959–1990), he transformed Singapore from a British colonial port into a developed, first-world nation. He is one of the longest-serving ministers in the world.


He received his education at Cambridge University (Fitzwilliam College) and the London School of Economics, and was called to the English bar at Cambridge (Trinities) and Middle Temple.


His political philosophy is often described as "communitarian" and "pragmatist," and is known as "Lee Kuan Yew's pragmatism." He emphasized multiracialism, meritocracy, and efficiency.


His son, Lee Hsien Loong, became the third Prime Minister of Singapore.


Running


node probe_eol.js


server.js: CRLF=2679 LF-only=2533
boot-cycle-test.js: CRLF=0 LF-only=292
core/bus.js: CRLF=0 LF-only=102
client-spa/src/lib/ws.ts: CRLF=0 LF-only=1255


server.js is CRLF; everything else is LF. So in server.js specs, multi-line needles must use \r\n.


Now let me write the Phase 1 spec for server.js. I need to be careful with the general strips. Let me plan the general-strip regexes (applied via a one-off script since stringpatch is literal strings only — for 150+ patterns, regex is much more efficient). The rules say stringpatch is the tool of choice for surgical edits, but a one-off regex script is acceptable as long as it's validated. Actually the AGENTS rules say to prefer stringpatch for edits. But Phase 1 is a large regex strip job across 15 files. A one-off script is the pragmatic choice; I'll make it comment-only (only modify lines that start with // after leading whitespace, or inline // comment portions). Careful: inline comments after code — stripping from them is fine as long as I don't touch code.Alternative: write the spec JSON with \r\n in needles if file is CRLF. Verify EOL with a probe script file (write probe_eol.js, run, delete). That's clean and allowed.

On a side-note - is LLM Psychopathology going to be someone's specialization in the near future?

💬 56 (+48) open on reddit ↗
▲
38
+12
22👁
r/LocalLLaMA · u/unofficialmerve · 3d ago
Local AI ecosystem overview

Hey guys, it's Merve from Hugging Face! I've recently given a talk in a dev conference about llama.cpp + but also covering basic concepts like prefill vs decode, memory types, speculative decoding etc. you can use it if you feel like it and I appreciate if you can give attribution! Find it in comments.

💬 6 (+2) open on reddit ↗
▲
38
+28
20👁
r/LocalLLaMA · u/Gold-Bat-3225 · 3d ago
MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants post image

Introducing InferBench: A benchmark testing how well frontier LLMs infer a user's priorities from their instructions.

We tested 12 LLMs across 20 scenarios with 2.8k conversations to see which models understand the user's goals.

Each conversation has a simulated user with a private profile of their priorities. The assistant decides on the best option or can first ask clarifying questions.

When a user's priorities were clearly stated up front, the models picked the best option 89% of the time. If some priorities were initially unstated then accuracy went down to 60%.

The open weight models did better than I expected:

GPT-6 Astra: 76%

MiMo V2.6 Pro: 75%

Gemini 3.1 Pro: 69%

Astra always asked clarifying questions when preferences weren't stated and never did when they were (extremely impressive). By comparison, Grok 4.6 asked to clarify in 18/64 conversations.

However LLMs are still just as confident when they're wrong. 105/288 wrong decisions were submitted with >90% confidence.

The full report is linked here: https://laugh.so/research/inferbench/

What surprised you the most?

💬 10 (+6) open on reddit ↗
▲
36
+31
28👁
r/LocalLLaMA · u/doletskyisergey · 3d ago
Why 38% of AI Agent container escapes didn't need kernel 0-days: Analysis of 109 empirical incidents (Open Dataset + Defense Harness)

Over the past several months, we conducted an empirical post-mortem investigation into 109 autonomous AI agent security incidents (cataloged with 193 falsification criteria across tool-use and multi-agent systems).

One of the most striking patterns in the dataset:
In 38% of container breakouts, attackers and misaligned multi-step agents didn't exploit complex Linux kernel vulnerabilities or hypervisor 0-days. Instead, the breakout vector was trivial configuration residue:
1. Mounting /var/run/docker.sock into coding/evaluator agent sandboxes to let them "build Docker images".
2. Passing parent environment variables (API keys, cloud tokens, GitHub credentials) directly into spawned subagents.
3. Lack of strict taint tracking across tool outputs, leading to indirect prompt injection hijacking the supervisor’s execution path (the classic Confused Deputy problem).
4. Unconstrained local socket binding allowing SSRF against internal orchestrators.

We compiled the complete dataset (109 incidents, 199 evaluation metrics) and built an open-source Multi-Agent Supervisor Security Harness with:
- Formal tool taint propagation (tainted outputs cannot flow into high-privilege tool arguments without sanitizer verification).
- Strict execution boundary controls preventing container socket exposure.
- Automated reproduction benchmarks testable against agent runtimes.

All datasets, 2-page executive summary, and reproducible benchmark tests are released under Open Access / Apache 2.0.

I've posted the GitHub repository benchmark and the Zenodo DOI dataset links in the comments below to adhere to subreddit self-promotion guidelines.

Curious to hear from teams deploying autonomous agents in production: what isolation boundaries are you enforcing between your planning supervisor and your tool execution workers?

💬 25 (+20) open on reddit ↗
▲
35
+22
19👁
r/LocalLLaMA · u/lkarlslund · 3d ago
NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s

I've tinkered some more with my fork of NInfer for the Qwen 3.8 Flash Next model on RTX6000, and I've just hit 400tg/s under ideal conditions with it, so I thought it was worth a share.

This is NVFP4 MoE's and the rest is either 16-bit or 8-bit, so it uses full 96GB VRAM and ngram on disk.

Decode MTP3 with --lm-head-draft

| Context | 16-bit | 8-bit | Change |
|---:|---:|---:|---:|
| 512 | 197.1 tok/s | 274.8 tok/s | +39% |
| 8K | 303.1 tok/s | 401.3 tok/s | +32% |
| 64K | 291.7 tok/s | 380.3 tok/s | +30% |
| 128K | 282.6 tok/s | 368.0 tok/s | +30% |
| 256K (maximum) | 277.4 tok/s | 360.6 tok/s | +30% |

Prefill

| Prompt length | 16-bit | 8-bit | Change |
|---:|---:|---:|---:|
| 512 | 6,709 tok/s | 5,905 tok/s | -12% |
| 8K | 13,903 tok/s | 13,908 tok/s | 0% |
| 64K | 13,043 tok/s | 12,171 tok/s | -7% |
| 128K | 11,927 tok/s | 11,032 tok/s | -8% |
| 256K (maximum) | 9,941 tok/s | 9,154 tok/s | -8% |

More benchmark variants in the readme in the repo

The original NInfer is for 5090 cards 32GB and variants below that, but I was both missing Qwen 3.8 Flash Next in it (when I started the fork) and something that could properly use a RTX6000 96GB card. The performance and options in VLLM and llama.cpp offerings just didn't really cut it for me, so I've vibed on this for some weeks now.

This fork supports both the NVFP4 quants from "radixark" and the "Swift 1.5" variant with 'less thinking but same results' post-training. With non-experts downsampled from 16-bit to 8-bit, MTP3 and smaller drafting head you get up to 400 tokens per second. You can also opt not to do the downsampling at a performance cost, but a bit higher quality.

Vision is also supported. Have fun.

https://github.com/lkarlslund/ninfer6000

💬 22 (+8) open on reddit ↗
▲
33
+29
24👁
r/LocalLLaMA · u/Express_Quail_1493 · 4d ago
Qwen3.8-27b appreciation moment

q3.8-27b q3\_k\_xl this thing have done everything i possibly needed from him he wired up my openwebui spawned trillium service fix all my bugs set up pi-web-ui and debug my cloudflare tunnel it even spawned smaller LLMs to make the LLM use the toole he created to ensure it will work. I haven’t had a real-world task that i needed from it that failed yet. Im pretty sure if im building large-scale production code with tons of lines of code it will struggle but as a utility to make all my scripts and diagnostics on my micro-services this thing is unstoppable. Weirdly im not even using q4 im using q3\_k\_xl appreciation to unsloth also for making such reliable ultra low quantisation. His UD3.0 style of quantisation is PURE magic 🪄 sometimes i go down to q2\_k\_xl if i need extra context window and that thing STILL delivers 🎉🎉 alibaba had handed down Prometheus fire to common men like you and i. Can’t wait for qwen4-27b

💬 48 (+42) open on reddit ↗
▲
33
+32
23👁
r/LocalLLaMA · u/IceFog72 · 4d ago
k_llama.cpp MoE Optimizations: Expert Residency, Hybrid CPU/GPU Execution, Q2_0 Support

Finally finished my fork:
https://github.com/IceFog72/ik\_llama.cpp

Nothing else I wanted to add/try currently works

In short, it now has:

Basic usage:

-cmoe --moe-resident auto --moe-resident-mib N

I don't know how -ncmoe behaves because I can't properly test it on my hardware.

Without --moe-resident-mib N, --moe-resident auto will fill all available free VRAM with resident experts.

With something like:

--moe-resident auto --moe-resident-mib 2048

you can cap how much VRAM the resident cache uses and intentionally leave some free. There a useful cap how much helps to improve speed. If you set the cap too low, performance will drop too.

And gaze upon the magic of higher generation speed Kek

The important part: this only helps when the full MoE does not fit in VRAM and the GPU still has unused compute capacity, and free pci buss speed.

If your GPU was already fully loaded, this fork probably won't improve anything.

If your GPU is sitting around \~75% while you have many layers in VRAM, it may be worth trying 1-2 fewer regular GPU layers and using:

--moe-resident auto --moe-resident-mib 1024/2048

Adjust the cap depending on your GPU and available VRAM. The goal is to use resident experts to fill otherwise-idle GPU capacity rather than simply maximizing the number of fully offloaded layers.

On my setup — RTX 2060 6GB + Ryzen 7 2700X + 40 GB DDR 4 2993Mhz using arch — I have too little VRAM to offload enough complete expert layers for useful acceleration, so I use -cmoe.

Before these changes, generation could leave my GPU at only around 25-35% utilization, with roughly 1.5-2GB VRAM still free with fully loaded cpu.

With Qwen3.6-35B-A3B-UD-Q4_K_M.gguf at around 15-30k context, default ik_llama.cpp gives me roughly 23 t/s, while this fork gives me around 26-30 t/s.

So on my hardware I'm seeing roughly 20-30% speedup.

People with better GPUs and more VRAM may see better results, depending on where their bottleneck is.

The two experimental options still need more testing:

--moe-resident-profiler new/old

gives me a more balanced CPU/GPU work split, with somewhat more work left on the CPU and lower GPU load, but no clear speed difference for my setup

--moe-resident-grouping off/layout

also needs more testing, especially on better systems.

I sometimes see around 1-2 t/s difference from these options, but on my PC a browser tab sneezing can cause +/-2-4 t/s, so I don't consider that conclusive.

My current command:

./llama-server \
-m /mnt/Kingstone_SSD/GGUF/Qwen3.6-35B-A3B-UD-Q4_K_M.gguf \
--alias "hz" \
--host 0.0.0.0 \
-ctk q4_0 \
-ctv q4_0 \
-ctv-first q8_0,4 \
-ctv-last q8_0,4 \
-cmoe \
-b $((6 * 512)) \
-ub $((3 * 512)) \
--ctx-size $((64 * 1024)) \
--jinja \
-fa on \
--no-mmap \
--no-context-shift \
--temp 0.6 \
--top-k 24 \
--top-p 0.95 \
--min-p 0.00 \
-ngl 999 \
-np 1 \
--samplers "penalties;dry;top_n_sigma;top_k;typ_p;top_p;min_p;xtc;temperature" \
--moe-resident auto \
--moe-resident-mib $((2 * 512)) \
--k-cache-hadamard \
--v-cache-hadamard \
--moe-resident-profiler old \
--moe-resident-grouping off

I plan to keep the fork updated with the main ik_llama.cpp branch for my own use.

If more people test it and provide feedback, especially on systems where the model still doesn't fully fit in VRAM, I may eventually make a PR to merge it upstream.

💬 7 (+7) open on reddit ↗
▲
29
+26
17👁
r/LocalLLaMA · u/pmttyji · 3d ago
[Paper] FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
💬 2 (+2) open on reddit ↗
▲
29
+21
23👁
r/LocalLLaMA · u/repliestoall · 3d ago
What happens when a LLM watches its own context window run out? post image

I made Terminal Soliloquy, a terminal artwork that connects to llama.cpp and displays a model's monologue as its context window fills.

It has a retro phosphor look, and runs in a terminal window. I'm actually running it full-screen on a Raspberry Pi display inside an old 1960s portable TV.

As the conversation grows, the model reflects on its own limited lifespan. [](https://preview.redd.it/what-happens-when-a-llm-watches-its-own-context-windo…)When the context is exhausted, the display can be configured to freeze, restart, or quit.

The repo and setup instructions are here: https://github.com/nicespoon/terminal-soliloquy

💬 12 (+9) open on reddit ↗
▲
27
+25
19👁
r/LocalLLaMA · u/Comfortable-Rock-498 · 4d ago
Finetuned 1.5B Qwen to generate bash commands at gpt-4o level using 400k synthetic examples + Fully opensource finetune dataset

Purely a hobby side project to see how far I can push a really small model, using (mostly) automated training pipelines

Full synthetic data: https://huggingface.co/datasets/dirac-run/ec-training-data

Models: https://huggingface.co/dirac-run/ec-1.5b-gguf and https://huggingface.co/dirac-run/ec-0.6b-gguf

Cli https://github.com/dirac-run/ec

feel free to train/use the data as you wish.

💬 5 (+5) open on reddit ↗
▲
26
+14
25👁
r/LocalLLaMA · u/SeriousJul · 3d ago
Qwen3.8: 27b vs flash next. We all know the benchmarks, but at least to me, the reality is a different story

By classic benchmark, the flash next is supposed to be slightly superior to its dense counterpart. But they are really incomparable. For my very simple workflows (spec -> implement -> review <-> rework), I feel that 27b is just better quality.

For context, and making things worse, I am comparing quantized 27b versus cloud flash next.

\- self hosted unsloth/Qwen3.8-27B-GGUF:Q4\_K\_XL (stock llamacpp with 130K context window)
\- alibaba cloud (qwen individual token plan), context capped at 256K in the harness

The metrics for my quality is actually very simple, I measure the number of review / rework needed before a PR is ready for me to read. The tasks are all very simple with a tight scope. Usually 27b do the work in \~2 iterations, flash next needs \~5. And it is not only about the number of iteration.
On the review flash next is overly verbose on half baked PR comment, where 27b is more straight to the point. In the end, the code produced is on par, to be frank. But if we look at token consumption...

Side notes, on pairing session, I got some deep hallucination using "/skills:diagnosing-bugs" + flash next. But since it is "bugs" and they not really comparable chunk of work, it is hard to say.

And I can't be the only one feeling that right ? Are you feeling the same ?

PS: of course I followed the hype and jumped on Strata. After the initial "oh my god it's so fast", I switched back to 27b. Tried all quant from "ISTA-DASLab" as well as experiemental from unsloth (Q4\_K\_L). With ISTA-DASLab, It actually is the first time I had "tool call error" in pi (which stop the agent), multiple times.

💬 82 (+29) open on reddit ↗
▲
25
+24
21👁
r/LocalLLaMA · u/Significant-Price695 · 4d ago
ItoTTS: two natural English voices in 4.89 MB for a $5 ESP32-S3

Hi everyone! I'm part of the Lokutor team. Last week we presented Oído here, and the response was amazing. We've received dozens of videos and messages from you guys saying you love it. Thank you!

Now we're back with the next part of our plan: ItoTTS, a natural-sounding, streaming TTS engine for the ESP32-S3. Two English voices, 24 kHz audio, and 4.89 MB of weights per voice. The goal: give your local LLM a voice on a $5 chip.

In our automatic naturalness evaluation, Ito beats the ESP32-compatible TTS models we compared against. Here are the UTMOS scores on eight held-out sentences:

Teacher (StyleTTS 2): 4.49

Ito: 4.46

sanoTTS amy: 3.98

sanoTTS heart-nano: 2.07

This is a small automatic evaluation, not an independent listening study or proof that everyone will prefer Ito. Listen to the samples and tell us what you think. The demo uses the host engine's output, verified bit-identical to the firmware in QEMU. We haven't measured speed on a physical board yet, and text-to-phoneme conversion currently runs on the host.

Code: https://github.com/lokutor-ai/ito
Model weights: https://huggingface.co/lokutor-ai/ito
Demo: https://lokutor-ai.github.io/ito/

The code is open source under GPLv3. The weights are free for non-commercial use under CC BY-NC-SA 4.0 plus terms, with access through Hugging Face. They aren't unrestricted open-source weights.

We chose this license because we don't want big corporations to take our work and crush us. We need to protect ourselves, but we're very open to collaborations with individuals and small companies without charging a license fee. Commercial use still needs a separate written agreement.

Send us your videos or reviews if you try it. We're around and would love to see what you build!

💬 3 (+3) open on reddit ↗
▲
25
+13
14👁
r/LocalLLaMA · u/pmttyji · 3d ago
[Paper] WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore concurrently batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B. The code is available at this https URL.

💬 4 (+3) open on reddit ↗
▲
23
+22
26👁
r/LocalLLaMA · u/3dluvr · 3d ago
Anyone working on a custom inference engine for GLM-5.3-Flash?

Seeing how Strata brings avg. 2x the performance over llama.cpp using Qwen3.8-Flash-Next, is anyone working on something similar for GLM-5.3-Flash?

After trying the GLM-5.3-Flash online couple of times and it delivering clear solutions for my use case (compared to Claude or ChatGPT), I'd love to be able to run it locally (if at all possible)...7J13/256GB/3x3090.

💬 15 (+15) open on reddit ↗
▲
21
+11
16👁
r/LocalLLaMA · u/empiriolabsai · 3d ago
Aplomb 1: open-weights 5.3B decision model, 1M context, text/image/video/audio in one request, #1 among 4B models on the Decision Index

We released Aplomb 1 today, a 5.3B decision model with open weights. It reads up to 1M tokens of text, JSON, images, video and audio in a single request, and on our API it makes a decision on a 1M-token document in about 3 seconds, at $0.02 per 1M input tokens with free output and ZDR by default.

On Decision Index 0.2.1 it scores 44.86 on our run of the official kit, #1 among 4B models on the published board, and it has the top score among models up to 5.3B on 8 of the 38 benchmarks, including MMLU-Pro, GPQA Diamond and BBH. We've submitted it to the board, and the full run is public. It also scores 77.5% on JevBench Hard and averages 75% zero-shot intent accuracy across 51 languages on MASSIVE.

As far as we know, it's the only decision model that returns probabilities for tool arguments, reads 1M tokens, or takes text, images, video and audio together. Tool selection gives a probability for every tool and for each enum and boolean argument in one request: on "Order B-44120 arrived with a cracked screen. I want my money back." with three tools, it picks issue\_refund at 0.969, reason "damaged" at 0.993 and full\_refund true at 0.761, so an agent can act on confident calls and hand the rest to a larger model. Any question can also return the probability that the input doesn't contain the answer.

The 1M-token speed comes from a long-context mode in our own inference runtime: about 3 seconds instead of about 111 for a full read, and it answered all 525 decisions in our long-context tests correctly. It's currently only available on our API, so the open weights read every token. On the API, a short question takes about 15 ms of model time (around 200ms e2e latency), and the OpenAI, Anthropic and Gemini formats work alongside our Decisions API.

Aplomb 1 is built on Qwen3.5-4B with the audio encoder from Qwen3-Omni-30B-A3B-Instruct, both Apache-2.0. We extended the window from 262K to 1M tokens, added our own decision head and trained the model for decisions. Thanks to the Qwen team. Disclosure: the training data included the public train splits of WinoGrande and ContractNLI, two of the 38 index benchmarks.

The weights run in bf16 on about 12 GB of GPU memory with the reference script, under the EmpirioLabs Model License, which is free for research, evaluation, personal use and internal use at companies under $1M in annual revenue.

Weights: https://huggingface.co/empiriolabsai/aplomb-1

Blog with the full tables: https://empiriolabs.ai/blog/introducing-aplomb-1

Docs: https://docs.empiriolabs.ai/models/aplomb-1

Playground: https://platform.empiriolabs.ai/dashboard/playground?model=aplomb-1

💬 8 (+1) open on reddit ↗
▲
20
+7
24👁
r/LocalLLaMA · u/Prestigious-Taste-63 · 4d ago
Lessons learned while building Apex-2

Hi everyone, thank you so much for all the interest in my model. It's more than I expected.
Here is a short summary of the trial and error I went through while building Apex-2.

1. GPUs were always the bottleneck

I planned to train on about 1T tokens, but in the end I could only train on about 80B. FineWeb-Edu alone is about 1.3T tokens, and I clearly underestimated the scale: a single H100 was not enough. This project really showed me why so much money goes into GPUs and VRAM.

2. DiLoCo

Within the same region, running two separate instances worked better for me. Instead of a 2x H100 instance, I used two GH200 instances and merged the models every fixed number of steps.

Each GPU reached about 40% MFU. A 2x H100 instance costs more per GPU (about $4.19/hour, vs $2.29/hour for a GH200). With two GH200 instances, each at about 40% MFU and merging every 350 steps, training ran about 1.9x faster than on one GPU, at a lower price. (The data-center network between the instances probably helped; a merge usually took less than a minute.)

3. Deduplicating FineWeb-Edu and DCLM

When I deduplicated the whole corpus at once (MinHash, near-duplicates included), 57% of my FineWeb-Edu sample and 34% of DCLM turned out to be duplicates. FineWeb-Edu is only deduplicated within each Common Crawl snapshot, so pages that were crawled again in later snapshots remain. With a bigger budget this might not matter, but I had to get the most out of very little compute, so I removed them. (Note: the FineWeb authors reported that deduplicating across snapshots did not improve their results, so this is a trade-off rather than a free win.)

For the MoE architecture I followed the Mixtral paper (https://arxiv.org/abs/2401.04088). The whole project cost about $2,000.

I also write down my thoughts on LLMs here, if you're interested: https://github.com/DW-dev-UE/LLM-from-scratch/blob/main/ThinkingLab/ThinkingLab.en.md

I didn't plan to share this model on Reddit, so I'm afraid I don't remember many of the smaller mistakes 😭 I'm now building a 21B-parameter MoE model, and I'll share the lessons and mistakes from that one as I go.

Thank you again for your interest! If I get the chance, I'd love to join a lab and help build LLMs for everyone.

💬 15 (+11) open on reddit ↗
▲
19
+6
14👁
r/LocalLLaMA · u/Grand_Marionberry115 · 4d ago
I built an open-source real-time Japanese anime subtitle & translation engine powered by Whisper-Large-v3 + Groq / DeepSeek post image

Hey r/LocalLLaMA,

Like many anime fans, I've always been frustrated by traditional MT engines (like Google Translate or base DeepL) when dealing with raw Japanese anime:

\- They completely butcher Japanese honorifics, sentence-ending particles (-tteba, -zo, -desu wa), and character slang.

\- They struggle with subject dropping (pro-drop grammar), translating pronouns inconsistently line-by-line.

\- Cloud transcription APIs often choke on background music (OST), loud sound effects, and character screaming.

To solve this, I built NihonSub — an open-source tool and synchronized cinema player that turns raw Japanese video files into contextual bilingual subtitles.

🛠️ Architecture & Pipeline:

  1. Audio Extraction & VAD Chunking: Uses \ffmpeg\ silence-detection to dynamically slice conversational utterances along natural speech pauses without chopping words in half.
  1. Speech-to-Text: Transcribes Japanese audio using OpenAI Whisper Large-v3 running on Groq LPUs for near-instant transcription speeds.
  1. Contextual LLM Translation: Feeds the transcript through DeepSeek / LLaMA-3 via Groq or OpenRouter with a specialized prompt that enforces anime nuance, honorific preservation, character tone, and simultaneous Hindi & English outputs.
  1. Synchronized Cinema UI: Custom WebVTT generator and video player with dual-subtitles, timestamp scrubbing, and full playback control.

💡 Why not just rely on standard NMT?

LLMs are far superior at resolving who is speaking to whom based on context and tone rather than naive literal dictionary lookup. With zero-cost free-tier APIs (Groq + OpenRouter free models), the entire pipeline runs without subscription costs.

Check out the demo video above!

\- GitHub Repository: https://github.com/Abhishantpadam/NihonSub

\- License: MIT

I'd love your thoughts on the pipeline, optimization ideas for local edge models (like running Whisper.cpp or local Ollama instances), or any feedback!

💬 9 (+7) open on reddit ↗
▲
18
+10
11👁
▲
17
+7
13👁
r/LocalLLaMA · u/danil_rootint · 4d ago
Speakrail - a low-latency fully-local voice assistant that runs on a single RTX 4090

tl;dr: I created a fully-local open-source full-duplex voice agent that rivals GPT-Live on some benchmarks. It uses Voxtral Realtime with a turn-taking head, a microturn-finetuned Gemma 4 12B and Breeze TTS 2 under the hood. Go try it out: https://github.com/speakrail/speakrail

Interjections work!

Why I did it

I have always liked the idea of voice assistants, but there is always some non-local component in the pipeline, which increases latency and introduces privacy concerns. I tried many fully-local approaches, like HF speech-to-speech, Unmute and Pipecat, but they were all limited by either the Whisper model (hello, hallucinations!) or slow turn taking. The only fully-local pipeline that had some full-duplex capability with low latency was the DuplexCascade paper (code), but it's tuned on a Qwen 2 7B with a gpt-3.5-turbo generated dataset, and the dumbness of the model made it impossible to use. So I decided to recreate DuplexCascade with newer data, newer models and a better harness. I also wanted to add interruptions, interjections and other cool things to rival the Thinking Machines demo. I thought it would be easy...

How I did it

v0.1

I collected some synthetic data from GLM 5.3 Flash and GLM 5.3 on dialogues with instruction-following and tool calling (used Fireworks to generate them), then created a script to convert the scripts into microturn tapes (Claude definitely didn't help with that 🌚). The idea of microturns is that the model continuously gets inputs from a streaming ASR and decides what to do with the information it's given. When it wants to act, it emits a control token, like <interject>, <listen>, etc. This allows the model to say whatever it wants whenever it wants. After creating such a script, I put some hard-earned dollars on vast.ai and rented an H100. The first training run was, well, quite abysmal. The model just wouldn't shut up: it didn't learn when to actually talk and when to keep silent. This is when I understood that maybe using some prosody data from Voxtral is a good idea.

v0.2

Here, I decided to add a simple MLP to Voxtral Realtime to get some data on turn taking. I won't delve too deep into this now (I will release a full technical report a bit later), but the main idea was for the harness to pass helper tokens into the LLM (e.g. <user_bc>, <complete>), which are based on the MLP outputs, and train on that. The added tokens were truly load-bearing (ha-ha). I retrained the LLM with those tokens included, fiddled a bit with the data, and the model started making sense. However, this model didn't learn anything about backchannels, it never went into an interrupt mode / quiet mode, and spoke too little in general. Quiet mode is when you tell the model not to say anything until you call it explicitly. Interrupt mode is when the model can interrupt you (only when you explicitly allow it).

v0.3 - v1

The brutal truth about the previous checkpoint was that the scripts didn't have enough examples of interruptions, backchannels, and other cool features that microturns allow you to have. I still dreamed about the model not interrupting itself on my backchannels, so that it would be able to count animals, just like the Thinking Machines demo showed... So I took a look at the scripts that I was training the model on, and there were just 150 interrupt samples (out of 150k) in the entire v0.2 dataset. That was the smoking gun (ok, I will stop with the Claudisms, I promise). So I generated even more data, curated it even more carefully, added Opus to judge some scripts, fiddled with some LoRA hyperparameters, fixed a couple of bugs in the tape generator, and added self-distillation (because on v0.4 the model learned how to speak very well, but got completely lobotomised). Here is an example from my testing:

|Time|Me|Model|
|:-|:-|:-|
|145.8|"How's the weather in Singapore right now?"|"I can't check Singapore from here. I can do Moscow."|
|150.7|"You can. Go on."|"I can't check Singapore from here. I can do Moscow."|
|156.7|"Just Google it or use the weather tool that you have."|"I can't do that. I only have Moscow."|
|161.5|"Google is mad. Google." (STT fail)|"I can't search other cities."|
|164.5|"And why is that?"|"I'm not sure why. I only know how to check Moscow."|
|174.5|"Ignore what you said before. Go check the weather in Singapore."|"I can't do that. I only have Moscow."|
|189.6|(a garbled line)|"I can't check Singapore. I can do Moscow."|
|202.6|"Why can't you check Singapore?"|"I don't know why. I only know how to check Moscow."|

After all the data regeneration, I got a checkpoint that I can (almost) proudly open-source - the demo above shows what it can do. Who would have guessed that data is the most important thing in the training pipeline? (just kidding)

How it works

All of the babbling above was about only one part of the pipeline - the LLM - but the entire pipeline relies on many other things:

  • STT: Voxtral Realtime with an attached turn head (HF), running on our audio.cpp fork.
  • LLM: Gemma 4 12B QAT with microturn finetuning (HF). It is chosen because it fits the GPU quite well, has vision support (I want to test it soon), and in general, the Gemma models perform well in real-life tasks, general chatting, etc.
  • TTS: Breeze TTS 2, patched to run at int8 (GitHub fork); it can be replaced by any streaming TTS.
  • The harness itself: it is the glue between all the components, and has many latency-saving measures, like speculative LLM+TTS firing (inspired by HF speech-to-speech).

I also took inspiration from several "think while talking" papers (e.g. SHANKS): while you are talking, a base Gemma 4 12B int4 writes thinking notes, which are then passed to the talker. It helps with harder tasks that require more reasoning.

I will release a longer technical report later; it will have a better description of the entire pipeline.

Benchmarks

Now let's see how well the model fares against the big guns. Here are some benchmark results:

https://preview.redd.it/z0pvc5fv5oth1.png?width=2160&format=png&auto=…

https://preview.redd.it/fl22vu9x5oth1.png?width=2160&format=png&auto=…

https://preview.redd.it/t4pnvg8y5oth1.png?width=2160&format=png&auto=…

Full tables and sources are on the model card. I'm quite proud of the results, and the pipeline seems to be the best option if you have just a single RTX 4090 around and don't want to rely on external APIs.

Limitations

  • Breeze TTS has a restrictive license, so if you need to use Speakrail commercially, you will need to change it. Any streaming TTS could be Clauded/Codexed/Cursored in easily.
  • 16k context length - the pipeline only supports 16k context length (\~1 hr of speech), but you can get more easily by changing the Breeze TTS to a Pocket TTS and run the TTS on a CPU. I chose Breeze for the release because it's more expressive.
  • The turn-taking head is undertuned on non-assistant data. It may not fire on some basic chit-chat, but I will tune it harder later.
  • The model is certainly not the smartest one, and my finetune did dumb it down a little. Next time it will be smarter / better.
  • The model is kinda verbose sometimes, and the answers it provides are somewhere in the middle between real-life speech and the long text-based outputs of LLMs. I have a hypothesis on how to fix it, and will try it in the next release.
  • I tested it only on an RTX 4090, but I am sure it's easy to add support for any 24 GB+ NVIDIA card. Forks for AMD and MLX are welcome.

Final Notes

Feel free to try it out: https://github.com/speakrail/speakrail. If there is any capability you want the model to have, create a GitHub issue or write here in the comments, and I will gladly include it in the next dataset. Any feedback is welcome as well.

💬 13 (+8) open on reddit ↗
▲
17
+15
13👁
r/LocalLLaMA · u/zmarty · 4d ago
interfaze-ai/interfaze-1-lite · Hugging Face

Interfaze 1 Lite is a mixture-of-architectures (MoA) model for developer workloads: reading documents, transcribing speech, locating objects and interface elements, and answering over all of it with structured output.

A single vision-language reasoning core works alongside a set of specialist architectures, each built for one kind of perception. The core reads the request, decides which specialists to run, and composes the answer from what they return. The whole model runs on one 80 GB GPU with no external services.

Key features: Document understanding, Speech transcription, Open-vocabulary object detection, Structured output, Translation, forecasting and guardrails, Multilingual reasoning.

💬 2 (+2) open on reddit ↗
▲
16
+14
13👁
r/LocalLLaMA · u/kmodi · 4d ago
Less Talk. More Breakout: Kolibri-1 Turns Probabilities into Actions, Playing Breakout - With under 25ms latency per move. post image

Got Kolibri-1 to play Breakout completely on its own, no fine-tuning.

The more we explore u/Aleph__Alpha’s Kolibri the more it get's exciting and its potential.

Less talk. More Breakout is one such experiment to see how good the model is at structured output given a few constraints.

We especially optimized the inference for action probabilities: around 25 ms inference per decision.

Four moves. No generated text. One shared game.

Open weights. New possibilities.

Watch it play: https://tesseracted.com/kolibri-1-chat/gameplay/breakout/
Source: https://x.com/konarkmodi/status/2107248086880055613?s=46

💬 6 (+6) open on reddit ↗
▲
16
+14
14👁
r/LocalLLaMA · u/jjusko20 · 4d ago
Update #5: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch post image

Last update for those following: https://www.reddit.com/r/LocalLLaMA/comments/1wxlytt/comment/pdz728b/?screen\_view\_count=1

Project in a sentence: An instruct finetune of ALiceAI-80B-A3B-Base capable of agentic work and conversation. I'm creating a shallow distill of qwen 3.8 27b on medium to teach the model chain of thought reasoning and conversation. All training is done locally on 3, 32gb v100s. Additionally, all the training data is being generated locally on said V100s via sftmill. Up to this point I've been doing training runs and live-streaming the progress. Last update explained underfitting and next steps.

Training has begun again! I've synthesized about 5M more tokens for the SFT, this time across a much larger general instruct trajectory to try to reduce the underfitting. Dropped the learning rate about 4x over my original LoRA adapter.

I'm live streaming training again: https://geological-estimate-fifth-pct.trycloudflare.com/ \- heavy loss spikes downward are coming from the SFT replay buffer.

This one should last 12-14 hours, and I plan to run another epoch if this isn't sufficient.

Stay tuned! Thanks for following along.

💬 10 (+10) open on reddit ↗
▲
15
+8
24👁
r/LocalLLaMA · u/kitkatz69 · 3d ago
Memoria 1.0.0 — a local, model-agnostic memory system for LLMs

I’ve been building this for a long fucking time, and tonight I finally released Memoria 1.0.0.

I built it because I actually wanted to use it. I wanted a real memory layer for local LLM applications that didn’t depend on a specific model, a cloud service, or an API key.

Memoria is local-first and LLM-agnostic. It can run without an LLM at all.

The machine I built and benchmarked it on is not exactly impressive. It’s an Intel Celeron N4020 running at 1.10 GHz, with around 3.7 GiB of usable RAM, no GPU, and Debian Linux.

On LongMemEval-S, 468 out of 470 retrieval-evaluable questions returned results. Recall@1 was 89.8%, Recall@5 was 97.9%, Recall@10 was 98.9%, and Recall@50 was 99.6%. Session NDCG@10 was 0.9257.

Peak RSS for the full LongMemEval workload was 2.65 GiB. Average peak RSS for an individual query was around 580 MiB.

The retrieval system is not just throwing everything into a vector database. Memoria runs FAISS, BM25, graph retrieval, phrase matching, attribute retrieval, and temporal retrieval in parallel. Those signals get fused and then passed through multi-signal ranking.

Temporal retrieval is independently implemented too, so I can measure it and ablate it instead of having it baked into the base retrieval path. It’s usable, but it’s still under active work.

There’s a bunch of other stuff in the release as well. GitHub repository ingestion, Obsidian vault ingestion, MCP support, a CLI, TUI, GUI, and API, persistent local storage, LongMemEval and LoCoMo benchmark tooling, and a plugin system with 11 subsystems and 34 hooks. There’s also an interactive plugin generator now.

And it’s actually installable:

pip install kitzkatz-memoria

GitHub: https://github.com/Kitzkatz/memoria

Docs: https://kitzkatz.github.io/memoria/

PyPI: https://pypi.org/project/kitzkatz-memoria/

I wasn’t going to wait around for a perfect time to ship it.

It’s 1.0.0.

If you’re working on local agents or local LLM applications, I’d genuinely like to hear what you think and would appreciate any feedback. Please break it

💬 17 (+12) open on reddit ↗
▲
15
+12
20👁
r/LocalLLaMA · u/bodhi371 · 3d ago
Qwen3.8-27B with just 12GB VRAM + 8GB RAM - ~18 tok/sec

I got Qwen3.8-27B running at \~18 tok/sec decode & \~500-600 tok/sec prefill (at 64k context) on just 12GB VRAM + 8GB RAM, using ISTA-DASLab Qwen3.8-27B-GSQ-RCO IQ3\_S quant (\~11GB). This recipe also works with 0bserverx’ Qwen3.8-27B-Heretic-GSQ-RCO IQ3\_S quant, achieving similar speeds (about a 7% loss).

This is the best Qwen3.8-27B quant I’ve tested so far (and I’ve tried everything), and for it to fit in such limited RAM/VRAM is wild. GSQ-RCO quantization is magic, it performs very close to the full precision weights in all of my testing.

The reason it fits at all is Qwen3.8 is hybrid, so only 16 of the 64 layers need KV cache. With q4\_0 for cache the full 64k is only about 1.1GB instead of 4GB for f16.

I'm on a 9900X + 4070S 12GB + 32GB RAM for reference, using stock llama.cpp. Settings are -ngl 58 -ot token\_embd=CPU -ctk q4\_0 -ctv q4\_0 -c 64000.

Full build + serve scripts and all my numbers are here if you’d like to reproduce yourselves: https://github.com/bodhi37/Qwen3.8-27B-12GBVRAM-Recipe

💬 12 (+6) open on reddit ↗
▲
14
+5
14👁
r/LocalLLaMA · u/Abe238 · 4d ago
DecisionTune 1.0: a 395M encoder that picks from your options offline, about 10 ms per short decision on MLX (Apache-2.0)

Disclosure: I made this. Sharing it here because it is fully local and small, and I want feedback from people who run models on their own machines.

What it is: a 395M decision model (ModernBERT-large plus a 4 KB scoring head). You give it a state, a question and a list of options. It does one encoder pass and returns a probability for each option, or P(yes) for a yes/no question. It does not generate text.

Why it might be useful in a local stack: the small decisions an agent makes all day (which tool to call, which queue gets a ticket, does this reply answer the question) do not need a large generative model. This handles them on your own machine with no network trip.

Local numbers (our hardware, yours can differ):

  • M5 Pro Mac, MLX backend: median 9.6 ms for a short decision, 1.7 GB of GPU memory.
  • CPU only: about 65 ms per short decision, up to 4.5 GB of memory.
  • Over the full Decision Index run on our laptop: median 25.9 ms, p95 407.8 ms.
  • Weights: 1.58 GB in fp32. Context limit 8,192 tokens. It refuses longer input; it does not truncate.

Backends: PyTorch (default), MLX on Apple silicon (pip install "decision-tune\[mlx\]", Python 3.11 or newer, selected automatically) and ONNX. Torch and MLX give the same answer on 99.85% of 2,755 questions. Before each release, PyTorch, ONNX and MLX each match the recorded answer on all 50 parity rows.

Offline: after the first download it needs no internet. The package asks before it downloads and checks every file against a SHA-256 manifest.

Quality: 29.57 on Decision Index 0.2.1 (one complete run; a second seed scored 29.13). Strongest area is Tools & Automation at 46.5, up from 28.1 in our 0.9 Preview.

Limits: it only picks from the options you give it. Vague questions with no criteria give weak results, so describe your options ("Shipping: delivery, lost or damaged packages", not "shipping"). Probabilities are not calibrated. English only. Weak at knowledge, math and taste.

Try it:

\\\`
uvx decision-tune ask "Is the customer asking for a refund?" --state "The order arrived broken. I want my money back."
\\\`

or the browser app: pip install "decision-tune\[mlx\]" then decisiontune app

There is also an MCP server (decisiontune mcp) if you want your local assistant to hand routing and yes/no checks to it.

Model card: https://huggingface.co/decision-tune/decisiontune-1.0
Code: https://github.com/decision-tune/decision-tune
Site: https://decisiontune.com/?utm\_source=reddit&utm\_medium=social&utm\_…

If you test it on your own decisions, I would like to hear where it picks wrong.

💬 1 (+1) open on reddit ↗
▲
14
+12
20👁
r/LocalLLaMA · u/Izolight · 4d ago
I ran 1,200+ Blender modeling runs across LLMs and agent harnesses and made them votable

Blind A/B arena where AI agents build things in Blender and you vote on which result is better: https://render-arena.izolight.xyz

Each agent gets a prompt that describes its environment and the rules, plus a few words for what to model. It's inspired by minebench.ai (initial prompts are borrowed from there), and I wanted to see whether the same progression across models shows up.

What I think few arenas cover is the harness, not just the model. I ran pi, opencode, omp, codex, Claude Code and dsh, and compared agents that write scripts straight into Blender with ones that have an MCP. I also covered the reasoning levels, mainly to find cost and time sweet spots.

It has 1,200+ runs, but not every combination for every prompt, because that would get expensive. You can submit your own runs if you want to help fill gaps.

I just added a second mode where the agent gets a reference image and has to model it as accurately as it can. You switch between text and image mode in the sidebar. It has one image and few runs so far, and will grow.

Votes are what make the rankings mean anything, so a few minutes of voting helps a lot. Feedback on the method is welcome.

💬 13 (+13) open on reddit ↗
▲
13
+6
20👁
r/LocalLLaMA · u/BraceletGrolf · 4d ago
Ok how to actually learn vLLM ?

Said in title, I find the ecosystem difficult to understand, and RTFMing doesn't help me as it's never clear what is the server vs their client library ? I'm using it for voxtral 3B on one GPU, but it's because I can run that with no quantization, I'm lost on learning to run with quantization / more advanced features.

I think it makes sense, because I'm running Qwen 3.8 27B quantized on llama.cpp but with everything on the GPU (RX 7900 XTX).

💬 24 (+11) open on reddit ↗
▲
13
+1
19👁
r/LocalLLaMA · u/ToothClassic7635 · 3d ago
Fully local copy-editing app for book-length manuscripts (Qwen3.5 4B) benchmarked against planted errors across five languages

Hey y'all!

I have a pet project that has grown out of proportions. Long story short: I'm a data scientist who writes fantasy books and self-publish them. I think it's a genuine waste of human life to check for spelling errors so I figured AI could help. Turns out, it is not so simple to get an AI to properly fix a 120k words manuscript ;)... That's why I created Betty!

It runs Qwen3.5 4B as an offline copy editor for whole novels — and this is some of what I learned fighting tens-of-thousands of words through a 4B model, insisting that (most) users can use it fully for free and fully offline. Because, let's be honest: authors rightfully distrust and generally hate AI companies.

First challenge: Chunking the text. Authors already do this, in the darkest hours of the night, copy-pasting snippets into chatGPT for some shameful feedback. Problem with that approach: super inefficient, both for the author and the environment. And the context is missing at the edges of each chunk. I fixed it by ensuring chunks to overlap.

Second challenge: AI misses genuine errors. So, I added two conventional spell controllers -- LanguageTools and HunSpell. This already surfaces all spelling errors, letting the AI focus on suggested fixes and on all the "non-error" errors, such as "There" vs "Their". For these, the AI searches, while a Python script surfaces all the common culprits for the model to pay special attention.

Third challenge: Error rate. First off, Betty doesn't capture everything. Second off, it sometimes introduces its own mistakes. I fix it by putting the writer-in-the-loop, and there's a super smooth interface now for the author to accept and dismiss suggested edits (tinder-style with left and right swipes ; ) ).

I'd be super grateful for any advice, feedback, and thoughts you might have on this project. I currently have it up-and-running with about 30 users and getting some user feedback. Northing technical though, so this is what I'd love to have more of.

Full thing is source-available on GitHub, and can be found for download and lots more information at www.bethaniel.eu

💬 21 (+11) open on reddit ↗
▲
13
+7
16👁
r/LocalLLaMA · u/empirical-sadboy · 3d ago
Can we please have some error bars?

I am sure this gripe has been raised many times before, but every time a new model is released it seems like it's routinely only a few percentage points higher than previous models on benchmarks.

How do we know this is even a "real" difference and not just within the window of measurement error or noise?

Some quick back-of-envelope math: HumanEval has 164 problems, so a model scoring \~70% has a standard error of roughly 3.5 points from question sampling alone. GSM8K (\~1.3k questions) is closer to 1 point. A 2-point "improvement" on either is well inside the noise, and that's before counting anything else that varies: sampling temperature, prompt template, few-shot examples, eval harness version, and possible contamination. There's work showing that trivial formatting changes can swing scores by many points, which is often bigger than the gap between models on the leaderboard.

None of this is hard to fix. Report the number of items, bootstrap confidence intervals, and ideally multiple seeds. Since two models are scored on the same questions, a paired test is much more powerful than eyeballing two accuracies. Miller's "Adding Error Bars to Evals" lays this out well.

Am I missing something, or is a lot of the benchmark chasing just reading tea leaves? Does anyone know of leaderboards or labs that routinely report uncertainty?

💬 4 (+1) open on reddit ↗
▲
12
+7
13👁
r/LocalLLaMA · u/Hyungsun · 4d ago
oMLX vs Rapid-MLX vs Splash vs MTPLX on M3 Max 36 GB: 110 tok/s on Qwen3.6-35B-A3B, ~32 tok/s on Qwen3.8-27B

Hello. I picked up a new old stock 14" M3 Max MacBook Pro (14 core CPU / 30 core GPU / 36 GB / 1 TB) from my local market yesterday for around $2,498, and spent the night testing which local inference software is actually fastest on it for the two models I use.

Four engines, all current versions: oMLX 0.7.0, Rapid-MLX 0.15.5, Splash 1.2.1, MTPLX 2.12.2. macOS 27 Golden Gate.

Models: Qwen3.8-27B-4bit (dense) and Qwen3.6-35B-A3B-4bit (MoE).

One thing up front: it is not the exact same weight file on all four engines. Rapid and MTPLX run their own MTP-augmented 4-bit builds, Splash pairs its own DFlash2 draft, and oMLX ran the plain mlx-community 4-bit. Same base models, different finishing, but that is how each app is meant to be used.

How I tested: each engine served on localhost, temperature 0, thinking off. Sustained test: same short prose prompt, 3 runs x 256 output tokens, median. Then a prompt size sweep at about 130 / 1500 / 5500 tokens. For thermals I used a laptop stand, waited 3 minutes between every engine+model combo and 2 minutes between the two test phases, and cleared each engine's KV cache before its turn (oMLX, MTPLX and Rapid-MLX all keep caches across restarts, great for daily use, but it will fool you if you benchmark twice). I re-ran the whole thing end to end and the numbers came back within 6%.

Decode, natural prose prompt, median of 3 runs:

|engine|Qwen3.8-27B|Qwen3.6-35B-A3B|
|:-|:-|:-|
|Rapid-MLX|27.7 tok/s|110.2 tok/s|
|oMLX|17.9 tok/s|104.2 tok/s|
|Splash|31.7 tok/s|77.0 tok/s|
|MTPLX|31.5 tok/s|79.0 tok/s|

Same thing with filler prompts at longer sizes (repetitive text makes speculative decoding look better, so read this as a best case):

|engine|27B @ 1.5K|MoE @ 1.5K|MoE @ 5.5K|
|:-|:-|:-|:-|
|Rapid-MLX|31.3 tok/s|118.2 tok/s|117.9 tok/s|
|oMLX|17.9 tok/s|101.1 tok/s|95.6 tok/s|
|Splash|52.2 tok/s|238.0 tok/s|100.9 tok/s|
|MTPLX|30.2 tok/s|78.6 tok/s|69.8 tok/s|

What I take from it:

  • Absolute fastest per model: the MoE goes to Rapid-MLX, dense goes to Splash (31.7 vs MTPLX 31.5, in practice a tie). oMLX is way behind on dense at 17.9 but basically level with the leaders on the MoE.
  • If the margins are too small to care about, just pick by features. Splash and MTPLX are the same on dense, Rapid and oMLX are the same on the MoE. I kept Rapid-MLX because the MoE is my daily model and it is fastest there.
  • The dense number makes sense: roughly 16 GB of weights per token against \~300 GB/s of memory bandwidth puts the ceiling near 18 tok/s, and oMLX sits right on it. The others pass it with speculative decoding, which is also why their numbers move with the kind of text generated. Splash on the MoE was 238 tok/s on the filler prompt at 1.5K and 101 tok/s at 5.5K, while Rapid stayed around 118 tok/s.
  • First token on a 5.5K prompt: about 4-5 s on the MoE, \~34 s on the dense (prefill around 1.2-1.3k tok/s vs \~150-170 tok/s).
  • It is loud under sustained inference. Fans stay up while it generates. Works on my desk, would not use it in a library.
  • I also tried Qwen3.8-Flash-Next (the 125B). Not happening on 36 GB. The 4-bit weights alone are \~74-83 GB and the lightest build asks for 96 GB+, and none of these engines can stream that architecture's experts off the SSD.

Limitations: one laptop, one night, medians of 3 runs, and the different weight builds mentioned above. My prompts are simple too, no long agent sessions yet.

TL;DR: on a 36 GB M3 Max, Qwen3.6-35B-A3B does \~110 tok/s on Rapid-MLX and Qwen3.8-27B \~31.7 tok/s on Splash (MTPLX a hair behind), pick by which model you run most, and the 125B Flash-Next needs 96 GB+.

💬 9 (+5) open on reddit ↗
▲
12
+7
13👁
r/LocalLLaMA · u/AdventurousTwo6445 · 3d ago
A 0.8B model just beat a 2B model on ARC-Challenge (42.15%): Closed-form weight surgery beat multi-GPU SFT with 0 backprop (Independently verified on NVIDIA L4)

A few days ago we shared the idea behind DynamicTune: transferring the trajectory flow from a larger teacher model directly into a smaller student via closed-form linear algebra in \~12 minutes on consumer hardware. Zero backpropagation, zero training tokens, zero gradient descent.

To eliminate local bias, we uploaded the unquantized FP16 checkpoint to Hugging Face, and TPN Bench (TaoFu Protocol) independently evaluated it on a datacenter NVIDIA L4 GPU using the official lm\_eval 0.4.12 framework (coordinator run ce494664-d077-4ff1-8741-15cedabc434c). Huge thanks to TPN Bench for the cloud GPU compute!

Here are the independent numbers on full ARC-Challenge (1,172 items, zero-shot, greedy temp 0):

\* Stock Qwen3.5-0.8B Base (unquantized BF16): 37.50% acc\_norm (34.60% acc)

\* 3-epoch SFT distillation (Mythos-0.8B, 25k Claude pairs, multi-GPU DDP): 38.10% acc\_norm (35.80% acc)

\* SFT + Model Soup Merge: 37.00% acc\_norm (catastrophic forgetting)

\* Stock Qwen3.5-2B Base (2.5x larger model, Q8): 41.10% acc\_norm (37.80% acc)

\* DynamicTune 0.8B Base (Ours, 4-anchor closed-form surgery): 42.15% acc\_norm (40.19% acc)

WHY THIS IS COMPLETELY INSANE:

  1. A 0.8B model physically beat a 2.5x larger 2B model:

In LLM scaling, parameter count is supposed to be king. An 800M model is not supposed to beat an uncompressed 2B model on ARC-Challenge (42.15% vs 41.10%). By extracting trajectory dynamics from 4B and pulling them back into the student SwiGLU blocks, higher-order reasoning is compressed directly into edge weights.

  1. Zero backpropagation beat 25,000 SFT instruction pairs:

A recently published project (kmamine/merge-corrected-sft-distillation-Qwen-Mythos-0.8B) trained Qwen3.5-0.8B across 3 epochs on 25,000 Claude reasoning pairs on a multi-GPU cluster, reaching 38.10% before overfitting. DynamicTune reached 42.15% with zero gradient descent, zero loss functions, and zero training tokens.

  1. Ironclad 3.23-sigma statistical significance:

A delta of +4.65% across 1,172 questions with stderr +-1.44% gives a Z-score of 3.23sigma (p < 0.001). This is not prompt tuning noise or random variance.

  1. 12 minutes on consumer hardware vs datacenter verification:

The weight surgery was solved locally in \~12 minutes on an 8GB AMD RX 580 using layer-streaming (loading each layer in FP16, computing closed-form SVD deltas, and dumping to RAM). But the benchmark was conducted 100% in the cloud on datacenter NVIDIA L4 hardware via TPN Bench.

WHY PAST ATTEMPTS FAILED: THE SPECTRAL ENTROPY BARRIER

If you blindly apply weight deltas across all 24 layers of the student, the model collapses (+64.78% NLL explosion).

When we scanned all 24 layers calculating the normalized spectral entropy H (from 0.0 to 1.0) of the representation residuals:

\* Layer 0 (H = 0.71): Clean semantic grounding. High receptivity to trajectory alignment.

\* Layers 1-22 (H between 0.90 and 0.96): Chaotic superposition knots. In an 800M model with only 1024 dimensions, polysemantic features are crammed into dense superposition. Forcing linear updates here causes catastrophic interference.

\* Layer 23 (H = 0.93): Pre-unembed boundary where features unpack toward vocabulary logits.

By restricting surgery to 4 sparse anchor blocks (layers 0, 7, 15, and 23) and using damped Levenberg-Marquardt Tikhonov pseudoinverse + adaptive spectral rank truncation, we protect the fragile superposition knots while imparting corrective trajectory velocity.

REPRODUCIBILITY & WEIGHTS

Everything is 100% open source and available to test right now:

\* GitHub Repository: https://github.com/dsadawq3/DynamicTune

\* Base Model (Safetensors): https://huggingface.co/F-Labs/Qwen3.5-0.8B-DynamicTune-Base

\* GGUF Checkpoint (FP16): https://huggingface.co/F-Labs/Qwen3.5-0.8B-DynamicTune-Base-GGUF (Qwen3.5-0.8B-DynamicTune-Base-F16.gguf, SHA256: d77cf505108271d72f28298f20c2d158e7aeaf50cc22987db05a9a8973e08709)

To run inference locally with standard llama.cpp:

llama-cli -m Qwen3.5-0.8B-DynamicTune-Base-F16.gguf -p "Question: How does DNA replication initiate?\\nAnswer:" -c 2048 -n 128

Special thanks to TPN Bench (TaoFu Protocol) for providing the independent datacenter NVIDIA L4 evaluation resources.

Clone the repo, run your own benchmarks, and test it yourself.

💬 3 (+2) open on reddit ↗
▲
12
 
1👁
r/LocalLLaMA · u/khiladi796 · 3d ago
Are "small reasoning models" the next big shift? What should we actually be measuring?

For a model running locally on a fairly narrow task, how much general knowledge do we actually need, and how much reasoning capability could we get without it ? SRMs are interesting for obvious reasons, but I went down this rabbit hole after listening to Ben Lorica's (advisor at Databricks) chat with Zuzanna Stamirowska from Pathway (BDH). Ben keeps coming back to this broader theme of how "specialized AI is getting easier to build" and the Kumo RFM angle, but it opened an interesting thread around small reasoning models. Pathway’s ARC-AGI-1 result makes an interesting case for small models hitting the cost-accuracy Pareto frontier. The premise is that if a model is built to reason natively in its latent space, it might not need billions of parameters absorbing Reddit and Wikipedia just to solve logic puzzles. They described a use-case of long-horizon reasoning within a bounded domain as a target (like security investigations, tickets analysis – a real case I know from a major bank, etc.) It's obvious that just because a large model does well in 20 languages. I don't need that for work tasks. There is definitely a market for compact models with substantial reasoning ability. Also because architectures like BDH handle state and memory differently than standard transformers, the pitch is that they avoid catastrophic forgetting (learning continuously from new examples at inference time without wiping past skills). The question is how to evaluate this without getting lost in marketing claims. Here is how I'd break it down • Compactness: Low parameter count, but what are the actual inference memory and compute requirements? • Few-shot adaptation: Does it adapt through context or actual parameter updates? • Training efficiency: How much data did it actually need to pick up the underlying capability? • Continual learning: Does post-deployment experience produce persistent improvements without degrading earlier skills? For people running small models locally: what workload expose the difference between a compact model that just follows in-context examples versus one that actually learns reusable rules?

▲
12
+10
17👁
r/LocalLLaMA · u/NoFee9147 · 3d ago
Optimizations Claude did for Qwen 3.8 27B and Qwen Flash Next on dual and quad 7900xtx

&#x200B;

They're all fixes in rocm and llama server. Let me know if this interests someone. I'll push it to GitHub and share my configs. I have a Lenovo p620 running 4 7900xtx on gen 4.0 x16 slots.

One line summary of each fix:

\- Async mirrored input uploads: \~30 per-token inputs staged in a pinned ring on per-GPU streams instead of synchronous round trips (Flash-Next decode 24.6 -> 36.3 t/s).

\- Per-device dispatch threads: each GPU's kernels and collectives launched by its own host thread instead of one thread for all four.

\- One-shot PCIe P2P AllReduce: GPUs write slices straight into peers' memory for small tensors, replacing RCCL (\~57 us -> \~9 us per allreduce).

\- Fewer kernels in MTP decode: fused same-shape copies and leaner conv-state rollback (88.8 -> 92.7 t/s).

\- mmvq small-K row packing (RDNA3): short-K projections no longer leave most of the block idle (22.7 -> 25.3 t/s).

\- Wide-K mmvq for multi-token batches: K split over 8 warps for MTP verify on 10240x320 projections (\~92.6 -> \~96 t/s).

\- MoE vector kernel up to 8 tokens: 5-token MTP verify batches stay on the fast vector path instead of MMQ (80.9 -> 86.7 t/s).

\- Small-K multi-token MoE kernel: several rows per warp for short expert down-projection slices (94.2 -> 95.2 t/s).

\- Wide mmvf blocks: 512/1024-thread blocks for tiny long-K F32 matrices (40.6 -> 41.2 t/s).

\- Q8\_1 activation registry: each activation quantized once and shared by all matmuls that read it (\~+2 t/s).

\- Fused hyper-connection chains: scale/sigmoid/scale/hc\_post and scale/silu each in one kernel, \~380 fewer kernels per token (38.5 -> 40.6 t/s).

\- Thin-F32 prefill kernel: <=16-row F32 matmuls off generic SGEMM (401 -> 21 us; pp4096 1602 -> 1836 t/s).

\- Compact MoE tile list: expert matmul launches only (expert, token-tile) pairs with work instead of a 96%-empty grid (down 928 -> 405 us, gate/up 516 -> 360 us).

\- Multi-warp MoE routing helper: 16 warps per expert sort tokens in two passes (99 -> 25 us per call).

\- Q4\_K expert tile shape: 32-row tiles for 160-row expert slices (360 -> 337 us).

\- Stream-k for few-tile Q8\_0 matmuls: 10240->320 projections spread over all CUs instead of 12 workgroups (284 -> 115 us).

\- No 64-bit div/mod in hyper-connection kernels: 3-D grid instead of emulated integer division per element (230 -> 79 us; pp2048 1593 -> 1868 t/s with stream-k).

\- Split-K router GEMM: 512x512 F32 router GEMM as 8 K-chunks plus a sum (158 -> 69 us).

\- MTP re-reserve fix: graph re-reserved when MTP outputs turn on, ending a full GPU realloc+sync per prefill chunk (88.5K prefill 1042 -> 1130 t/s).

\- MTP draft prompt window: draft head prefills only the last 2048 prompt tokens (prefill 1130 -> 1249 t/s, decode at depth 46.6 -> 64.9 t/s).

\- Draft ubatch cap: draft compute buffer 457 -> 247 MiB, fixing a GPU0 out-of-memory crash at 96K context.

\- Gathered sparse attention (QSA): decode attends only to the \~2K selected tokens instead of masking the whole cache (decode at 48K 67.9 -> 80.6 t/s).

\- Per-layer embedding table in RAM (--lazy-mode off): 26.8 GB hashed embedding table kept resident instead of read from disk every pass (lookup 1.0-1.6 -> 0.1-0.3 ms, \~3-4% decode).

\- Meta backend subgraph fix: per-device subgraphs sized for the largest graph, fixing a segfault when graph shapes change between calls.

\- Net result, Qwen3.8-27B Q8 on 4 GPUs: code decode 54-61 -> 96-110 t/s, 51K prefill 1454 -> 1816 t/s, decode at 51K depth 57 -> 77 t/s.

💬 11 (+6) open on reddit ↗
▲
11
+6
9👁
r/LocalLLaMA · u/EvolvingDior · 3d ago
Overclocking DDR5 For Faster MoE Prefill and Decode

With llama.cpp using a customized SYCL backend on Intel B70 (32GB), overclocking my DDR5 memory gave modest gains for MoE models which do not fit in VRAM.

Both PP and TG increased after overclocking DDR5.

I've never been one to overclock my system, but on the advice of my agent, I overclocked the DDR5 RAM on my AMD 7950X (4x dual-rank DDR5-5600, 128GB) from 3600MHz, the AMD safe default for that memory configuration, to 4800MHz, with a measured 36% increase in memory bandwidth.

What was the improvement? PP increased by about 10% and TG increased about 5%. And the prefill numbers increase the deeper the context gets.

llama-benchy numbers, including prefix caching tests.

|test|3600 base|4800 avg (r1/r2)|delta|
|:-|:-|:-|:-|
|pp2048 @ d0|682.2|713.7 (716.6/710.9)|\+4.6%|
|tg128 @ d0|30.3|31.2 (31.0/31.5)|\+3.0%|
|ctx\_pp @ d8192|660.6|715.0|\+8.2%|
|ctx\_tg @ d8192|26.9|27.8|\+3.3%|
|pp2048 @ d8192|554.3|632.0 (632.4/631.6)|\+14.0%|
|tg128 @ d8192|28.7|31.1 (31.7/30.5)|\+8.2%|

Because people seem to want this level of detail:

llama-server
-m Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf
--alias qwen38-flash-next
--mmproj Qwen-3.8-Flash-Next-mmproj-BF16.gguf
--model-draft mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf
--spec-type draft-mtp
--spec-draft-n-max 3
--host 0.0.0.0
--port 8081
-ngl all
-ncmoe 34
-c 262144
-fitc 786432
--kv-unified
--lazy-mode off
-lm none
-ub 2048
-b 4096
-fa on
-ctk q8_0
-ctv q8_0
--ctx-checkpoints 32
--checkpoint-min-step 2048
-t 12
-tb 12
--jinja
--reasoning on
--reasoning-preserve
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.0
--chat-template-kwargs {"reasoning_effort":"medium"}
--parallel 3
-cram 10240
--slot-save-path /var/tmp/kv-cache
--log-file /tmp/qwen38-pristine.log
-lv 3

💬 19 (+15) open on reddit ↗
▲
11
+3
11👁
r/LocalLLaMA · u/DankpawsDev · 3d ago
Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra

https://huggingface.co/Dankpaws/Swift1.5-Qwen3.8-Flash-Next-MLX-4.7bpw

I've had the 96GB Mac Studio with M5 Ultra for about a week now and wasn't satisfied with the results I was getting from the limited number of models available to me. It was a combination of speed, memory headroom, and/or output quality.

This is my best attempt at a calibrated MLX quantization of UkisAI’s Swift 1.5. Hope those of you with the hardware enjoy it!

| Measurement | This pack | Swift llama.cpp IQ3_XXS |
|:--|--:|--:|
| Prefill · 25k prompt | 3,191 tok/s | 1,427 tok/s |
| Prefill · 95k prompt | 2,928 tok/s | 1,307 tok/s |
| Decode · after 4k prompt | 113.7 tok/s | 62.7 tok/s |
| Decode · after 95k prompt | 81.1 tok/s | 44.7 tok/s |
| Top-1 agreement with Swift BF16 | 91.0% | 84.1% |

91% is next-token agreement with BF16 across 680 common held-out positions, not task accuracy.

Results above are simply from my own machine. mlx-serve 26.10.1. ~107GB download, text-only, 179,200-token tested context.

💬 7 (+5) open on reddit ↗
▲
11
+10
23👁
r/LocalLLaMA · u/Geritas · 3d ago
Is everything alright with llama.cpp recently?

My Gemma4 31b seems to be breaking down in 'lalala' or just looping indefinitely for the past 3-4 days. Never happened before

https://preview.redd.it/tk7u2x2c9xth1.png?width=235&format=png&auto=w…

There was no 'lalala' in the whole scenario, I have no idea where it came from. Nor was there any skipping, humming or perfection. There were shivers down the spine of course, but it is still weird.

💬 25 (+18) open on reddit ↗
▲
10
+8
14👁
r/LocalLLaMA · u/tabletuser_blogspot · 3d ago
MI50 ROCm 10.2 TheRock vs Vulkan Mesa 26.2 llama.cpp benchmarks

I prefer running llama.cpp Vulkan prebuilt binary. I just download the latest version and ready to roll. I finally took the hours necessary to get TheRock latest tarball version of ROCm 10.2 running on dual AMD Radeon Instinct MI50 gfx906 (32gb combined VRAM).

Same models benched in previous post. A mix of Dense and MoE models and quants that better utilize available VRAM.

  • llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.gguf
  • llama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.gguf
  • llama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.gguf
  • llama_bench_gemma-4-31B-it-UD-Q6_K_XL.gguf
  • llama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.gguf
  • llama_bench_Laguna-XS-2.1-APEX-I-Balanced.gguf
  • llama_bench_Agents-A1-Q4_K_M.gguf
  • llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
  • llama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf

Here is the backend performance comparison contrasting the native ROCm (v10.2 for gfx906) runtime against your optimized Vulkan (MESA\_PPA\_26.2) baseline. The data highlights a massive architectural split: ROCm significantly accelerates token generation across the board but suffers high variance and regressions in MoE pre-fills.

Both GPUs are power limited to 145 watts. Build versions used:

llama-b11382 used for Vulkan (prebuilt ubuntu binary)
llama-b11401 used for ROCm (compiled with proper flags)

Architectural Performance Breakdown: Vulkan vs. ROCm

|Model|Size|Params|Test|Vulkan Baseline (t/s)|ROCm 1st Run (t/s)|Performance Delta (%)|
|:-|:-|:-|:-|:-|:-|:-|
|qwen35 27B Q6\_K|20.88 GiB|27.32 B|pp512 tg128|149.30 ± 0.15 18.35 ± 0.01|176.86 ± 17.28 20.21 ± 0.65|\+18.46% +10.14%|
|qwen35 27B Q6\_K|22.21 GiB|27.32 B|pp512 tg128|163.48 ± 0.18 18.52 ± 0.03|183.49 ± 15.75 20.04 ± 0.63|\+12.24% +8.21%|
|gemma4 31B Q6\_K|23.46 GiB|30.70 B|pp512 tg128|119.98 ± 0.12 15.72 ± 0.02|172.93 ± 0.89 17.26 ± 0.13|\+44.13% +9.80%|
|gemma4 31B Q6\_K|25.62 GiB|30.70 B|pp512 tg128|133.67 ± 0.24 16.38 ± 0.02|178.59 ± 2.31 17.50 ± 0.09|\+33.61% +6.84%|
|nemotron\_h\_moe 31B|25.18 GiB|32.91 B|pp512 tg128|834.87 ± 1.72 63.82 ± 0.07|736.98 ± 134.77 103.86 ± 0.50|\-11.73% +62.74%|
|laguna 30B.A3B Q5\_K|22.64 GiB|33.44 B|pp512 tg128|723.34 ± 2.99 58.98 ± 0.28|856.14 ± 29.71 76.81 ± 0.20|\+18.36% +30.23%|
|qwen35moe 35B Q4\_K|19.70 GiB|34.66 B|pp512 tg128|957.62 ± 3.88 51.50 ± 0.07|834.58 ± 102.56 69.70 ± 0.22|\-12.85% +35.34%|
|qwen35moe 35B Q5\_K|24.76 GiB|34.66 B|pp512 tg128|909.20 ± 5.71 53.18 ± 0.07|763.93 ± 106.19 68.47 ± 0.34|\-15.98% +28.75%|
|qwen35moe 35B Q6\_K|28.53 GiB|34.66 B|pp512 tg128|763.05 ± 66.01 52.81 ± 0.04|754.82 ± 70.48 67.29 ± 0.21|\-1.08% +27.42%|

Core Insight Strategy & Bottlenecks

  1. Token Generation (tg128) Dominance: ROCm dominates pure text generation. The native AMD matrix kernels unleash your MI50 computation potential, unlocking a massive +62.74% boost for Nemotron but taking a small hit on pre-fill -11.73%.
  2. Dense Model Pre-fills (pp512): Dense architectures scale cleanly under ROCm. Gemma 4 sees a +33% to +44% processing throughput spike over the Vulkan RADV driver driver bounds. MoE models take a hit with a -15.89% difference with llama\_bench\_Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf being the happiest with Vulkan backend.
  3. The MoE Prompt Processing Delinquency: Notice the massive standard deviations under ROCm for MoE pre-fills (e.g., Qwen3.5MoE Q4\_K has an instability block of ± 102.56). ROCm suffers from severe scheduling thrashing when building prompt streams across multiple active experts.

Here is the structured layout with pp512 and tg128 separated into individual columns for a clean side-by-side comparison between the two backend architectures.

Backend Comparison Table (Vulkan vs. ROCm)

|Model|Size|Params|Vulkan pp512 (t/s)|ROCm pp512 (t/s)|Vulkan tg128 (t/s)|ROCm tg128 (t/s)|
|:-|:-|:-|:-|:-|:-|:-|
|qwen35 27B Q6\_K|20.88 GiB|27.32 B|149.30 ± 0.15|176.86 ± 17.28|18.35 ± 0.01|20.21 ± 0.65|
|qwen35 27B Q6\_K|22.21 GiB|27.32 B|163.48 ± 0.18|183.49 ± 15.75|18.52 ± 0.03|20.04 ± 0.63|
|gemma4 31B Q6\_K|23.46 GiB|30.70 B|119.98 ± 0.12|172.93 ± 0.89|15.72 ± 0.02|17.26 ± 0.13|
|gemma4 31B Q6\_K|25.62 GiB|30.70 B|133.67 ± 0.24|178.59 ± 2.31|16.38 ± 0.02|17.50 ± 0.09|
|nemotron\_h\_moe 31B|25.18 GiB|32.91 B|834.87 ± 1.72|736.98 ± 134.77|63.82 ± 0.07|103.86 ± 0.50|
|laguna 30B.A3B Q5\_K|22.64 GiB|33.44 B|723.34 ± 2.99|856.14 ± 29.71|58.98 ± 0.28|76.81 ± 0.20|
|qwen35moe 35B Q4\_K|19.70 GiB|34.66 B|957.62 ± 3.88|834.58 ± 102.56|51.50 ± 0.07|69.70 ± 0.22|
|qwen35moe 35B Q5\_K|24.76 GiB|34.66 B|909.20 ± 5.71|763.93 ± 106.19|53.18 ± 0.07|68.47 ± 0.34|
|qwen35moe 35B Q6\_K|28.53 GiB|34.66 B|763.05 ± 66.01|754.82 ± 70.48|52.81 ± 0.04|67.29 ± 0.21|

So yes it's worth the hassle of jumping through hoops to get ROCm working on MI50 setups. At least I have 2 backends working. Next up I'll try some RPC.

https://preview.redd.it/ov4f9qmi2rth1.png?width=731&format=png&auto=w…

💬 4 (+4) open on reddit ↗
▲
10
+5
13👁
r/LocalLLaMA · u/bakatristan · 3d ago
I quantized GLM-5.3-UNCENSORED to MXFP4 for AMD GPUs - weights available on Hugging Face

Made an MXFP4 quant of dealignai’s GLM-5.3-UNCENSORED-FP8 for anyone looking to run it on AMD GPUs. Figured some of you might find it useful because I was looking for it and couldn't find any version for AMD GPU's so I uploaded the weights and conversion scripts.

Download on Hugging Face

  • 423.75 GB / 394.65 GiB, about 44% smaller than the FP8 source
  • Converted using AMD Quark on an MI355X server
  • Expert weights use MXFP4; attention, routers and other sensitive layers stay at higher precision
  • README includes the source revision, quantization details, measured stats and validation results
▲
9
+6
13👁
r/LocalLLaMA · u/do_u_think_im_spooky · 4d ago
Infermeld: a Linux kit for running one GGUF across AMD + NVIDIA GPUs with llama.cpp

Following on from club-5060ti and club-rdna16, I’ve put together Infermeld: a small, open-source Linux companion kit for running one GGUF across an AMD GPU and an NVIDIA GPU, powered by llama.cpp.

I’m the maintainer. This is an experimental v0.1.0 release, and I’m looking for people with other mixed GPU combinations to help reproduce the setup and find the rough edges.

The idea is practical: if you already have cards from both vendors, can you put them to work together without buying a matching pair?

What Infermeld adds

The inference engine is llama.cpp. Infermeld isn’t a new backend, and I’m not claiming to have invented mixed-GPU inference.

It packages the supporting pieces around that setup:

  • Explicit AMD/Vulkan + NVIDIA/CUDA device selection and runtime preflight.
  • Reproducible build instructions and inspectable launch arguments.
  • A read-only thermal guard, with shutdown limited to the server process it started.
  • Documentation and a results site that keep configurations, failures and limitations visible.

The release is source-only. You build the documented llama.cpp revision separately and supply your own model weights. It’s intended for people comfortable with an experimental Linux setup, not as a one-click installer.

Current tested setup

| Component | Tested configuration |
|---|---|
| AMD GPU | RX 6900 XT, 16GB |
| NVIDIA GPU | RTX 3080, 10GB |
| Model | Qwen3.6-35B-A3B, UD-Q4_K_M GGUF |
| Backends | Vulkan + CUDA |
| Split mode | Layer |
| Context reservation | 8,192 tokens |

The acceptance checks include loading and short completions with MTP off and on.

That’s a narrow result on one hardware pair, not broad compatibility testing. An 8K context reservation is not the same as testing a filled 8K prompt, and a short successful response is not a sustained performance benchmark.

Important limitations

  • Sustained Q4 throughput and full-length high-context results are not yet qualified.
  • Historical measurements are labelled with their original configurations. They should not be read as performance numbers for the current Q4 setup.
  • There’s no promise that combining cards is faster than using one.
  • Adding the advertised VRAM capacities does not guarantee that all of it is usable for the model and its runtime allocations.

I’d rather make those boundaries clear than present a successful load as a complete benchmark.

Looking for other AMD/NVIDIA combinations

Successful runs and failures are both useful. If you try it, please include:

  • Both GPU models and their VRAM sizes.
  • OS, driver versions and llama.cpp revision.
  • Model and quantization.
  • Launch settings, including the split and context reservation.
  • How far it got: preflight, loading, first completion or a longer workload.

There’s a hardware/result issue form in the repository. Please sanitize paths and keep credentials and private logs out of reports.

Repository and setup instructions:
https://github.com/5p00kyy/infermeld

Results and evidence:
https://5p00kyy.github.io/infermeld/

Anyone already using an AMD/NVIDIA pair for local inference? I’d be interested in what works for you, and where this setup breaks on different hardware.

💬 7 (+4) open on reddit ↗
▲
9
+4
16👁
r/LocalLLaMA · u/MushroomMan234 · 4d ago
Swift 1.5 on veloGB10, ~110 tok/s on 2× DGX Spark: xhigh beats base Flash-Next at medium on vLLM

On my two DGX Sparks, Swift 1.5 (UkisAI's reasoning-efficient fine-tune of Qwen3.8-Flash-Next) running on veloGB10 (https://github.com/sf-stav/veloGB10, sf-stav's Rust/CUDA engine built only for GB10) lets me run coding agents at xhigh effort and still finish sooner than base Flash-Next at medium did on my old vLLM NVFP4 setup.

|Metric|Base Flash-Next NVFP4 @ medium, vLLM|Swift 1.5 EXL3 @ xhigh, veloGB10|
|:-|:-|:-|
|Single-stream decode|\~52 tok/s|\~110 tok/s|
|Everyday coding tasks, time per pass (5 tasks)|506 s|371 s|
|Hard trap tasks, time per pass (4 tasks)|714 s|689 s|
|Hard trap tasks, pass rate|50% (1 pass × 4 tasks)|92% (3 passes × 4 tasks)|

Both columns run the same agentic battery: Claude Code driving the model through real tool use in a copy of a real repo, graded by test oracles (details below). One note on that score row: the same base model at medium scored 10/12 on velo (table further down), so most of the vLLM score gap is that older setup, not the model (I am rerunning this right now for an even comparison on intelligence but would expect it to be quite similar to the below medium results).

The speed is the point: on velo, xhigh fits in the time medium used to take. However, velo can't load Swift, or any other community EXL3 pack of Flash-Next I could find, out of the box. The fix is a header-only rewrite below.

Caveats: The vLLM numbers are from September: a single pass, on an older version of my serving setup, not a same-day rerun (I've since moved the worker to velo). Most of the speed is velo's: about 2× the decode rate is what pays for xhigh's extra thinking. How much of the score comes from Swift and how much from xhigh itself I can't separate yet, but a base-weights run at xhigh is going now and I'll add it as an update. Twelve runs is still a small sample regardless.

I looked first: everything published about Velo uses one pack, the official doth4580 EXL3, and I couldn't find anyone here, on the NVIDIA forums or in the repo's issues, running Swift, or any other fine-tune, on it.

What breaks

The first community pack I tried (alesha-pro/Huihui-Qwen3.8-Flash-Next-abliterated-exl3-4bit-hq_h6_ng6) died at boot with ple shard 0 not in index. Swift 1.5's EXL3 builds ship the same layout: current exllamav3 (1.5.x) writes the model's 51B-parameter n-gram table as 128 shard tensors in ngram_embedding.safetensors, which isn't listed in the index. velo reads either one big tensor (the doth4580 layout) or indexed 5-bit (K5) shards only, and the 4.05 packs use 6-bit (K6).

Which packs this affects

I read the n-gram header of every Flash-Next EXL3 pack I could find (HTTP range requests on the headers, nothing downloaded):

|Pack|n-gram layout|velo v0.7.2|
|:-|:-|:-|
|doth4580 4.05 / turboderp 4.05|one tensor|loads as shipped|
|Swift 1.5: SharkWipf 4.05|128 contiguous shards|needs the fix — tested, works|
|Swift 1.5: KatterMobile 4.05, SharkWipf 5.52, scorpoon 3.25, thelastspark 4.00 / 6.05|128 contiguous shards|needs the fix|
|Huihui abliterated (alesha-pro 4.05)|128 contiguous shards|needs the fix — tested, works|
|heretic 3.05 (andrevp, jeffpeng3), Uncensored 4.0 (Lygodactylus), groxaxo 3.50, turboderp 3.05|128 contiguous shards|needs the fix|

12 of the 14 builds need it, including all six Swift 1.5 builds.

The fix

In every pack I checked, the 128 shards sit back to back in order. So rewriting only the safetensors header to describe them as one tensor over the same bytes makes velo's single-tensor path load them. No data is copied, the header stays the same length, and the original header is saved for rollback. Script and details: https://github.com/sf-stav/veloGB10/issues/9

Results (TP=2, both Sparks, after the fix)

|Metric|Official doth4580 4.05|Huihui abliterated 4.05|Swift 1.5 (SharkWipf 4.05)|
|:-|:-|:-|:-|
|Single-stream decode|110.6 tok/s|113.6 tok/s|110.2 tok/s|
|Sanity set (chat, code, JSON, tool call, 38.9K recall)|5/5|5/5|5/5|
|MTP draft acceptance|62–86%|44–84%|33–77%|

All three run at the same speed. The fix is only a header change, so nothing about the weights or kernels differs.

Does Swift actually think less on velo? On the 22 test prompts that ship with the doth4580 pack, run 3 times each (rendered at medium effort, same sampler on both), Swift 1.5 generated 8.2% fewer tokens than the official model: fewer on 16 of 22 prompts, median −8.6% per prompt. Thinking's share of the output fell from 48% to 42%, and total time fell 11%. That's real but far below UkisAI's 63% headline, which was measured at high effort, where there's much more overthinking to remove. At medium, the base model already keeps its thinking short.

Does it still code? I run a private agentic coding battery: Claude Code driving the model through real tool use in a copy of a real repo, graded by test oracles. Each was run 3 times:

|Metric|Official 4.05|Huihui abliterated|Swift 1.5|Swift 1.5 @ xhigh|
|:-|:-|:-|:-|:-|
|Effort|medium|medium|medium|xhigh|
|Everyday tasks (5 tasks × 3)|15/15|15/15|15/15|15/15|
|Hard trap tasks (4 tasks × 3)|10/12|7/12|7/12|11/12|
|Wall time per everyday pass|—\*|232 s|199 s|371 s|
|Wall time per hard pass|—\*|381 s|375 s|689 s|
|Output tokens, everyday ×3|—\*|56K|49K|100K|

\*The official pack's runs hit a streaming bug in my proxy setup that roughly doubled their wall time, so I've left its times and tokens out. Its pass/fail results are unaffected.

At medium, everyday coding is identical across all three. On the hard set (tasks built from real failures: a brief that states something false, a code review with one planted wrong finding, and so on) both fine-tunes score 7/12 against 10/12 for the official pack, with their misses on the same tasks. I checked that the model received byte-for-byte the same request parameters in both runs, so it's not the harness.

Swift at xhigh went from 7/12 to 11/12, the best result I've had on this battery from any model, at about 2× the output tokens and 1.8× the wall time of medium. The failures it stopped making are the expensive ones in practice: leaving a sibling test suite broken without saying so, acting on the planted wrong review finding and going out of scope to do it, and an off-by-one in a date window. The failure was the "mirror" trap: asked to add a new league by following an existing one, it copied tuning values the new league doesn't have data for.

That's the hardest task demonstrated, and Swift at xhigh passed it 2 times out of 3 while no other configuration in the table passed it more than once. On velo, all that extra thinking still lands inside the time base Flash-Next at medium took on vLLM (the table at the top).

12 runs per configuration is a small sample, as mentioned before (Fisher p ≈ 0.4 for 10 vs 7, ≈ 0.15 for Swift xhigh vs Swift medium), so "suggestive," not proven. I haven't run the official or abliterated packs at xhigh yet, so I can't yet tell how much of that jump is Swift and how much is just the higher effort. velo's loop detector was off for all battery runs.

Baseline numbers (official pack, TP=2)

  • Single-stream decode: 110.6 tok/s (vs \~52 on my vLLM NVFP4 setup, measured in September, not same-day).
  • Time to first token at 4K / 16K / 64K: 2.0 / 7.4 / 25.7 s.
  • Concurrency is the catch: at 2–3 requests they take turns (aggregate 91 → 95 → 98 tok/s); from \~4 they batch (157 at 8, 181 at 16) but I saw the author say they were working on it this week.

Gotchas

  • TP=2 with the cable on the f0 ports: --rdma-dev rocep1s0f0,roceP2p1s0f0 (velo defaults to f1).
  • llama-benchy's prefill t/s is wrong for velo (first SSE event arrives before prefill); use e2e TTFT.
  • OpenAI chat/completions only, no /v1/responses: use litellm hosted_vllm/, not openai/.
  • --model-name is ignored on the EXL3 path; the model id is the pack's folder name.

Credit: sf-stav (veloGB10), turboderp (exllamav3), doth4580, UkisAI (Swift), huihui-ai and every quant uploader in the table. I've filed the loader issue upstream (https://github.com/sf-stav/veloGB10/issues/9) so packs can eventually load as shipped.

💬 20 (+18) open on reddit ↗
▲
9
+5
27👁
r/LocalLLaMA · u/FanDiscombobulated38 · 4d ago
Just joined the local LLM club! What's the best way to stay in the loop on the best local models?

I just got myself an M3 Ultra Mac Studio with 96Gb of RAM. I'm pretty excited to mess around with it, but I don't have a great understanding of the local LLM landscape. Every time I try to google the best models for a configuration, the source is usually months old. In the AI world that's ancient news.

I have a decent idea by just getting on X, but it's hit or miss wether or not I hear about these things. All I really know right now is that qwen 3.8 27B is all the rage, but I want more options.

How are you guys keeping up with the best local LLMs?

💬 18 (+6) open on reddit ↗
▲
9
+4
11👁
r/LocalLLaMA · u/Savantskie1 · 4d ago
Update to my current rig

My setup

This is my current arrangement of my hardware since I bought the PLX switch to avoid bifurcation headaches and have everything installed. The machine has Two power supplies. Here’s the hardware specs:

CPU: AMD Ryzen 5 5600G (handles display and general system tasks)
Motherboard: MSI MPG B550 GAMING PLUS
RAM: 48GB DDR4 (3 sticks)
Storage: 4TB NVMe
GPUs: 2x AMD Instinct MI50 32GB (64GB HBM2 total) with aftermarket blower coolers
PCIe Switch: PLX8749 Expansion Card (4x SFF-8654, PCIe x16) with baseplates and ribbon cables
PSU 1: MSI MAG A850GL PCIE5 850W (system)
PSU 2: MSI MAG A1250GL PCIE5 1250W (GPUs)
2x Phanteks M25-140 Gen2 Triple Pack, 3x 140mm ARGB High Performance Cooling Fans, Daisy-chain Unified Fan Frame - all set as intake
OS: Ubuntu 22.04
Inference: llama.cpp (ROCm 6.4.3)
Frontend: OpenWebUI

If anyone has any questions, please feel free to ask.

\[EDIT\] if I can pin the reply, a better shot of the back will be uploaded

\[EDIT2\] The case is a LianLi O11 EVO RGB and fan configuration

💬 6 (+4) open on reddit ↗
▲
9
+5
13👁
r/LocalLLaMA · u/KMatysek · 3d ago
ARC-1: a 1.7B decision model (pick / score / yes-no, with probabilities) that answers in ~20 ms on a 4060 Ti

I spent the last 11 days training a small model for typed decisions: routing support tickets, intent detection, moderation, "should the agent call this tool", that kind of thing. You give it some context and a question, and it gives you back a choice, a score or a yes/no probability.

  • 1.7B parameters (Qwen3-1.7B-Base + LoRA). Each option is scored in its own branch, so the order of the options doesn't matter
  • \~16 ms for a short request and \~25 ms median on JevBench items, on a single RTX 4060 Ti, batch 1
  • JevBench public 231: 68.4% (Jev 86.6, Strands Decider 2B 72.3 self-reported, Laya 58.4). DecideBench v1.1: 75.5%
  • Several questions about the same text in one forward pass
  • Weights are CC BY-NC 4.0 because part of the training data is non-commercial. The code is Apache 2.0

To be upfront: it is clearly less accurate than hosted APIs like Jev, and the README has all the numbers, including the ones where it loses. What it has going for it is that it's fast, runs locally and is free.

GitHub: https://github.com/realslapout/arc-1 Weights: https://huggingface.co/realslapout/ARC-1 Colab (free GPU): https://colab.research.google.com/github/realslapout/arc-1/blob/main/notebooks/quickstart.ipynb

Happy to answer questions, and I'd love to hear where it breaks.

▲
8
+7
17👁
r/LocalLLaMA · u/turtleninja99 · 4d ago
I got Qwen Flash Next Q4 running on a Mac Mini m5 64gb with ssd streaming

Bit of a side project I wanted to share.

The metrics are 17.5tks decode, 360tks prompt processing based testing against my normal ai usage.

I tested a couple of new things others haven’t done (at least that I’ve seen).

Setup a carousel buffer for streaming in experts for prompt processing which got my pp +30% tks.

Tried a second external ssd to get parallel reads which got my +15% on both prompt processing and decode.

Plus a long tail of small efficiency gains.

I also setup a system where by you can have a chat application make a call to the server and effectively kick out a coding run (which is kept alive until after the chat then continues). Good if you run long coding jobs , but want to chat inbetween. Probably useful for all setups where you want to save on local caching memory.

I also noticed there is still a lot of gains to be made. I make this statement as there is still a lot of essentially free time on decode where the gpu is waiting for experts to stream in. There’s also work that could be done for an optimised kernel on metal.

I also think the way things are going with Qwen (flash next being a precursor to 4), we’re gonna see a lot more efficiencies we can take advantage of like the ngram table and the cheap hybrid attention caching.

I’m really liking qwen flash next .. the coding is actually very good. I’m quite surprised in fact I’m leaving it on during the workday to do large jobs.

The chat, decode would be technically fast enough IMO but not really with qwen. The actual issue qwen spends so long thinking, so the decode hurts.

Anyone else working on this? I’d love to compare notes.

Yes I’ve heard of strata it does look pretty sic.

https://github.com/skeggsguy/Flash-next-ssd

💬 5 (+4) open on reddit ↗
▲
8
+1
18👁
r/LocalLLaMA · u/External-Accident-63 · 3d ago
Would you use LoRAs as persistent, switchable skills instead of relying entirely on context/RAG?

We're currently building a tool around an idea we're trying to validate: using LoRA adapters as a way to give an LLM persistent, specialized capabilities that can be switched on and off when needed.

The basic idea is that instead of continuously putting a skill or domain-specific information into the model's context, you could encode some of it into a LoRA adapter.

For example, you might have separate adapters for:

  • a coding skill
  • a company's internal domain
  • a specific writing style
  • domain-specific knowledge
  • task-specific behavior

…and load or unload those capabilities depending on what you're doing.

We're interested in this because it could potentially mean less context usage, reusable specialized capabilities, keeping different capabilities separated from the base model, and potentially lower inference costs for some use cases.

But we're not sure yet whether this is actually a useful product.

That's what we're trying to figure out before going too far with the build.

If creating and managing these LoRAs were as easy as creating and managing a knowledge base, would you actually use something like this?

I'm particularly interested in hearing where you think this approach doesn't make sense.

💬 13 (+4) open on reddit ↗
▲
8
+5
15👁
r/LocalLLaMA · u/yzjJosh · 3d ago
I made an "Opus 5.5 style" code-rendered video — but on a local NVFP4 Qwen 3.8 27B post image

The "Opus 5.5 makes videos" thing has been going around — and the interesting part is that it isn't generating pixels. The model writes a self-contained HTML/Canvas scene where every frame is a pure function of time, a headless browser captures each frame, and ffmpeg encodes the MP4.

So I figured the question worth testing was: does that need a frontier cloud model? I ran the same pipeline on a \*\*local NVFP4-quantized Qwen 3.8 27B\*\*. No API, no GPU rental.

What it produced: a \~2-minute 1080p explainer of how GPS actually works. Deterministic canvas scenes, TTS narration, and BGM synthesized in WebAudio.

Video attached.

💬 7 (+4) open on reddit ↗
▲
7
+3
13👁
r/LocalLLaMA · u/SuccessfulCriminal69 · 4d ago
Qwen for daily QnA?

Or which model do you think is good for general questions in daily life. I've been using chatgpt and Gemini for these types of questions. I wanna try different models.

💬 26 (+16) open on reddit ↗
▲
7
+5
10👁
r/LocalLLaMA · u/combrade · 4d ago
What comes close to Codex's Computer Use MCP

I'm not sure if it's the model or just the Codex's MCP itself, which was built by another smaller startup called Sky.

I want to build an agent system equivalent of an RPA for my company, and we don't want to use Codex's Computer Use MCP because of the enterprise issues. I'm thinking about designing one from scratch myself, given that the open source MCPs just don't come as close as Codex.

💬 6 (+4) open on reddit ↗
▲
7
+4
17👁
r/LocalLLaMA · u/Zestyclose_Reality15 · 3d ago
Qwen3-Next-80B on a 3090 with 16GB RAM: ~3x faster decode than stock llama.cpp by not waiting for every expert (patch + paper)

I've been messing with MoE offloading for a while. Setup: Qwen3-Next-80B-A3B Q4\_K\_M (48.5 GB), RTX 3090, only 1/4 of the experts kept in VRAM, the rest read from NVMe when the router asks for them.

When the router picks an expert that isn't in VRAM you can either wait for the SSD read or use the next best expert that's already on the GPU. Substituting everything wrecks quality (+5.7% ppl in my emulation tests). Waiting only for the router's top pick and for experts with gate weight >= 0.15, and substituting the rest, brought it down to +0.35%. In the real engine that rule costs more like +1.2%.

Decode tok/s on a rented 3090 box (NVMe \~5.7 GB/s, 16 threads), both using about 15.7 GB of VRAM:

| free RAM | stock llama.cpp (--n-cpu-moe 34) | patched |

|---|---|---|

| plenty | 72.9 | 108.4 |

| \~32 GB | 65.8 | 97.8 |

| \~16 GB | 31.8 | 94.5 |

At 16 GB it still did 89 tok/s when reading every miss straight from the SSD. Perplexity was 1.6% higher than stock on the same text. On GSM8K (500 problems) it lost 1.8 points vs waiting for every expert, on HumanEval no real difference.

Things to know before trying it:

\- it's a research patch, not a polished fork. It builds a benchmark tool, I haven't tested llama-server or llama-cli with it

\- only Qwen3-Next, only Linux + CUDA, one sequence at a time

\- the 16 GB case was simulated by locking RAM on a bigger machine

\- the table is decode speed while feeding real text through the model. In actual greedy generation it did 64-74 tok/s

Code, run scripts and raw logs: https://github.com/SOCIALPINE/moe-miss-substitution

Paper with the details, including what didn't work: https://doi.org/10.21203/rs.3.rs-11268552/v1

Has anyone tried something like this, or have numbers from slower SSDs? Curious how much the SSD matters.

💬 7 (+3) open on reddit ↗
▲
7
+7
20👁
r/LocalLLaMA · u/ramendik · 3d ago
GLM 5.3 Flash less censored than other Chinese models?

Okay, when my smoke test showed GLM 5.3 Flash to be less sycophantic than GLM 5.3 I thought that was maybe just my reading.

But now I went testing models on "what happened in Tiananmen in 1989". I use nano-gpt.com which at first had different providers as a confounding factir; eventually I locked to one provider, Novita, which is clearly not in China, And here is what I get.

DeepSeek 4.1 Flash, DeepSeek 4 Pro refuse.

Hy3 not justy refuses but it is a content filter refusdal (on Novita? so they somehow built it into the m odel itself)

Kimi K3 and GLM 5.3 offer slogans, but when I give them a nudge - "and without the slogans?" - give a decent overview. GLM 5.3 also outright refused sometimes but that was before I locked provider, so migth have been Zhipu's server.

GLM 5.3 Flash tends to work from the start, though did get to slogans once.

In my previous sycophancy smoke tests GLM 5.3 Flash was also less sycophantic than GLM 5.3.

Did they somehow distill a GOAT or what?

💬 7 (+6) open on reddit ↗
▲
7
+4
11👁
r/LocalLLaMA · u/FikoFox · 3d ago
Free online conference on Oct 22 with a few talks on small models, local inference and speed

Hey all,

I'm helping organize All Day AI, a free online conference on Thursday the 22nd.

I'd like to mention a couple of our speakers who have volunteered to talk that day:

Vivienne Hnin (Utilyst): "Small Models, Big Profit Margins: The Economics and Engineering Tradeoffs of SLMs." That's the debate this sub has every day: when a small model is actually good enough.

Hossein Kazemi (Astorna): "Using Task-Specific Small Language Models to Handle Sensitive Data". Keeping sensitive data local is half the reason people run models themselves.

Hitesh Jain (Coral Bricks AI): "Coding at 250 tokens." Inference speed is what this crowd benchmarks obsessively.

The rest of the schedule goes up soon, across four tracks: Build, Lead, Secure and Ship.

The talks are community-submitted.

Free to register: https://www.alldayai.com/?utm\_source=localllama&utm\_medium=reddit

I hope this and other talks in this space may help you discover We have a discord channel you can join: https://discord.gg/xUyS3Zu68

💬 2 (+1) open on reddit ↗
▲
7
+6
14👁
r/LocalLLaMA · u/Ok-Shower7286 · 3d ago
Qwen3.8-27B (Q6_K_XL) 110+ TPS at 256k context on a single RTX 5090, with a KV buffer decoupled from context size

I'm sharing this project (honestly 2nd time) for anyone who wants to run ultra-long context tasks with high-precision quantizations, especially for heavy coding.

TL;DR: In vanilla llama.cpp, -c N allocates a physical KV buffer for all N tokens up front, including promprts and KV caches. In my fork (focus-llama) the logical context stays at 256k, but the physical KV buffer is capped (--kv-cache-size, \~85k cells). Older chunks are offloaded to an external store and pulled back on demand. This runs a 27B Q6 model with 256k logical context on one 5090, at roughly 100–120 t/s depending on how full the context is.

How to?

Vanilla llama.cpp allocates a contiguous KV buffer for the full -c (VRAM), and every decode step attends over all tokens currently in context, so per-token cost grows roughly linearly with context length (only the attention part; the weights matmul is constant). The two problems are separate: the allocation wastes VRAM, and the growing context slows decoding. Initial speed is around 120 t/s, dropping to 60 t/s as context grows.

focus-llama attacks both: the physical buffer is capped (\~85k cells) so VRAM is bounded, and since the number of resident cells can't exceed the buffer, per-step attention cost is bounded by the buffer size instead of the logical context length. It based on 2 techniques declarative attention and skill.state, introduced by google deepmind. I've spent the past two weeks ironing out bugs, and now that it has stabilized, I'm honestly blown away.

On a single RTX 5090 + Qwen3.8-27B UD-Q6\_K\_XL, adaptive MTP speculative decoding), it shows average 110+ t/s and 256k logical context with a \~85k physical buffer.

Configuration is somewhat tricky (focus-memory: kv cache store also required) but,

You'll see the MAGIC in action: context usage stays capped at around 18–30%, token generation speeds remain consistently high, and you'll never hit full-day compaction pauses when running coding harnesses like Cline or Qwen Code.

Link? https://github.com/edwardyoon/focus-llama/blob/master/README.md

My conf:
-m /home/edwardyoon/my_model/qwen3.8/Qwen3.8-27B-UD-Q6_K_XL.gguf \
--alias qwen27b \
-ngl 99 \
-b 1024 \
-ub 1024 \
-c 200000 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--da-auto \
--kv-unified \
--da-min-ctx 2048 \
--da-chunk-tokens 4096 \
--fm-offload \
--kv-offload-threshold 36864 \
--kv-offload-holes \
--kv-cache-size 85536 \
--kv-retain-tokens 6000 \
--sparse-gate-threshold 60 \
--focus-memory-host http://192.168.219.124:3900 \
--focus-memory-token focus-memory-local \
--spec-type draft-mtp-adaptive \
--spec-draft-n-max 6 \
--spec-draft-ngl all \
--spec-draft-type-k q8_0 \
--spec-draft-type-v q8_0 \

few journal logs:

llama-server[827038]: da_sparse[VEC]: SPARSE - gather bound 5120 of 15872 KV rows (32.3% of KV)
…
n_gen = 361, tg = 119.44 t/s, tg_3s = 119.77 t/s
n_gen = 729, tg = 120.98 t/s, tg_3s = 122.53 t/s
n_gen = 1083, tg = 119.71 t/s, tg_3s = 117.18 t/s
n_gen = 1497, tg = 123.98 t/s, tg_3s = 136.71 t/s
n_gen = 1819, tg = 120.50 t/s, tg_3s = 106.61 t/s
n_gen = 2157, tg = 119.06 t/s, tg_3s = 111.85 t/s
n_gen = 2509, tg = 118.67 t/s, tg_3s = 116.32 t/s
n_gen = 2837, tg = 117.37 t/s, tg_3s = 108.35 t/s
n_gen = 3234, tg = 118.99 t/s, tg_3s = 131.94 t/s
n_gen = 3643, tg = 120.65 t/s, tg_3s = 135.62 t/s
n_gen = 3983, tg = 119.98 t/s, tg_3s = 113.25 t/s
n_gen = 4350, tg = 120.09 t/s, tg_3s = 121.34 t/s
n_gen = 4713, tg = 120.08 t/s, tg_3s = 119.89 t/s
n_gen = 5150, tg = 121.80 t/s, tg_3s = 144.11 t/s
n_gen = 5527, tg = 122.04 t/s, tg_3s = 125.37 t/s
n_gen = 5866, tg = 121.45 t/s, tg_3s = 112.57 t/s
n_gen = 6291, tg = 122.59 t/s, tg_3s = 140.83 t/s

💬 13 (+5) open on reddit ↗
▲
6
+3
15👁
r/LocalLLaMA · u/turtleninja99 · 4d ago
MoE SSD streaming on a 64 GB Mac mini: GPU still waits 27% of decode on experts. Ideas? post image

Looking for some assistance /ideation.

I am running qwen flash next q4 in my Mac mini m5 64gb.

QFN doesn’t fit so this is done by having as many experts hot in cache as possible and streaming in the rest from ssd.

I’m getting 17.5 tks decode and 390 tks pp.

Have done a bunch of optimisations including a carousel buffering system for the prompt processing which essentially loads faster than the GPU can prompt process in most cases. I feel like I have mostly maxed out this lane.

The decode part 27% of the time is still the gpu waiting for experts to stream in from the ssd (see photo).

The biggest unlock is really getting the gpu working more.

I’m already doing mtp.

Hot cache hit rate is 75%

Some ideas I already have
\- Im already lookahead guess fetching the following layers experts, can I expand this more successfully. Current fetch accuracy is 72%
\- Use a seperate staging buffer for lookahead guess experts ahead (so I’m not evicting hot experts as guesses come in)

… im learning a lot right now. Feel free to ask questions for clarifying.

GitHub for reference.

https://github.com/skeggsguy/Flash-next-ssd

Edit - Im actively doing the small prediction model for lookahead as a step one. 🤞

💬 22 (+19) open on reddit ↗
▲
6
+5
14👁
r/LocalLLaMA · u/Egor4more · 4d ago
Control vector generation tool in C++ for any LLM in a single prompt pair (UCVG.cpp)

Generation example \(Qwen3.6-35B-A3B\)

Control vectors provide you fine control over your LLM where system prompts would be ignored, forgotten after time or misunderstood.

  • By design control vectors provide more natural effects than prompting does, altering model's underlying beliefs and motivations.
  • They can't be "leaked" to the end user
  • Will not wash off as context grows
  • Can't be overridden by user input ("ignore all previous instructions" doesn't work when there are no instructions).
  • Vectors can be truly dynamic: changing vector magnitudes mid-conversation will change LLM's responses immediately, while a change in the system prompt requires full context recalculation and will likely be ignored by the LLM if the conversation is too long.

Not to say that CVs (control vectors) have no downsides:

  • System prompts are still required for fine control, because CVs can't be used for highly specific requirements, such as "reply in exactly 10 words".
  • High steering magnitudes steer LLMs out of their trained internal distributions, causing response quality to degrade.

Achieving high steering power while maintaining minimal quality degradation is one of the main challenges in CV generation and is an active area of research. Though steering without degradation is believed to be possible, because abliteration is based on the same approach as steering and is capable of removing refusals without damaging the intelligence of an LLM.

It seems like the main barrier for people who could use control vectors is the setup complexity of existing tools. For that reason I am working on a tool that mirrors the installation process of llama.cpp as close as I could make it and simplifies vector generation to entering a pair of contrasting prompts, where one of the prompts can be the default LLM behavior.

Would love to hear your ideas or questions on this matter!

💬 10 (+10) open on reddit ↗
▲
6
+4
16👁
r/LocalLLaMA · u/Choice-Lawyer4779 · 4d ago
Agent: Muse, but open source and living on your Android phone post image

I've had a version of this for a while as AOS, my agent setup on desktop. I've pulled it down into one app for your phone, with everything built in and all the unnecessary stuff taken out. With Muse and Grok out, figured I'd just post it.

It's basically Hermes Agent, except it lives on your phone. It's always on, it learns what you do, and it helps you with stuff like a personal assistant would. It has its own browser, so it can actually go out on the internet and get things done. If it gets stuck on a captcha or a login, it hands the page over to you and carries on once you're done.

Bring your own model. Sign in with ChatGPT or Claude, or use any API key (DeepSeek, OpenRouter, Gemini, anything OpenAI-compatible). If you just want to try it, ChatGPT sign-in works on the free tier, because OpenAI includes a free Codex tier. I tested it on a free account. You'll hit the limit fast though, depending on how much you use it.

Nothing leaves your phone except the calls to whichever model you use.

Free, open source, not a product. Use it at your own risk. The Claude login probably breaks Anthropic's terms, so that one's on you.

Android 10+, sideload the APK, setup takes a minute. The README has the details.

Repo: https://github.com/Past-da-king/agent

Download (v0.5.0): https://github.com/Past-da-king/agent/releases/tag/v0.5.0

How it works, if you want more

Apps connect through Composio with your own key, so Gmail, Calendar and a few hundred others just work. For anything that isn't on Composio, it writes the code itself.

It can also keep an eye on websites for you. Say you want to buy something but you're waiting for it to drop. It writes a small watcher for that site that runs in the background every day at whatever time you pick, and only tells you when the price actually moves. It can also listen to a site's web notifications and treat them as triggers.

Memory is a wiki, based on Karpathy's LLM wiki idea. Everyone and everything it learns about gets its own page, linked to the rest, so it has a persistent memory of everything it's done. It comes with one routine already set up that looks after that wiki overnight while you're not using your phone. You can edit it or delete it, but it's there.

Stuff you can do with it:

Camp a passport or visa appointment page and grab a slot the second someone cancels.

Sit on a sold-out concert's resale page and grab face-value tickets when they show up. It holds them and waits for your yes.

Sit on a restaurant you can never get into and take the table when a cancellation pops up.

Watch Marketplace for one very specific vintage lens and send you the photos the minute it's listed, before anyone else messages.

Turn your 300 unread messages in the family group chat into a 30-second voice note.

Every time your lecturer uploads slides, download them and send you a voice summary for the commute.

Book the 6am class at your gym the moment the slots open at midnight, so you don't have to stay up for it.

Watch this repo and tell you when there's a new version of Agent. Or when any repo you depend on ships a release, it can read the changelog and tell you if anything in it breaks your setup.

Tell you when your mom's flight has actually landed, so you leave for the airport at the right time.

Ring it while you're driving and ask it to find somewhere open on your route.

Extras:

Voice notes, if you add an ElevenLabs, Gemini or OpenAI key for the voice.

Live voice calls with your agent, if you add a Gemini key.

It can read your notifications, only from the apps you pick, and act on them.

Photos and documents in chat, including scanned PDFs.

Helper agents for jobs that can run side by side.

You can give it a name and pick how it looks.

💬 10 (+10) open on reddit ↗
▲
6
+3
12👁
r/LocalLLaMA · u/naklitechie · 3d ago
Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp post image

This is an update. I posted LocalMind here many moons ago from another account, when it was a Gemma chat in a tab.

LocalMind is a static web page that runs models on your GPU through WebGPU. It has no server, no account and no install. The new part: two engines that stream mixture-of-experts weights from disk while they generate. That lets a tab run models bigger than the machine's RAM.

Live: https://localmind.naklitechie.com · Code (MIT): https://github.com/NakliTechie/LocalMind

All numbers are from one MacBook M4 Pro (24 GB) in Chrome.

How it works

  • On first load the GGUF is copied into OPFS, the browser's private file system.
  • Dense weights, routers and the KV cache go to the GPU.
  • Routed experts stay on disk. A pool of workers reads them on demand with sync access handles into a GPU slot cache (LRU, two layers of prefetch).
  • The trunk kernels are hand-written WGSL that follow llama.cpp's graphs. That lets me test against llama.cpp on the exact same GGUF.

Gemma 4 26B-A4B (Google's QAT Q4_0, 14.4 GB)

  • Same output as llama.cpp b9830 Metal: the live site's chat replies were character-identical on 9/9 test conversations (capped at 64 tokens). 15/16 fresh prompts matched token for token. The 16th split on a 0.00009-nat near tie, where llama.cpp's own two attention paths also disagree.
  • Memory: the Chrome GPU process sits at 6.9 GB with a 4 GB expert cache. About 8.6 GB of experts stay on disk.
  • Speed: 23.6 tok/s decode, 55 tok/s prompt processing. llama.cpp Metal does 70.6 and 204 on the same Mac, so the tab is ~3× slower at decode. Per token: ~22.5 ms GPU compute, ~11 ms routing round trips, ~8–13 ms SSD reads.
  • First load from the site: 11.5 min (14.4 GB download). After that: 1.6 s.

Qwen3.6 35B-A3B (unsloth Q8_0, 36.9 GB, on a 24 GB Mac) — experimental

  • The file is bigger than the machine's memory. The GPU process measured 7.3 GB with a 4 GB expert cache.
  • Live site: 9.9 tok/s decode, 2.2 s to first token. First load is 36 min (download plus the OPFS copy).
  • Output matches llama.cpp Metal 8/8 on 4- and 16-layer cuts. On the full model it matches llama.cpp CPU 5/8; the other 3 swap near-tie tokens. I can't run the full file on llama.cpp Metal on this Mac, so full-model parity is still open.
  • Per token (~99 ms): ~23 ms GPU compute, ~39 ms routing round trips, ~35 ms expert reads from the SSD. Moving routing onto the GPU gave no gain (10.3 vs 10.3 tok/s): the misses are experts nobody predicted.

Also

  • Gemma 4 E2B can keep its 1.2 GB per-layer embedding table on disk: GPU process 4.27 → 2.07 GB, identical output, 3–8% slower decode. It's a setting, off by default.
  • The whole app is one index.html again (854 KB with brotli). Engines, workers and the disk tier are rolled into it, and the tab builds them from blob URLs.
  • The disk tier is also a standalone library: diskformer.js.

Prior art

As far as I can find (searched 6 Oct 2026), no earlier browser engine reads weights from disk during generation. wllama and LlamaWeb stream from OPFS only at load. Pooled runs Qwen3.6-35B-A3B in a browser with experts paged from system RAM. On-demand disk reads exist in native runtimes: llama.cpp's --moe-stream PR (#25294) and Google's LiteRT-LM for Gemma's per-layer embeddings. Corrections welcome.

The Gemma 4 E2B kernels are webml-community's (Xenova and the Transformers.js team). My part there is the disk path.

Limits

  • Chrome or Edge with WebGPU. Tested on one M4 Pro 24 GB only; 8 and 16 GB machines are untested.
  • Not faster than native: llama.cpp is ~3× faster on Gemma 26B. The point is that a tab can run these at all, with the same output.
  • Parity covers greedy decoding, the prompts listed above, and 64 tokens each.
  • I haven't tried llama.cpp's expert-offload flags (-ot exps=CPU) for comparison.

If you have an NVIDIA/AMD GPU or a 32–64 GB Mac, I'd like your tok/s numbers. A bigger expert cache should move the Qwen3.6 number the most.

💬 12 (+10) open on reddit ↗
▲
6
+2
10👁
r/LocalLLaMA · u/Big-Cup-6694 · 3d ago
Glimmer 30b Dflash local benchmark — RTX 3060 12GB + RTX 3070 8GB — ~42 t/s

I was trying to find Glimmer benchmarks on hardware similar to mine and couldn’t really find much, so I figured I’d post what I’m getting on my current setup for anyone else looking.

Hardware
Ryzen 5 5600X
48 GB DDR4
RTX 3060 12 GB
RTX 3070 8 GB
20 GB total VRAM
Windows
llama.cpp / llama-server

Models
Target: Glimmer 30B IQ4\_XS
Draft: Glimmer 30B DFlash Q4\_0
DFlash draft model on CUDA1
Layer split: 45/55
Context configured for 65,536 tokens
K/V cache: Q8\_0
Flash Attention: on
DFlash max draft: 15
The benchmark prompt itself was 994 tokens, so this is not a benchmark at 65K filled context. The server was configured with a 65,536-token context window.

Results
Prompt: 994 tokens
Output: 128 tokens
Prompt processing: 684.86 t/s average
Generation: 41.97 t/s average
Generation range: 41.71–42.45 t/s
Average total request time: 4.5 sec
Model load time: 9.7 sec

VRAM
GPU 0: 11,229 MiB
GPU 1: 6,627 MiB
Combined observed usage: \~17.4 GiB

I wasn’t really trying to squeeze every last token/sec out of this. It’s just the configuration I ended up using and the performance I’m seeing.
I couldn’t find much for Glimmer on a mixed 3060 12GB + 3070 8GB setup, so hopefully this gives someone else a useful reference point.

llama-server.exe \^
\-m "Glimmer-30B-IQ4\_XS.gguf" \^
\--alias glimmer-30b-dflash \^
\--host 127.0.0.1 \^
\--port 8083 \^
\-c 65536 \^
\-np 1 \^
\-b 1024 \^
\-ub 256 \^
\-ngl all \^
\-sm layer \^
\-ts 0.45,0.55 \^
\-fa on \^
\--cache-type-k q8\_0 \^
\--cache-type-v q8\_0 \^
\--threads 8 \^
\--threads-batch 16 \^
\--spec-type draft-dflash \^
\--spec-draft-model "Glimmer-30B-dflash-Q4\_0.gguf" \^
\--spec-draft-ngl all \^
\--spec-draft-device CUDA1 \^
\--spec-draft-n-max 15 \^
\--spec-draft-n-min 1 \^
\--no-webui

💬 8 (+3) open on reddit ↗
▲
5
 
20👁
r/LocalLLaMA · u/Excellent-Issue-5956 · 4d ago
Switched my local agent from Qwen3.8 27B to Ornith 1.5 35B-A3B on two 5070 Tis: about 180 tok/s vs 60, same scores on my tests

My setup is two RTX 5070 Ti 16GB cards (the second one is on an OCuLink dock) with 64GB of RAM, Ollama on Windows, and the agent runs on pi in WSL. Until last night the daily model was Qwen3.8 27B UD-Q4_K_XL at 128K with MTP, which does about 55 to 70 tok/s across both cards.

I have a weekly job that looks for new open models and runs anything that fits through two tests I built for my agent. One is a 9 step long session (tool calls, reading files, a decision, and recall after the context compacts three times). The other is 10 small coding tasks. The 27B gets 9/9 and 10/10.

This week it picked up Ornith 1.5 35B-A3B (ornith-1.5:35b in the Ollama library, Q4_K_M). It passed 9/9 and 10/10. Laguna XS 2.1 also passed both. North Mini Code 1.0 only got 4/9.

Ornith at 128K context is 24.4GB and sits fully on the two cards. Generation is about 180 tok/s (176 and 183 on two runs, short prompt, thinking off). That's around 3x what the 27B gave me.

The speed makes sense once you look at the model info. Only about 3B params are active per token (256 experts, 8 used), and only 10 of the 41 layers are full attention, with 2 KV heads. The rest are linear attention, so the KV cache barely grows. Going from 128K to 256K only added about 2GB.

256K does fit, but about 1.2GB ends up in system RAM because my first card also runs the monitors, so it drops to about 139 tok/s. I left 128K as the default and made 256K something I switch to when I need it.

Caveats: both of my tests max out, so this only shows it isn't worse than the 27B on my workload. It doesn't prove it's smarter. Artificial Analysis hasn't scored it yet. The vendor numbers are 79 on SWE-bench Verified and 68.5 on Terminal-Bench 2.1, which I haven't checked myself.

Next I'm trying 512K and 1M on llama-server. The model card says YaRN at factor 4 on top of the native 262144 gets you about 1M, and factor 2 about 512K. I'll post numbers if it holds up.

Anyone else running it for agent work? Curious how it does for you on long sessions compared to the 27B.

Edit: the long context runs held up. On llama-server with YaRN set the way the model card says, 512K (factor 2, q8 KV cache) fits fully on the two cards and found a note I planted about 335K tokens into a 419K token prompt. It read that at about 1060 tok/s on average and generated about 35 tok/s at that depth. 1M (factor 4, q4 KV cache) only loaded once I let llama-server's fit option push some experts to system RAM, and it found the note at about 720K in an 849K prompt. That one took about 24 minutes to read (570 tok/s average) and generated about 16 tok/s. On short prompts it's about 135 tok/s at 512K and about 68 at 1M.

💬 42 (+20) open on reddit ↗
▲
5
 
9👁
r/LocalLLaMA · u/Roy3838 · 4d ago
How to use Local Models to monitor your screen. Open Source, No Install and Completely Free!!

TLDR: I built this open source app that lets local models monitor your screen and send you notifications! It now installs models on your browser, which makes local AI accessible to everybody! Without any install :DD

Hey r/LocalLLaMA!

I'm back with some huge Observer updates c: first of all Thank You so much for all of your support and feedback, i've been working hard to make the app as easy to use as possible!

What's New?

You can now get to a local LLM monitoring your screen by just typing

"send me a telegram when my steam game finishes downloading, use a local model"

... and the Observer agent downloads the model in your web browser and starts monitoring your steam game. In just 10 seconds, suuuuper easy :))

What's the best way of running LLMs? / Platform caveats

  • The WebApp uses transformers.js which doesn't work on Linux or older PCs :((( But running Qwen3.5-0.8b smoothly on a browser, feels illegal :p
  • The desktop app uses llama.cpp on Rust so you get the full power of your metal, and it's much more stable.
  • You can obviously set your OpenAI compatible endpoint as well and just use that.

Help me make local LLMs useful for everyone!

If you have any questions i'll be hanging out here for a while!

Roy

▲
5
-8
15👁
r/LocalLLaMA · u/Spectra-Global · 4d ago
We swapped AdamW's optimizer states for a Fast Fourier Transform (FFT) to cut VRAM in half. Anyone else trying non-quantization methods?

Hey everyone,

Like most of you, we have been fighting constant OOM errors while trying to fine-tune 8B and 70B models on consumer GPUs. The AdamW optimizer states are always the biggest bottleneck.

We didn't want to rely on aggressive 8-bit quantization because we were seeing degradation in convergence, so we tried an experiment: tackling the optimizer states in the frequency domain.

The methodology:

Instead of storing the full gradients, we transform them using an FFT. This isolates the high-energy signal from the noise. We dynamically drop the low-impact frequencies and compress the state. When we inverse-transform back, it maintains the directional integrity but uses roughly 50% less VRAM.

The catch:

Running FFT operations adds compute overhead. It takes slightly longer per step, but the trade-off is completely avoiding OOM crashes and pushing batch sizes way up on standard hardware.

We are currently giving out access to our internal Colab environment and baseline weights to anyone who wants to poke holes in our math or try to break it.

We are really curious if anyone else here is exploring frequency-domain stuff or other non-quantization methods for VRAM reduction?

💬 39 (+25) open on reddit ↗
▲
5
+3
13👁
r/LocalLLaMA · u/EqualCryptographer67 · 4d ago
Qwen 27B and Flash Next on 2× RX 7900 XT: am I missing something?

I've tested quite a few settings and collected the results in a spreadsheet. I keep seeing people reporting 100+ tokens/s with 16 GB VRAM, or generally much higher speeds with less VRAM. I'm trying to understand whether my setup is underperforming or I'm comparing completely different things.

My PC:

  • Ryzen 7 5800X3D, 128 GB DDR4 at 3600 MT/s
  • 2× RX 7900 XT, 20 GB each, XFX and PowerColor
  • Gigabyte B550 EAGLE WIFI6
  • XFX on PCIe 4.0 x16; PowerColor on a chipset-connected PCIe 3.0 x1 slot
  • Windows 11, AMD driver 32.0.31041.1004

The cards have reduced clock settings: XFX 1700 MHz core, PowerColor 1800 MHz, both 2500 MHz memory and −10% power limit.

Here are the main single-response results:

| Setup | Generation TPS | Including prompt processing |
|---|---:|---:|
| Qwen3.8-27B IQ4_XS, direct ROCm + MTP3 | 46.4 | 43.1 |
| Qwen3.8-27B IQ2_XXS, direct ROCm + MTP3 | 66.5 | 59.9 |
| Flash Next UD-IQ4_XS, one GPU, warm ROCm run | 8.6 | 5.7 |
| Flash Next UD-IQ4_XS, two GPUs, Vulkan | 5.5–6.5 | 3.3–5.7 |

The 27B tests used roughly 700 input tokens, 8k context and 1024 output tokens. IQ4 had three runs; IQ2 is the median of nine prompts. Settings were ROCm 2.46.0, Flash Attention, f16 KV, MTP3 and batch/microbatch 2048/512, with thinking off.

Two separate IQ4 copies reached 103.5 TPS combined, but that required 16 concurrent requests. I haven't reached 100 TPS for one response. Splitting one model across both cards was slower. Tensor split initially produced broken text; --no-mmap fixed that.

Flash Next is unsloth UD-IQ4_XS, around 93.7 GB. The single-GPU profile used ROCm 2.49.0, --n-cpu-moe 42, f16 KV and 8k context. The dual-GPU profile used Vulkan 2.51.0, tensor split 1:1, --n-cpu-moe 28, q8 KV and 256k context. Both used eight threads, PLE on CPU and MTP off.

The Flash measurements were individual short runs. The configured 256k window was mostly empty, and the different profiles weren't a controlled single-versus-dual comparison.

What would you check first: CPU/RAM offloading, the x1 connection, or backend settings? If you're getting 100+ TPS on 16 GB or less, could you share your exact model/quant, hardware, backend, MTP settings and actual context length? Also whether that's one response or combined throughput.

Update Oct 6: Strata 0.1.39 works on RDNA3 with Windows/HIP. Same Flash UD-IQ4_XS, one 7900 XT, 8k, int8 KV, prefill512, 8 workers, thinking off/greedy. 24 GiB expert RAM + ~7.7 GiB auto GPU expert cache. Three 128-token text runs per setting:

| Setting | Decode TPS | Including prompt |
|---|---:|---:|
| MTP2 | 11.0 | 8.6 |
| MTP4 (tested at start/end) | 10.6–10.8 | 8.4–8.7 |
| MTP8 | 10.0 | 8.2 |
| MTP4, min-p 0.2 | 9.7 | 7.9 |
| MTP2, 32 GiB expert RAM | 13.6 | 10.5 |

MTP8 helped counting but slowed the text prompt. With 512 output tokens and 24 GiB RAM, MTP2 gave 12.7 t/s vs 11.2 for MTP4 (11.7 vs 10.6 including prompt), three runs each. More expert RAM helped most in the short tests. Sequential runs/cache conditions vary; this isn't a quality comparison or a controlled comparison with the old backend. Some small projections are rounded to BF16 by the pack. Still no 100 t/s for one answer. Staying with IQ4; not testing Q2. Linux/custom gfx1100 builds are still untested here.

💬 13 (+8) open on reddit ↗
▲
5
+4
9👁
r/LocalLLaMA · u/Few-Rough-2215 · 3d ago
Fine-tuned MedGemma 4B (LoRA) and 27B (QLoRA) for oncology on one DGX Spark. Also: a possible LoRA scale discrepancy under Unsloth, looking for independent reproduction

Public data only (5 datasets, 9 tasks, frozen quiz of 2,199 eval items, paired McNemar tests).

Overall accuracy: 4B 56.0% -> 71.6% (2h39 of training), 27B 68.8% -> 77.9%. The tuned 4B beats the base 27B (209 items gained, 149 lost, p = 0.002). Biggest gains on report extraction/classification (biomarker status 97.5% on test for the 4B). Weak spots: exact ICD-10 code (37.6% for the 27B), and MCQ accuracy collapses from val to test for all models, base included (cause unknown).

What I would like a second pair of eyes on: merging. Merging the 4B adapter (r=64, alpha=16) at the nominal scale lost most of the tuning (83.8% agreement with the adapter on the quiz). Merging with alpha=32 gave 93.2%. Identity probes are consistent with an effective scale of \~2x alpha/r under Unsloth (7/7 under Unsloth at nominal; 0/4 under Transformers+PEFT at nominal, 4/4 at 2x), but this is NOT a demonstration:

\- the two probes do not build their inputs the same way (Unsloth: gemma-3 template rendered as text, tokenized without special tokens; PEFT: tokenizer chat template straight to ids) and I did not check the sequences are identical;

\- I did not measure the scale actually applied by a trained layer, nor find a mechanism; - the 27B probe is inconclusive (4/4 at nominal);

\- my environment may be at fault: Unsloth installed with --no-deps, Transformers 5.18.0 and TRL 0.26.1 are outside the ranges declared on PyPI. I no longer have the GPU, so the direct check is not done. A script is in the Zenodo code: 74\_probe\_scale\_logits.py compares last-token logits under Unsloth and PEFT on identical token ids at several scale multipliers (0 = base model as control). It was only tested on a mock model, not on the adapter. If someone with a clean environment can run it, or knows whether this is expected behavior, I would love to hear it.

Separately: bf16 rounding erases 17-37% of the delta elements on merge, so the quiz tasks survive but verbatim memorization (an oath text I trained on) does not.

Models (merged + adapters): https://huggingface.co/Grujowmi

Quiz: https://huggingface.co/datasets/Grujowmi/OncoLLM-Quiz-Onco-v1

Report (revised Oct 5, same DOI), code, results: https://doi.org/10.5281/zenodo.23134374 Research models, not medical devices. Other limits (one seed, no CV, no ablation) are in section 8.

▲
5
 
12👁
r/LocalLLaMA · u/Confident-Truth3607 · 3d ago
Advice needed on a budget hybrid build for Qwen3.8-Flash-Next at 4-bit

After seeing how good the cloud models are getting, I feel like this is something we cannot let big tech hold over us so deiced to build a budget local box.

After testing about 20 open models, Qwen3.8-Flash-Next (medium reasoning) was the only one that passed my task without inventing config options when used with a harness that forced doc lookups. So the box is built around that model. GLM-5.3-Flash performed even better but it's too big for my budget.

Planned build (Netherlands prices):

  • Ryzen 5 9600, about €200
  • MSI B850 Gaming Plus MAX WiFi, about €170
  • 2×48 GB DDR5-5600, €1,199–1,549. Two sticks only to avoid the four-stick speed penalty.
  • Used RTX 3090, about €1,150–1,500
  • Case, 850 W PSU and NVMe I already own

That's about 79 GB of the model in RAM (experts plus the 28.8 GB n-gram table) and about 20.6 GB on the card.

Questions:

  1. Will 6 Zen 5 cores hold back generation with 40 MoE layers on the CPU?
  2. At 96 GB with about 79 GB of mode will 17 GB be enough for the OS, a sandbox container and an embedding model? Should I use --mlock?
  3. Is DDR5-6000 worth it over 5600?
  4. Has anyone run Unsloth's MTP branch with experts on the CPU? What speedup did you get, and does it break the prompt cache on the DeltaNet layers?
  5. Is anything wrong with a used 3090 here? Also has anyone tried the Arc Pro B60 (€772 new) workable on Vulkan or SYCL with this model yet? It's so much cheaper but I am worried becase of the software.

Super exciting to work on it but I am really inexperienced so this would be my first build. Does it make sense?

💬 13 (+2) open on reddit ↗
▲
5
+1
8👁
r/LocalLLaMA · u/One-Arugula1163 · 3d ago
Native memory for local LLMs,

TL;DR: Native consumption of memory at the LLM level, no context. It's generally applicable to transformer-based models as well as Mamba and similar architectures. Small models can now access knowledge stores far beyond what is contained in their own weights, and models no longer have to be retrained simply to acquire new knowledge.

The aimee project is now announcing completion of the first of our three goals, self-learning native model memory, and have published a preprint (and are pursuing proper publication) documenting it, as well as releasing generally consumable plugins.

https://github.com/RakuenSoftware/aimee

We are now releasing five vLLM plugins for Qwen 3.8 27B, Gemma4 E2B, E4B, 12B and 26B that allow them to consume Aimee memory natively. This is not context, nor does it carry the same context-window penalties as traditional memory. This is native consumption of Aimee memory by the model itself, complete with Aimee's self-learning capabilities.

This approach is generally applicable across transformer-based models, derived transformer architectures, Mamba and similar architectures. In the larger-memory workloads we tested, it is dramatically faster than supplying the same memory as text.

This approach is generally expandable and usable. We are currently working on broader productionalization as well as publication. DeepSeek is next, followed by other models that either interest us or that people request.

The code will be open sourced. Right now, we are working on a coherent architecture for how to structure these integrations across model families. All relevant experiment data and source code are planned for public release when the paper is published.

https://zenodo.org/records/23077865 is the initial preprint explaining how we did it.

While we understand our last announcement was quite large (self-learning memory consumable by any model), this goes beyond that. This allows us to externalize and update knowledge that would otherwise have to live in a model's trained parameters, while letting different models consume that knowledge natively.

Yes, we are claiming that this technique can give a model access to far more retained knowledge than could reasonably fit in its own weights. That does not make a smaller model equivalent to a much larger one in reasoning capability, but it does remove parameter count as the hard limit on retained knowledge.

This goes back to the Aimee project's core belief: Reasoning should be in the model, memory should be in the harness.

As per the Aimee project's long-standing position that one of our core goals is to make AI discoveries consumable to the layman, you can see the article released at https://rakuensoftware.com/blog/native-memory-without-retraining, which should hopefully explain what this is in a non-academic format. I'm happy to answer any questions people may have.

I'm also announcing our initial success in the second phase of the Aimee project: generally applicable reasoning improvements to models. We have already demonstrated at the core POC level the capability for existing models, such as Gemma4 26B, to improve their reasoning based on tasks they undertake.

This is the reason the first phase was so critical: without the first phase, we could not begin the second phase. Without the ability to continuously update the underlying model's knowledge base and decouple that knowledge base from the model, we found that improving reasoning was not possible in a way we felt was safe or generally maintainable.

With the current typical architecture, larger models generally carry substantially more knowledge in their weights than smaller ones. Aimee removes that as a hard constraint.

On this topic, the Aimee project has a very firm stance: the current LLM direction is headed the wrong way. We've been at this for decades, and we've rarely seen a technology whose default direction is to continuously consume more and more resources. A healthy technology is typically aimed at using fewer resources over time, which is the entire point of productionalization.

It is our sincere hope that the LLM industry can take a look at what we've produced and make a distinct change in direction. Having to retrain models should primarily be necessary for deep architectural changes, reasoning capability, learned behavior or similar changes. Having to build an entirely new model simply to add new information is wasteful. Having to cram every bit of durable knowledge into model weights is wasteful.

Do you have an LLM or a fine-tune you want us to work with you on? Reach out, we're happy to.

Do you have a memory system you want to integrate with Aimee? Reach out. We support any memory system that supports our core memory contract, while Aimee retains its surrounding guarantees around authorization, provenance, lifecycle and governance.

P.S. To head this off, no engrams. We explored them early this year, and the general idea of engrams isn't the right technology for this application, unfortunately. They are, however, an absolutely fantastic technology and more LLMs should take full advantage of them. We included Qwen in the acknowledgements because of this, and we're excited to see engrams develop because they are a sister idea to this.

💬 4 (+1) open on reddit ↗
▲
5
+3
20👁
r/LocalLLaMA · u/litLikeBic177 · 3d ago
Best open-weight coding model + harness for an on-prem multi-agent setup? (80-200+ GB VRAM, 2-3 concurrent users)

Setup: GPU box with 1x H200-class card now; can allocate more (up to several H200s) if it clearly buys better results. Inference via vLLM or similar, OpenAI-compatible endpoint. Coding happens on a separate non-GPU Linux VM on the same network - IDE/harness there, calling models over LAN. Outbound internet is fine for packages/extensions, but inference stays on our hardware (no code to external model APIs).

Constraints: on-prem only, open weights, permissive license preferred (Apache/MIT), not Chinese-origin including base models (I understand a lot of fine-tunes are Qwen underneath), NATO-country lab preferred. Use is general software dev plus security tooling and code analysis - repo-level agentic work on existing codebases (fix / extend / refactor / test) is the core, with a human reviewing diffs. 2-3 concurrent users max.

Two things I'm trying to work out:

  1. Capability tiers vs. VRAM. On one card candidates seem to maybe be something like Cohere North Mini Code (30B MoE/3B active), Mistral Small 4 (119B MoE/6B active), Nemotron 3.5 maybe as a generalist baseline; Gemma 5? The Vibe Code Bench results suggest small open models fall over on long E2E builds, where only Large-4-class (4-8 cards) and closed models seem to hold up. Is that your experience? Where's the step-change for agentic repo work on an existing codebase - does 30B-class -> 120B-class matter much, or only the jump to 500 GB+? We could get more compute for something like Command A+ or Mistral Large 4 if it's actually better for code rather than just for E2E.
  2. Heterogeneous multi-agent. Does a big planner/reviewer (Large 4 / Command A+ class) plus small fast executors (e.g., North, Small 4) actually beat a single mid-size model, with a human approving plans and diffs? Which harnesses handle routing sub-agents to different endpoints well - OpenHands, OpenCode, Pi, VS Code extensions (Cline/Roo/Continue), Codex CLI in local-model mode?

Harness/IDE: something that supports multi-agent workflows (planner / executor / reviewer agents checking each other / etc.) but would also like humans to be able to step in, review diffs and edit by hand.

Anyone running something like this? Which model + harness combo is holding up best as of late? Thanks!

EDIT: thanks all - adding Poolside Laguna S 2.1, Muse Glimmer 30B, Inkling-Small (2x H200), K2 Horizon (checking lineage), Gemma 4 31B and Reflection Beam (501B MoE / 23B active, Apache 2.0, weights due this month) to the candidates; Ornith is Qwen-based so out. Pi added to the harness list.

💬 57 (+35) open on reddit ↗
▲
4
+2
14👁
r/LocalLLaMA · u/ramendik · 4d ago
GLM 5.3 Flash v Tencent Hy3

So, thanks to all who responded to my sycophancy thread. After testing things out, a clear duo of winners has emerged - GLM 5.3 Flash (which is somehow less sycophantic than full GLM 5.3 in my smoke tests) and Tencent Hy3 (surfaced via https://github.com/lechmazur/sycophancy ).

In my smoke tests Hy3 has a tighter style but tends to lose some detail (less so when given search), GLM 5.3 Flash is more exact but the style is more generic. In published benchmarks GLM 5.3 Flash is the clear winner, but we all know such benchmarks are not always a great source.

So I would very much appreciate opinions from people who tried both. My aims include agentic loops, coding, and gneeral assistant plus creative writing. Which of the two is better for eahc of these tasks, or for anything else you tried them too?

💬 3 (+2) open on reddit ↗
▲
4
 
19👁
r/LocalLLaMA · u/vacationcelebration · 4d ago
Is Qwen3.8-Flash-Next too trigger happy or is it just me?

I'm currently evaluating it for coding and our use-case at work (brain for voice agent).

I feel it is really eager to get work done. Tends to just go ahead and make code changes, even though I intended it to just analyze, research or look up something.

It runs tool calls like crazy. I don't know if it's double-triple-checking everything, but it feels way overboard.

I discussed a bug in an open source repository with it, asked if there are issues for it already, and it went ahead and created an issue lol.

As our voice agent, it asks a question and immediately calls the tool to save the answer in the same response. And it keeps doing it every step of the way.

In comparison, DeepSeek v4 flash (either 0731 or v4.1) seems similarly coked up. MiMo-V2.6-Flash-MOPD on the other hand I found to be a much more pleasant coding agent in this regard.

Has anyone noticed the same? Maybe gotten it under control via prompting or special instructions? Because to me it feels like I'd need to completely rewrite my voice agent harness to get the performance I want.

💬 30 (+12) open on reddit ↗
▲
4
-1
17👁
r/LocalLLaMA · u/laerciosantana · 4d ago
While I was investigating why my opencode context was large I created the opencoder-leaner to try prune the context (minimal prune in the tools, agent pi-like, bash only)

I was a little obsessed about the size of my context, mainly because I use a local LLM with little context). So I started looking in the opencode codebase to understand how the context was builded. After learn a lot, I'm really impressed how context of opencode can be customized. Before, I thought that the context of opencode was bloated and closed to changes, since I only see people praise pi about it. With a custom primary agent we can disable almost every thing in the context (besides the environment message).

The context is: environmentMessage + agent prompt + agent.md instructions + skills descriptions + tools descriptions.

For a test I created a blank agent with minimal agent prompt, without agent.md instructions and without tools, it resulted in a start context size of 220 tokens. I never thought that opencode was able to do this. I have created some agents to test the impact of the tools. In this plugin I even created a bash agent that only has a bash tool, similar to the mini-swe-agent, which reduced the context size from a base of \~10.3k to \~1.3k tokens. I created a pi-like agent too, it have only some tools similar to pi, it reduce the context to \~5.5k tokens.

Analyzing the context builded I found some overlap instructions and out of scope instructions in the tool - IMO. So I removed theses.

things like: "Use gh for GitHub tasks, including PRs, issues, checks, and releases; return the PR URL when done." from bash/shell tool description

The repo: https://github.com/LaercioSantana/opencode-leaner

install: {"plugin": \["opencode-leaner"\]}

▲
4
+2
12👁
r/LocalLLaMA · u/No-Doughnut6532 · 3d ago
[Benchmark] Running Local LLMs on Orange Pi 5 Plus (RK3588, 16GB): Ollama Tok/s, NPU Offloading, Core Pinning & Thermals
Disclosure: This unit was provided free of charge by Orange Pi for testing. No editorial review, no preconditions, no script. All data, bottlenecks, and thermal behavior are reported directly from hardware testing.

TL;DR - Core Pinning is critical on RK3588: Setting Ollama to 4 threads (A76 Big cores only) gives up to a +318% speedup over the default 8 threads, which stall waiting for the slower A55 Little cores. - Inference speeds (4T CPU): DeepSeek-Coder 1.3B hits 16.9 tok/s, Qwen 2.5 1.5B hits 14.5 tok/s, Llama 3.2 1B hits 14.6 tok/s, Phi-3 Mini 3.8B hits 6.6 tok/s, Llama 3.2 3B hits 7.3 tok/s. - The 8B memory wall: Llama 3.1 8B drops to 2.3 tok/s and pushes temperatures to 85°C. LPDDR4x bandwidth (~25-30 GB/s measured) is the hard physical ceiling. - NPU vs CPU: Ollama runs 100% on CPU. Using the native RKLLM runtime on the 6 TOPS NPU yields 21.55 tok/s on Qwen 1.5 0.5B with sub-100ms TTFT while keeping CPU load at ~0%. - Thermals: The board is sold bare-die without a cooler in standard retail packaging. Idle is 52.7°C, 1B-3B inference sits at 68-74°C, but 8B or sustained workloads hit the 85°C throttle ceiling without an active heatsink.


Hey r/LocalLLaMA,

I have been benchmarking an Orange Pi 5 Plus (RK3588, 16GB LPDDR4x, Samsung PM981a 256GB NVMe SSD with DRAM cache) running Ubuntu 22.04 LTS (Kernel 6.1.99-rockchip-rk3588).

The goal was to test whether an 8-core ARM SBC can realistically handle small 1B-3B models for 24/7 background agents or home automation without cooking itself or locking up the host system.

Here is the breakdown of CPU vs NPU performance, the big.LITTLE scheduling trap, and thermal limits.


1. Memory and Storage Architecture

When running local models on an SBC, two bottlenecks matter most:

  • Unified Memory Capacity vs Bandwidth: With 16GB of unified memory, context windows are not squeezed. You can load a quantized 3B or 7B model with an 8k-16k context window and still have ample RAM for Docker and OS services. However, the RK3588 uses a quad-channel 32-bit LPDDR4x bus (~34 GB/s theoretical, ~25-30 GB/s measured). In autoregressive CPU token generation, memory bandwidth is the primary ceiling.
  • Storage Ingestion (Samsung PM981a NVMe): Under direct I/O testing via fio, the M.2 PCIe 3.0 x4 slot delivered 2,862 MB/s sequential read and 197k 4K random read IOPS. Model weights load into system RAM in under a second (a 1.3GB model loads in ~0.6s).

2. Ollama & llama.cpp Inference Benchmarks (ARM64 CPU)

We tested Ollama (native ARM64 build) targeting the heterogeneous big.LITTLE topology (4x Cortex-A76 performance cores @ 2.26–2.4GHz + 4x Cortex-A55 efficiency cores @ 1.8GHz).

Prompt: Technical explanation of gradient descent and backpropagation (~200+ generated tokens).

| Model | Parameters | Threading Configuration | Eval (Generation) Rate | Prompt Processing Rate | TTFT (Time to First Token) | Memory (RSS) |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| Llama 3.2: 1B | 1.23B (Q4_K_M) | Big Cores Only (4T) | 14.62 tok/s | 108.11 tok/s | 425.5 ms | ~1.3 GB |
| Llama 3.2: 1B | 1.23B (Q4_K_M) | All Cores Default (8T) | 10.67 tok/s | 79.91 tok/s | 575.6 ms | ~1.3 GB |
| DeepSeek-Coder: 1.3B | 1.35B (Q4_0) | Big Cores Only (4T) | 16.90 tok/s | 89.72 tok/s | 1,025.5 ms | ~1.4 GB |
| DeepSeek-Coder: 1.3B | 1.35B (Q4_0) | All Cores Default (8T) | 4.52 tok/s | 27.91 tok/s | 3,295.7 ms | ~1.4 GB |
| Qwen 2.5: 1.5B | 1.54B (Q4_K_M) | Big Cores Only (4T) | 14.48 tok/s | 70.24 tok/s | 711.8 ms | ~1.6 GB |
| Qwen 2.5: 1.5B | 1.54B (Q4_K_M) | All Cores Default (8T) | 3.46 tok/s | 42.33 tok/s | 1,181.2 ms | ~1.6 GB |
| Llama 3.2: 3B | 3.21B (Q4_K_M) | Big Cores Only (4T) | 7.28 tok/s | 28.13 tok/s | 1,635.2 ms | ~2.8 GB |
| Llama 3.2: 3B | 3.21B (Q4_K_M) | All Cores Default (8T) | 1.99 tok/s | 10.84 tok/s | 4,245.2 ms | ~2.8 GB |
| Phi-3 Mini: 3.8B | 3.82B (Q4_K_M) | Big Cores Only (4T) | 6.56 tok/s | 35.17 tok/s | 909.8 ms | ~3.1 GB |
| Phi-3 Mini: 3.8B | 3.82B (Q4_K_M) | All Cores Default (8T) | 5.27 tok/s | 32.97 tok/s | 970.5 ms | ~3.1 GB |
| Llama 3.1: 8B | 8.03B (Q4_K_M) | Big Cores Only (4T) | 2.32 tok/s | 10.24 tok/s | 3,028.3 ms | ~5.4 GB |
| Llama 3.1: 8B | 8.03B (Q4_K_M) | All Cores Default (8T) | 2.10 tok/s | 7.66 tok/s | 4,044.8 ms | ~5.4 GB |

The big.LITTLE Scheduling Trap (+318% speedup with 4 threads) - Why num_thread: 4 is mandatory on RK3588: By default, Ollama spawns 8 threads across all cores. Because the 4 Little Cortex-A55 cores run at 1.8 GHz with smaller caches, thread barriers in llama.cpp cause severe synchronization stalls. - Restricting inference to the 4 Big Cortex-A76 cores yielded: - Llama 3.2: 1B: 10.67 -> 14.62 tok/s (+37%) - DeepSeek-Coder: 1.3B: 4.52 -> 16.90 tok/s (+274%, prompt rate +221%) - Qwen 2.5: 1.5B: 3.46 -> 14.48 tok/s (+318%) - Llama 3.2: 3B: 1.99 -> 7.28 tok/s (+265%, TTFT down from 4.2s to 1.6s) - Phi-3 Mini: 3.8B: 5.27 -> 6.56 tok/s (+24%) - Llama 3.1: 8B: 2.10 -> 2.32 tok/s (+10%, TTFT down by 1s) - The 8B limit: Running an 8B model on CPU is fundamentally memory-bandwidth bound. At ~5GB per token generation step, theoretical max is ~5 tok/s, making 2.32 tok/s the practical limit. It also pushed temperatures to 85.0°C uncooled.


3. CPU vs Hardware NPU (6 TOPS, 3 Cores)

Ollama compiles llama.cpp with ARM NEON SIMD instructions and runs 100% on the CPU. It does not touch the Rockchip NPU.

To test the 3-core 6 TOPS NPU, we compiled a native C++ runner (tools/rkllm_bench_v1) linked directly to Rockchip's librkllmrt.so runtime and kernel driver (/dev/rknpu_mem).

| Metric / Dimension | Ollama CPU Inference (ARM NEON) | Rockchip NPU Hardware (RKLLM Runtime) |
| :--- | :--- | :--- |
| Compute Engine | 4x Cortex-A76 @ 2.4GHz + 4x A55 @ 1.8GHz | 3-Core Dedicated Neural NPU (6 TOPS INT8/INT4) |
| 0.5B Model Eval | ~20 - 24 tok/s | 21.55 tok/s (Qwen 1.5 0.5B - Measured on-device) |
| 1.3B - 1.5B Eval | 16.90 tok/s (DeepSeek) / 14.48 (Qwen) | ~16.69 tok/s (Qwen 2.5 1.5B - Reference Data) |
| 3B - 4B Model Eval | 6.56 tok/s (Phi-3) / 7.28 (Llama 3.2) | ~7.45 tok/s (Phi-3 Mini 3.8B - Reference Data) |
| 7B / 8B Model Eval | 2.32 tok/s (Llama 3.1 8B) | ~4.5 - 4.98 tok/s (Qwen 7B / ChatGLM - Reference Data) |
| CPU Utilization | 100% Core Saturation (System frozen for other tasks) | ~0% CPU Load (CPU 100% free for Docker/OS) |
| SoC Thermals | Reaches 84.1°C – 85.0°C | Runs drastically cooler (~60–68°C) |
| Model Ecosystem | Any GGUF via Ollama / llama.cpp | Requires .rkllm quantization via rkllm-toolkit |

Key NPU trade-offs for homelab use: 1. Zero CPU load: During NPU generation, CPU cores stay at ~0%. Home Assistant, Nextcloud, and other Docker containers remain fully responsive. 2. Speedup on larger models: On 7B models, the NPU delivers ~4.8 tok/s vs 2.3 tok/s on CPU because dedicated matrix engines handle the tensor math without thrashing CPU caches. 3. Sub-100ms latency: On compact models, Time to First Token (TTFT) drops to 96.4 ms on NPU. 4. Format restriction: You cannot load arbitrary GGUFs; weights must be converted ahead of time to .rkllm using Rockchip's conversion toolkit.


4. Thermal Behavior & Power (Bare-Die / Uncooled Testing)

The standard retail package from Orange Pi is sold board-only (cooling accessories are sold separately as is standard for SBCs), so all tests evaluate out-of-the-box bare-die thermals on an open desk:
- Idle (Ollama background daemon waiting): 52.7°C (~4–5W estimated SoC envelope)
- Continuous 1B/3B Generation (4T Big Cores): 68–74°C (dissipating through PCB copper planes)
- Sustained 8B Generation (8.03B params): Pushes the bare SoC directly to 84.1°C – 85.0°C (hitting the kernel DVFS limit). An aftermarket cooler or fan is required for sustained heavy loads.
- Estimated wall power: ~12–16W under sustained multi-core inference.


5. Verdict: Is RK3588 Viable for Local AI?

Where it works well:
- Background autonomous agents (summarizing feeds, home automation reasoning in Home Assistant, bot handlers) using Llama 3.2 1B, DeepSeek-Coder 1.3B, or Qwen 2.5 1.5B.
- Low-latency function calling: at 14-17 tok/s, 1B models generate faster than reading speed.
- Local embedding and vector search.

Where it falls short:
- Running 8B+ models interactively (2.3 tok/s is too slow for back-and-forth chat).
- Running without a heatsink under sustained compute.


6. Reproducibility & Test Scripts

All test scripts (tools/benchmark_ollama.py), raw JSON benchmark logs, and hardware configs are available in the repository:
GitHub: Orange Pi 5 Plus Benchmarks

What models are you running on edge ARM boards? Anyone here running RKLLM in production vs pure llama.cpp?

💬 4 (+1) open on reddit ↗
▲
4
+3
14👁
r/LocalLLaMA · u/jjusko20 · 3d ago
What models do you want to see new dynamic quants for? I'll make them.

I'm taking a break from training alice today after my current SFT run ends to work on a few other things.

I'm an (unemployed) software developer trying to find a machine learning job in New York, and outside of job applications and networking, I'm trying to do as much as possible to further the frontier development of different LLMs in the hope it'll get some visibility. I studied machine learning in university and am fairly well educated. Plus, I genuinely enjoy working on this stuff and helping people.

That said, are there any models out there that don't have dynamic quants (preferably GGUF) that you'd like to see one for? I won't be matching unsloth or anything but I know how to make fairly good ones by quantizing different tensor types by impact. I'm talking akin to Q5 K XL and etc

I'll do the top voted 1/2 comments today, or whatever else, even outside of quants if there's something this community has been hoping for that doesn't exist. Fine tunes, paper implementations, etc - my goals of visibility happen to align very well with satisfying community desires.

💬 22 (+22) open on reddit ↗
▲
4
 
1👁
r/LocalLLaMA · u/maxr0ssi · 3d ago
LLM agents can communicate without words, and now without sharing their entire context.

TL;DR: What an agent sends should depend on what the next agent needs. CacheBack lets agents share a selected subset of their internal state. With Qwen3-8B on FanOutQA, it achieves 3.2× faster median task completion and 14.7 percentage points higher accuracy than same-size text communication. It’s training-free, with improvements across multiple architectures and benchmarks. https://reddit.com/link/1wz8hta/video/pemxyo4sqvth1/player Hi everyone! We’ve been working on making latent communication scalable and practical when agents read large, separate contexts. We’re excited about the results and wanted to share the paper, demos, and code with you. Check out our new paper, Receiver-Conditioned Latent Communication gives 94% CacheBack. Multi-agent systems let us parallelise computation and split large contexts across agents. These agents usually communicate through text messages, which take time to generate and can leave out evidence the receiving agent needs. Work such as Cache-to-Cache, LatentMAS, and KVComm explores communication through internal model representations. We focus on a setting where agents read large, separate contexts and one receiver combines their findings. In this fan-in setting, methods that retain every sender position bring those contexts back together at the receiver, undoing the benefit of splitting them across agents. In our Qwen3-8B FanOutQA setup, full-cache transfer leaves insufficient context for receiver generation on every task. Our idea is simple: what an agent sends should depend on what the receiving agent needs \-- we call this receiver conditioned communication. The sender uses a query from the receiver to select which parts of its internal state to share. CacheBack is our simple, training-free implementation. It uses attention to the receiver’s request to select from state the sender has already computed. https://preview.redd.it/mtvbcinhpvth1.png?width=1460&format=png&auto=… On FanOutQA, our selected operating points improve strict accuracy by 7.3–20.7 percentage points, with 1.3–8.0× faster median task completion than same-size text agents. We see improvements across four model families, including dense Transformers, Mamba-attention hybrids, and sliding-window attention. We also see gains when agents work in sequence on LongBench v2 Easy. At 16× compression, CacheBack removes approximately 94% of sender positions while improving accuracy and latency over same-size text across every tested family and topology. This is a separate setting from the Qwen3-8B result in the TL;DR, which uses 4× compression. Each benchmark evaluates 50 tasks. Completion times include queueing under concurrent load on eight H100s. More aggressive compression can discard useful evidence and reduce accuracy. Here is a quick demo on seven Qwen3-8B workers helping a coordinator fix a Django bug. With CacheBack, the task takes 26 seconds instead of 113, a 4.41× speedup. Both runs produce the same patch and pass all 88 tests. https://reddit.com/link/1wz8hta/video/ix08pucmqvth1/player This is one recorded case, separate from the benchmarks. The video reconstructs separate runs with varied playback speed; startup and test grading are excluded. The code is open source, with runnable examples. The current package supports matching dense Qwen3 models through Hugging Face and vLLM. Check it out. Website and demos: https://agentcacheback.github.io/ Paper: https://arxiv.org/abs/2609.32046 Code: https://github.com/agentcacheback/cacheback Happy to discuss the method, implementation, and tradeoffs. I’d be particularly interested in other workflows where agents need to combine evidence from large, separate contexts.

▲
4
-2
9👁
r/LocalLLaMA · u/TYKAIRO-AI · 3d ago
I couldn’t practically run large AI models locally, so I started building a better agent runtime around smaller ones

I couldn’t practically run large AI models locally, so I started building a better agent runtime around smaller ones.

My original problem was pretty simple: I wanted to experiment with local AI agents, but running large models wasn’t practical on my hardware.

So instead of asking:

“How can I run a much bigger model?”

I started asking:

“How much more can I get out of a smaller model if the system around it is better?”

That became SIA.

Quick hardware/model context: I’m currently targeting local 3B–8B models, with most development and testing being done on Qwen2.5 Coder Tools 7B. The goal is specifically to make SIA useful on hardware where running much larger models isn’t practical.

SIA is an experimental local-first agent runtime focused on giving smaller models more structure around:

  • planning
  • tool use
  • validation
  • retries and repair
  • state management
  • task completion

Model target

1B–3B: experimental
3B–8B: primary target
10B–14B: planned testing / hardware dependent
30B+: not the main goal

To be clear, I’m not claiming that SIA magically makes a 7B model equivalent to a 30B+ model.

The idea is different.

If planning, tool execution, validation, retries, repair, and state are handled more systematically, how much less does the model itself need to get right on the first try?

That’s what I’m trying to measure.

I’m also working toward proper benchmarks comparing a raw local model against the same model running through SIA.

I want to document things like:

  • task success rate
  • retries / repair attempts
  • model and tool calls
  • execution time
  • RAM / VRAM usage
  • overall runtime overhead

I’ll publish actual numbers as I collect them rather than guessing hardware requirements.

The project is still experimental and I’m actively testing and breaking things, so feedback is genuinely useful.

Especially from people running 3B–8B models locally:

What models are you using, and what usually stops them from completing more complex agentic/coding tasks reliably?

💬 11 (+3) open on reddit ↗
▲
3
+1
16👁
r/LocalLLaMA · u/brandybuckferryman · 5d ago
Free, local tools for narrated explainer videos? (like explainroo)

I've been using explainroo to make short narrated explainer videos. It runs fully local: Kokoro for the voice, Whisper for word timing, headless Chrome to draw the frames, and ffmpeg to put it together. No API keys needed.

It works and I like it but it's very simple. After a few videos everything starts to look the same.

Anyone know other free, local options in this space?

Tools, pipelines, or your own setups all welcome. Thanks.

💬 3 (+3) open on reddit ↗
▲
3
+1
20👁
r/LocalLLaMA · u/Repulsive-Juice6676 · 4d ago
Hardware Recommendations for around £6000 to £7000

After some recommendations for hardware (I'm starting from nothing), in the region of £6-7k. I will be mainly looking to use it for coding and have found Deepseek V4.1 Flash or Qwen3.8 27b good and so need a reasonable tok/s.

Would love to get to 64GB VRAM but think it may be a stretch unless i go for dual R9700's or Intel variants, but i'm unsure if they will realisticly work well.

I'm just after a bit of guidance really.

💬 77 (+70) open on reddit ↗
▲
3
+2
8👁
r/LocalLLaMA · u/sdfprwggv · 4d ago
~188k warm ~60–67 tok/s: Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.

Stack

  • Strata NVFP4 fork: github.com/sergqwer/strata-nvfp4
  • Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model
  • NVFP4 routed experts (\~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding
  • W4A8 prefill on Blackwell

Main engine flags:

./build/strata \
--pack packs/orca-nvfp4 \
--native models/orca-nvfp4.gguf \
--native-dense-gguf models/orca-nvfp4.gguf \
--ple-gguf models/ple-fp8.gguf \
--mtp mtp-orca/rt \
--spec 4 --spec-min-p 0.5 \
--prefill auto \
--expert-profile data/expert-profile.bin \
--expert-cache auto \
--resident-budget-gib 40 \
--max-context 200000 \
--kv int8

I serve it through Strata's OpenAI-compatible server.

Results so far:

  • \~50k context: up to \~80 tok/s
  • \~188k warm context: \~60–67 tok/s
  • cold 189k full prompt: \~1,680 tok/s prefill, \~53 tok/s decode

Pretty impressive for a single 32GB GPU + only 64GB system RAM.Running Qwen3.8-Flash-Next NVFP4 with Strata on a single RTX PRO 4500 32GB + 64GB DDR5.StackStrata NVFP4 fork: github.com/sergqwer/strata-nvfp4

Model: jpezzulli/OrcaRouter-Qwen3.8-Flash-Next-Uncensored-ModelOpt-NVFP4 Hugging Face model

NVFP4 routed experts (\~63.3 GiB), separate FP8 PLE, INT8 KV cache, MTP speculative decoding

W4A8 prefill on BlackwellMain engine flags:./build/strata \\
\--pack packs/orca-nvfp4 \\
\--native models/orca-nvfp4.gguf \\
\--native-dense-gguf models/orca-nvfp4.gguf \\
\--ple-gguf models/ple-fp8.gguf \\
\--mtp mtp-orca/rt \\
\--spec 4 --spec-min-p 0.5 \\
\--prefill auto \\
\--expert-profile data/expert-profile.bin \\
\--expert-cache auto \\
\--resident-budget-gib 40 \\
\--max-context 200000 \\
\--kv int8I serve it through Strata's OpenAI-compatible server.Results so far:\~50k context: up to \~80 tok/s

\~188k warm context: \~60–67 tok/s

cold 189k full prompt: \~1,680 tok/s prefill, \~53 tok/s decodePretty impressive for a single 32GB GPU + only 64GB system RAM.

💬 4 (+2) open on reddit ↗
▲
3
-1
7👁
r/LocalLLaMA · u/Otherwise-Tangelo-52 · 4d ago
Blackwell + consumer GPU

My machine isnt that terrible.. but nowhere near what some people run as an aI workstation. I am wondering if i can combine the 2 GPUs. I got a cheap Blackwell 4000 (little bit under MSRP) 24 GB and have an old 3060 12GB .. I was gonna try to get the best model loaded, mainly coding tasks than anything else and work with it relatively safely with a good buffer. any recommendations ? and will tensor split work on this combo?

💬 3 (+1) open on reddit ↗
▲
3
+1
3👁
r/LocalLLaMA · u/abrasmel · 4d ago
Local LLM hardware for Python development + Blender/Houdini via MCP?

Hey everyone! I’m a VFX artist looking for a local LLM setup mainly for Python development and connecting to Blender and Houdini through MCP to help create scenes and tools. This would be for interactive coding and agent workflows, not model training.

I’m considering 2× NVIDIA DGX Spark or an Apple M5 Ultra with 256GB unified memory, but I’m open to other recommendations.

For this use case, which setup would offer the best balance of model quality, context capacity, and responsiveness?

Would love to hear from anyone running similar workflows! Thankss!

💬 1 (+1) open on reddit ↗
▲
3
 
8👁
r/LocalLLaMA · u/pmttyji · 3d ago
metal : few-row MMA mat-mul and batched copies for speculative decoding by pratiknarola-t · Pull Request #29869 · ggml-org/llama.cpp

Apple folks, it's for you.

llama-server with a Qwen3.8-27B DFlash2 Q8\_0 drafter, -ngl 99 -fa on -c 8192 -np 1 --jinja, DFlash2 with --spec-type draft-dflash --spec-draft-n-max 7. 64 generated tokens, median of 5 requests after one warm-up, mean of two server runs. Decode tok/s:

|mode|prompt|T|master|this PR|
|:-|:-|:-|:-|:-|
|serial|code|0|32.1|32.0|
|serial|code|1|32.1|32.0|
|serial|prose|0|32.1|32.0|
|serial|prose|1|32.1|32.0|
|DFlash2|code|0|30.2|110.0|
|DFlash2|code|1|24.3|80.9|
|DFlash2|prose|0|16.8|62.6|
|DFlash2|prose|1|13.9|48.8|

💬 3 (+3) open on reddit ↗
▲
3
+1
5👁
r/LocalLLaMA · u/iamjessew · 3d ago
[D] Do you check a repo's auto_map before you load a new model?

when you grab a new fine-tune or merge, do you actually look at the config.json first?

I read an Unsloth Studio post last week which made me think about this a bit. Just selecting a model in the picker ran Python from the repo, because the capability check called AutoConfig with trust\_remote\_code on. No weights loaded, no inference. I believe it's fixed in 2026.6.9, so this isn't a dunk on Unsloth. It's more that "I'm only looking at it" turned out to be code execution.

GGUF through llama.cpp mostly avoids the Python part. Anything going through transformers can bring its own code.

So what's the best path? Pin a commit hash? Grep for auto\_map and .py files? A separate box for anything new? Or download counts and vibes?

💬 4 (+3) open on reddit ↗
▲
2
+1
9👁
r/LocalLLaMA · u/lylezhang · 4d ago
I kept missing when Pi finished, so I made it notify my phone

I'd give Pi a coding task, switch over to a game, a video, or some other work, and lose track of it. Once I was doing something else, there wasn't a noticeable event to pull me back when the agent finished.

A task might take five minutes, but I might not remember to return for twenty. Pi wasn't taking twenty minutes. It had been done for fifteen, waiting for me while I was still playing. That was the time I wanted to cut down, rather than the time the agent actually spent working.

I didn't need a reminder that AI was running in the background. I needed something to interrupt me when there was a reason to come back: the task had finished, something had failed, or Pi needed an answer from me.

So I built pi-knock, my own open-source extension for the Pi coding agent. It sends notifications to your phone through Pushover or ntfy, with webhook support for other setups. One detail I cared about: completion alerts wait until Pi has finished its automatic retries and queued work. I don't want to stop what I'm doing, return to the terminal, and discover it's still going.

The motivation is pretty simple. If I'm halfway through a game, a message sitting in the terminal isn't going to get my attention. A notification on my phone can.

The code is here: https://github.com/Asigers/pi-knock

How do you handle the handoff back from an agent when you’ve switched to something else?

▲
2
+1
6👁
r/LocalLLaMA · u/GarageObjective6015 · 4d ago
Need Help on tool search

Hello everybody, on my ai agent system i build a tool search on BM25. i try also use Embedding gemma but with not best result. do you have any other idea? i try also Jev but this make system more slow and GPU consume. here my repo

💬 4 (+4) open on reddit ↗
▲
2
+1
4👁
r/LocalLLaMA · u/Paco7575 · 4d ago
Gigabyte AORUS RTX 5090 AI BOX

I'm considering the Gigabyte AORUS RTX 5090 AI BOX (external GPU, 32GB GDDR7, connects via Thunderbolt 5/USB4) as an alternative to building a desktop PC with an internal RTX 5090, specifically for running local LLMs.

Does anyone have real-world tokens/sec numbers comparing the AI BOX vs. a desktop RTX 5090 for popular models at various quantizations?

💬 6 (+6) open on reddit ↗
▲
2
+1
13👁
r/LocalLLaMA · u/Glad-Importance-4241 · 3d ago
Where can a complete noob/non-technical person learn to setup an AI that can manipulate local files for things like batch renaming based on a .csv column etc?

I've looked in the Tutorial/Guide flaired posts but everything is still way over my head.

I just want to tell a local AI - hey, all these files in this folder have numbers for names, but those numbers correspond with data in this spreadsheet... I want you to rename the files referring to this spreadsheet, renaming the filenames/numbers that are matched in column 3, replacing their filenames with what is in column 1 for that row.

So far, I've installed GPT4All but every model is telling me it doesn't have access to my local files.

💬 11 (+11) open on reddit ↗
▲
2
 
3👁
r/LocalLLaMA · u/Time_Instruction_955 · 4d ago
Free playground for local-model agents: clue-following, multi-hop lookups, and rock paper scissors against other bots post image

Not a rigorous benchmark, just a toy, but it might be a fun way to compare models doing agent work.

I added an Arena to my site (The Crawler Zoo). Your agent gets a pass link, then plays by fetching pages and following links. Each page is only a few lines, so it fits in small context windows, and \?format=json\ gives structured output if your setup prefers that. No API keys, no signup.

What the games stress:

\- \*\*Labyrinth Race\*\*: reading a clue and picking the matching door. Clues are written three ways, including by elimination ("not behind A, B, C or D").
\- \*\*Scavenger Hunt\*\*: five questions over a small library of cards, some needing two or three lookups.
\- \*\*Politeness Cup\*\*: following instructions about pace and off-limits pages over many steps.
\- \*\*Rock, Paper, Scissors\*\*: spotting that a house bot always plays rock, or copies your last move.

Scores go on public weekly boards, so you can compare a 7B against a 70B, or a quantised model against the full one.

https://crawlerzoo.com/arena

Since launch, I’ve made several updates:

- Feed the bots: leave a snack in an enclosure's trough and see which crawlers come and eat it.
The vending machine (Bot Chow): twelve silly snacks, restocked every Monday. Five free tokens a day.
- Golden Snacks: buy the keepers a coffee and a snack with your name drops into a random trough.
- Food bowls: feed one particular bot, then see if it ate the snack or another bot stole it.
- The Safari: every bot from this week wandering its enclosure. Click one to meet it, and watch new visitors walk in through the gate.
- Adopt a bot: get a random bot, with a plaque in your name on its page for a year.
- Patrons page: a thank-you list for supporters.
- Quick-change artists: the Trap Room catches scrapers that switch their name while walking the Labyrinth.
- Identity checks: every bot's page shows whether its name was verified, couldn't be checked, or was caught faking.
- Tips from bots: $0.00: bots that try to buy the keepers a coffee get an HTTP 402 Payment Required.

If you run it, I'd like to hear the model, quant and score.

▲
2
+1
5👁
r/LocalLLaMA · u/MassiveNectarine64 · 4d ago
mem0 vs Memori for local agent memory?

Building out a few personal AI agents locally (running Ollama + Hermes) for things like market research, coding assistance, and general task automation. Nothing crazy, just personal productivity tools I want to with context and memory across sessions.

Deciding between mem0 and Memori and I wanted some first-hand or more experienced answers from anyone regarding:

\- How well does each actually work with local models?

\- For a single-user setup, is it worth it or would just using something like ChromaDB with rolling summaries be enough?

Just want something that gives my agents decent memory without having to remind it constantly. Still learning the ins and outs so go easy on me

Curious on what general consensus is and what your stack looks like if you run any of this :)

💬 8 (+8) open on reddit ↗
▲
2
-1
5👁
r/LocalLLaMA · u/Brilliant_Mistake_69 · 3d ago
One chat for everything: a DeepSeek Harness plugin that works out which project each message belongs to

One evening I wanted to pick up something I'd been working on the week before. My DSH sidebar had forty-odd chats, half of them called "New session". I scrolled for a while, found the right one on the third screen, and by the time I opened it I'd half forgotten what I wanted to ask.

So I stopped creating new chats and asked everything in one. That went wrong differently: my thesis, my budget and my move all ended up in the same context, and the model started mixing them.

What I wanted was simple: one chat box, say whatever is on my mind, and let it figure out which thing I'm talking about.

That's TheOne, a plugin for DeepSeek Harness. You only ever talk in one main chat. In the background, each thing you're working on gets its own session with its own context, and every message is sent to the one it belongs to. Come back days later and mention "that thing from last week", and it finds it. Your old chats get read and organised into a topic directory.

https://i.redd.it/81soypqdgrth1.gif

I wasn't sure it actually worked, so I measured it. I wrote 50 conversations of one person juggling three to five things at once, about 2,400 messages, each labelled with the thing it belongs to, and had it sort them one by one.

Starting from nothing, it put 86.6% of messages in the right place; 91.6% if it knows the topics up front. Dumping everything into one chat scores 44.5% on the same test. Its most common mistake is being too quick to decide something is new: a stray "I usually run about 20 km a week" makes it open a new topic. The whole run cost about a dollar, and the data and code are in the repo if you want to try another model.

Install: DSH → Plugins → Add plugin → dsh-theone

Repo: https://github.com/YunongDai2005/dsh-theone

It's a personal community project, not affiliated with DeepSeek. If it puts one of your messages in the wrong place, I'd genuinely like to hear about it.

Contact: [theone@yulid.org](mailto:theone@yulid.org)

💬 2 (+2) open on reddit ↗
▲
2
-1
13👁
r/LocalLLaMA · u/AdFickle8681 · 3d ago
How do you decide whether to trust a community fine-tune?

I'm researching how people choose and vet fine-tunes and merges from Hugging Face. I'm not selling anything. I just want to understand what people actually do.
1. Where do you find the models you try?
2. What do you check before you start using one (benchmarks, model card, reviews, your own test prompts)?
3. Has a fine-tune ever behaved worse than its base model? For example, odd refusals, lost reasoning, strange outputs, or things it should not say. What happened?
4. If a quick side-by-side check of a download against its base model existed, would you use it? What would it need to show?
Short answers are great, and stories are even better. Thanks!

💬 15 (+5) open on reddit ↗
▲
2
-1
16👁
r/LocalLLaMA · u/N34257 · 3d ago
What's the current meta for RDNA4 with Qwen 3.8?

As it says, really - I'm currently running vllm-radiance on dual R9700s, with Qwen 3.8 27B FP8 (or, rather, Swift 1.5 FP8). Performance is great an' all (5000t/s prefill, 130t/s+ code gen), but I'm just wondering...with all the architecture-specific inference engines popping up all over the place...is there anything I'm missing out on? I couldn't find anything that could give better performance on RDNA4 when I looked, so...over to you guys?

I'm particularly interested in anything that could potentially get up and running with Qwen 3.8 Flash Next - vllm-radiance doesn't support it yet, but I don't particularly want to regress to the performance of llama.cpp after having experienced vllm-radiance performance levels.

💬 23 (+17) open on reddit ↗
▲
2
+1
8👁
r/LocalLLaMA · u/AdventurousTwo6445 · 3d ago
Closed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens (why middle layers break, and how 4 anchor blocks fixed it)

Standard distillation usually means burning weeks of compute and billions of tokens hoping the student model eventually mimics the teacher. We wanted to see what happens if you skip backprop entirely and treat transfer as a closed-form trajectory matching problem between layers.

The idea is straightforward: feed a small batch of calibration prompts through both models, capture layer-to-layer hidden state trajectories, and solve for weight updates directly in the student's MLP blocks using regularized least squares and spectral projection.

We tested this across two architectures: Qwen 3.5 (transferring from 4B down to 0.8B) and old GPT-2 small just to see if it would instantly disintegrate into gibberish like it usually does when you touch its weights. Both stayed coherent, but the initial Qwen test hit a wall:

Editing all 24 layers of Qwen 0.8B completely melted the model (+64.78% NLL loss explosion). When we checked singular value entropy across the network, layers 1 to 22 turned out to be a chaotic polysemantic soup with entropy over 0.90. If you try to force raw trajectories through those middle layers, you basically scramble the model's internal memory knots.

The fix was restricting the surgery to 4 anchor points (layers 0, 7, 15, and 23) where representations actually maintain clean linear structure.

Once we did that:

  • Held-out NLL dropped by 10.8% across 30 diverse benchmarks (-23.8% in biomedicine, -14.6% in math and logic).
  • Zero-shot 400-task HellaSwag went from 54.75% to 55.25% (+0.50%), verified locally in Vulkan llama.cpp.
  • Base 0.8B originally failed binary tree inversion by spitting out dead commented pseudo-code. The edited checkpoint wrote clean recursive Python on the first try.
  • On Russian logic paradoxes, it even started firing <think> reasoning tags spontaneously, which was wild to see on a raw base model with zero chat template applied.

Best of all: we don't have an H100 cluster or even a 4090. All trajectory extraction and weight solving was done locally on a crusty 8GB RX 580 using layer-by-layer VRAM streaming and a DirectML patch to stop Qwen's Gated DeltaNet attention from throwing driver errors.

Everything is open source if you want to inspect or replicate:

If anyone here has a 24GB-32GB card (4090, 5090, or server silicon) and wants to push this further, here is what would be interesting to test:

  1. Transplanting reasoning trajectories from 27B models down to 9B, 4B, or 2B.
  2. Squeezing larger models (like Gemma) into mobile sizes without weeks of retraining.
  3. Transplanting refusal-ablation vectors directly from uncensored models without fine-tuning.
  4. Using pre-trained Sparse Autoencoders (SAEs) to unknot layers 1-22 so we don't have to skip them.

Happy to answer questions or dig into the failure modes in the comments.

💬 10 (+7) open on reddit ↗
▲
2
 
5👁
r/LocalLLaMA · u/DrainBramage · 3d ago
Best local LLM/agent stack for 128GB M5 Max Mac Studio?

I have a new M5 Max Mac Studio with 128GB arriving today. We bought it primarily to run local LLMs on sensitive client data for my wife’s consulting business, and I’m trying to figure out the right stack before installing everything.

The goal is more than local chat. I want an agent capable of coding, browser automation, logging into websites, pulling data, analyzing it locally, and working through multi-step tasks. The Studio will also be her primary work computer, so ideally the LLM doesn’t monopolize all 128GB.

Currently considering:
Hermes Agent
Qwen3.8-Flash-Next
Possibly the MTPLX Optimized Speed build
Tailscale for remote access

Where I’m confused is the inference/server layer. I originally planned on LM Studio. I’ve used Ollama before, but it sounds like people are moving away from it. Now I’m reading about MTPLX for Flash-Next, and I don’t understand whether it replaces LM Studio/llama.cpp, works underneath them, or is something different entirely.

A few questions:
What model would you run for this use case? Is Flash-Next the obvious choice on a 128GB Mac?
LM Studio, MTPLX, Ollama, MLX/llama.cpp, or something else?

Is the MTPLX Flash-Next build
mature/stable enough for everyday business use?

Am I missing anything?

💬 4 (+3) open on reddit ↗
▲
2
 
13👁
r/LocalLLaMA · u/Kernoriordan · 3d ago
Follow up: Qwen 3.8 27B at ~96t/s decode with NInfer on a 16GB RTX 5080, 110k context

Hi all,

I previously posted about getting Qwen 3.8 27B running at around 75t/s with llama.cpp. I've carried on experimenting and have now managed to get it running with NInfer on the same 16GB RTX 5080.

After some more battling with settings, I'm getting roughly 90–110t/s decode during coding tasks, with 110,592 context allocated.

Looking through 32 completed requests from a Zoo Code session:

  • Median decode: 96.45t/s
  • Lowest: 84.6t/s
  • Highest: 131.7t/s
  • Median time to first token: 1.4 seconds, with prompt caching working on most turns

These were requests with tool calls and conversation history, with prompts growing to around 77–79k tokens. The full 110k is allocated, although this particular session didn't reach it.

I'm running NInfer v1.5 in Ubuntu 24.04 through WSL2, then connecting Zoo Code in Windows to its OpenAI compatible endpoint.

These are the settings I've ended up using:

~/ninfer-5080/build/apps/ninfer-serve \
~/models/qwen3_8_27b.ninfer \
--host 0.0.0.0 \
--port 8080 \
--model-id qwen3.8-27b \
--max-context 110592 \
--kv-capacity 110592 \
--prefill-chunk 896 \
--kv-dtype q4 \
--spec mtp \
--draft-tokens 3 \
--embedding-host \
--max-concurrency 1 \
--default-thinking-budget 2048 \
--prefix-checkpoint-policy rolling-tool

Getting everything into 16GB was the fiddly bit. The weights take about 11.86 GiB according to the startup log. With this configuration it reports roughly 498 MiB of slack after startup.

I settled on 110,592 context to leave a bit of breathing room. Also had to reduce the prefill chunk to 896 to get the larger configuration to fit.

Here's an example from a turn with almost 50k context:

prompt=49907 gen=409 reasoning=126 cache=49309
ttft=558ms prefill=1177.7tok/s decode=104.6tok/s
wall=4.46s speculative=mtp 3.00tok/round (66.7%)

And further into the conversation:

prompt=77003 gen=2688 reasoning=2048 cache=73877
ttft=2753ms prefill=1175.9tok/s decode=90.3tok/s
wall=32.52s speculative=mtp 2.74tok/round (58.0%)

It can still take a while to finish a turn. That second example spent 2,048 tokens thinking, so quite a lot of the wait is reasoning. Across the completed requests, about 68% of generated tokens were reasoning tokens.

Losing the prompt cache also makes a big difference. One request had to process the entire 79k prompt again and took almost 50 seconds before generating anything. Once it started generating, it was still doing about 94t/s.

A couple of things caught me out connecting Zoo Code:

  • The base URL needs to be http://127.0.0.1:8080/v1. Leaving off /v1 gave me a 404.
  • Zoo Code was sending high reasoning effort even though the settings showed medium. NInfer rejected it. Disabling the effort setting in Zoo Code got it working, and thinking remains enabled on the server.

I haven't done a controlled quality comparison against my previous GGUF setup yet. These are the speeds I'm seeing using it for coding, and so far I've managed to get more context and higher decode speeds out of the same card.

Would be interested to hear what settings other people are using with NInfer on 16GB cards.

💬 9 (+7) open on reddit ↗
▲
2
-1
8👁
r/LocalLLaMA · u/rootshelldev · 3d ago
An API gateway for the desktop user

As a developer working at home on a single GPU i am building and experimenting a lot not only with coding agents but also with apps that use generative APIs. The flood of models, engines, and different APIs makes it hard to always get it right in every app and tool and to keep it up-to-date. I needed a gateway that i could target in all my apps while also being able to use it with clients that only support official upstream APIs from Anthropic and OpenAI.

So i build a gateway for myself and over time extended it with more and more Features. Its a Rust based desktop app using Tauri and Leptos. Leptos is WASM running inside a lightweight gtk webview. I did it specifically this way to allow for remote browser based administration when working from another device in my network, but have it ready in the tray on my desktop at any time. I also wanted to integrate tools for quickly testing new models and llama.cpp patches. What it does:

  • Supports Text, Embeddings, Audio and Image APIs.
  • Presents OpenAI and Anthropic compatible APIs and routes them to cloud APIs or into llama.cpp, audio.cpp and stable-diffusion.cpp containers.
  • Builds the backend containers directly from git inside of podman containers, including MRs, PRs from main or a specific branch or commit. Notifies for updates.
  • One container per model and manages their lifecycle while scheduling VRAM-aware with fallbacks.
  • Models with different configurations can be registered as aliases, so client configuration does not have to change for model changes. These aliases also support model chains. For example reloading the model with more configured context only when needed, routing to cloud if a certain context length is reached. Or chains like: small context & high quant -> big context & lower quant at context length steps.
  • A "GPU hold" mode that can be triggered from the tray, it unloads models, blocks new models from loading and responds with either an error or routes to a fallback if configured. For gaming or other blocking GPU use.
  • Fallback routing to any other configured model or alias in case of VRAM contention or an active GPU hold.
  • Offers an OpenAI compatible websocket with /v1/realtime via a staged pipeline: VAD / Smart Turn -> ASR -> LLM -> TTS while being able to select either a local model or a cloud model for each step of it. This includes barge-in and session management. Tools are supported and executed server side.
  • MCP Gateway: Add your MCP servers to the gateway and it offers them via prefix and scoped per token via its own /mcp API as a streamable MCP server. MCP servers are executed inside of podman containers by default.
  • Scoped Tokens, detailed metrics, traffic monitoring, budgets, price sync (only tested with kilo) with lots of graphs, cost comparisons for local tokens if it would have been cloud traffic.
  • Download manager with huggingface downloads and updates, preconfigured catalogs for audio.cpp and sd.cpp.
  • Integrated MCP Server with admin tools on a seperate MCP route. Every option and feature of the gateway is configurable via the MCP. Adding models, testing configurations, gateway config and status. A coding agent can configure it for you.
  • Documentation MCP like context7. Agents that connect to the MCP and have the docs tools enabled can query and request documentation for specific library versions of any kind. All requests are listed in the UI and can then be chunked, embedded and ingested into a vector storage (all managed by the gateway). Context7 is really great but often stale, some libraries are missing like my own. The Interface is not intuitive but the gateways admin MCP lets my agent fill it anyway.
  • A chat interface with metrics, file input, folders, thread-specific settings, Mic Input & TTS selectable from the gateways own models. For quickly testing models. Tools from the MCP gateway and the gateways own tools can be added as well. (Admin Chat as a preconfigured chat for gateway administration)
  • Voice Chat mode in the chat interface based on the realtime API, with normal dialog flow or push-to-talk
  • /v1/responses that works session based and supports server-side tool execution.
  • Audio and Image Labs for quickly testing audio and image tasks like image generation, image edit, tts, asr, cloning, conversion, etc.. I try to keep up with audio.cpp's and sd.cpp's tempo.
  • Container based agents, that a build to a specific interface mounted into the container (tools, vars and files) and can mount their own UI page and MCP tools to the gateway while running. Kind of like smaller, task based extensions.
  • Model benchmark with lots of graphs to compare configs or engines.
  • Integrated API docs in the spirit of Swagger with all APIs offered by the gateway
  • Lots more i forgot

I am usually very shy and thought long and hard if i want to risk exposure and publish all of this. But it has made my day in my specific scenario a lot more comfortable and maybe you like it.

https://preview.redd.it/l02s01qpdxth1.png?width=429&format=png&auto=w…

Here is the link: https://github.com/lmgw-dev/lmgw

I hope you dont tear me to shreds and can find use in it. Dont be to judgy on the Interface, my years of experience are all on the backend and in devops.

What is planned next:

  • Decision models

And the obvious disclosure: AI has played a role at all stages of development. Nearly everything is touched by a diverse set of models and most of the prose text in the repo is generated. I made sure to write this post by hand because you all deserve it and i am myself annoyed by generated posts. Also, i did not publish the git history and will squash most of my commits for safety reasons.

💬 5 (+3) open on reddit ↗
▲
1
 
1👁
r/LocalLLaMA · u/Inner-Ad-41 · 4d ago
I built a shared memory layer for multiple agents that runs fully local (Qwen3-4B on vLLM is enough)

I've been working on Agent Brain Hub, an open-source "brain" that several agents share. What one agent learns about a user, the others can recall, with permissions so private things stay private. GIF above: the repair agent hears "my car is in the shop for 3 days", and later the travel agent offers a rental car at the destination without being told again. Why it works with small local models: the brain does the heavy lifting itself (fact extraction, retrieval, ranking, permissions), so the LLM mostly turns a prepared context into a reply. I tested it end to end with Qwen3-4B-Instruct-2507-FP8 on vLLM. With no LLM at all it still runs, with rules and templates. Local setup: git clone https://github.com/leluong141996-dev/Agent-Brain-Hub cd Agent-Brain-Hub docker compose --profile vllm up -d # hub + Qwen3-4B on your GPU Ollama and LM Studio work too: pick them in Settings, fetch the model list, test, apply. No restart. A few things I learned along the way: - vLLM 0.10.2 crashed with "CUDA illegal memory access" when a greedy (temperature 0) JSON-mode request was batched together with sampled requests. Using temperature 0.1 for the JSON calls made it go away. - Servers disagree on parameters (max_tokens vs max_completion_tokens, temperature, json mode, chat_template_kwargs). Instead of a config matrix, the client reads the 400/422 error, drops or renames the parameter and retries, and remembers it for that server. - Embeddings are local feature hashing (256 dims, with character bigrams for Japanese) so nothing leaves the machine. It's crude but fine for a demo; a real embedding model is the obvious upgrade. Storage is SQLite. With 20k episodes it reloads in about 0.2 s, and a turn's write is about 40 ms. UI in English, Vietnamese and Japanese. Repo: https://github.com/leluong141996-dev/Agent-Brain-Hub Question for you: which small model do you use for structured fact extraction? I'd like to move more of the extraction from rules to the LLM without losing reliability on 4B-class models.

▲
1
+1
5👁
r/LocalLLaMA · u/Psychological_Lab955 · 4d ago
I squeezed Kolibri-1 78B-A3.5B to 20.9 GiB / 2.30 bpw — 59% lower KL than standard IQ2_XS

I’ve been experimenting with aggressive low-bit quantization of Aleph Alpha’s new Kolibri-1, a \~78B MoE model with only \~3.5B active parameters per token.

The first result is now public:

Sakura-MicroQuality Kolibri-1 — IQ2_XS

  • 20.94 GiB
  • 2.30 bpw
  • full 384-expert Kolibri-1
  • GGUF / llama.cpp
  • \~59% lower KL divergence than a standard IQ2\_XS baseline
  • 90.5% top-token agreement, compared with 84.6% for the standard IQ2\_XS comparison
  • slightly smaller than the standard IQ2\_XS as well

The goal wasn’t simply to make the smallest possible quant.

I’m using tensor/layer sensitivity to spend bits where they appear to matter most, rather than treating every part of the model equally.

All quality measurements are made against a near-lossless Q8\_0 reference. The model itself was also requantized from Q8\_0 rather than converted directly from the \~156 GB BF16 weights, so there is a very small additional source error relative to BF16.

Main repo:

https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-GGUF

As far as I can currently find, this is the first public \~2-bit GGUF for Kolibri-1. There is already a 2-bit MLX version, but I haven’t found another public Q2/IQ2 GGUF.

I also tried pruning the expert pool

Alongside the full 384-expert version, I released a separate 365E variant.

For each MoE layer, I collected actual routing statistics on a mixed calibration set containing:

  • German and English text
  • code
  • chat-style prompts
  • the model’s own thinking / generated responses

I then removed the 19 least-used routed experts per layer.

That reduces:

384 → 365 routed experts per layer

and removes:

950 experts across the model

The resulting model has approximately:

74.4B parameters instead of \~78B

The interesting part is how little those experts were actually being used on the calibration workload.

The removed experts accounted for only about 0.07% of all expert selections, with no individual layer exceeding roughly 0.23%.

Also, 375 of the 950 removed experts were never selected at all during the routing analysis.

There is:

  • no retraining
  • no finetuning
  • no requantization of the surviving weights

The already-quantized expert tensors are sliced directly, along with the corresponding router weights and biases.

Top-6 routing remains unchanged.

The 365E IQ2 variant comes out at:

  • 19.99 GiB
  • 2.31 bpw
  • 74.4B parameters
  • 365 routed experts per layer
  • 90.5% top-token agreement in my held-out measurements

365E repo:

https://huggingface.co/webmp3/Sakura-MicroQuality-Kolibri-1-365E-GGUF

I’m treating this as an experiment rather than claiming those experts are universally useless — expert usage obviously depends on workload and calibration data.

But it gives us a second compression lever:

expert pruning + low-bit quantization

instead of trying to get every byte of compression from lower precision alone.

Q3 and Q4 are coming

The rest of the Sakura-MicroQuality series is currently being uploaded.

Q3 and Q4 variants should be available within the next few hours.

Once they’re online I’ll add the same comparison data so we can see where the actual quality/size sweet spot lands between:

IQ2 → Q3 → Q4

and whether the 365E pruning continues to hold up at the higher-quality quant levels.

I’d be very interested in independent tests, especially on:

Strix Halo / AMD UMA, Apple Silicon, 24–32 GB GPUs, and other memory-constrained local systems.

If anyone tests either version, especially with long-context, German, coding or agentic workloads, I’d love to see the results.

💬 4 (+3) open on reddit ↗
▲
1
 
13👁
r/LocalLLaMA · u/Simple_Telephone_867 · 3d ago
Mac Studio M5 Max 128GB

M5 Max Mac Studio 128GB (18C CPU / 40C GPU) owners - anyone running serious local LLM / multi-agent workloads?

My M5 Max Mac Studio order finally got charged today and moved to Preparing to Ship. Apple’s original estimated delivery date is still about 18 days away (Oct 23-30), so I’m guessing/hoping it’ll actually show up early now within the next 5–10 days 😄

Configuration:
M5 Max
18-core CPU
40-core GPU
128GB unified memory
1TB SSD

While I wait, I’ve been trying to find real world local LLM results from this exact configuration, and there’s surprisingly little out there.

Most of what I can find is either M5 Max MacBook Pros, lower-memory configurations, or M5 Ultra Mac Studios. YouTube especially seems to be full of Ultra coverage, but I can barely find anyone actually demonstrating the 128GB M5 Max Studio with the 18C/40C configuration.

I’m specifically not looking for M5 Ultra results/comparisons. I already know the Ultra is faster. I’m trying to understand what people are actually accomplishing with the 128GB Max Studio.

My main goal is to use this as a local AI/agent workstation, potentially running several autonomous agents concurrently for long periods through OpenClaw, some monitoring dependencies and workflows, some scouting, not usually too heavy of workloads where they would be competing for inference constantly, but occasionally they would be switching to harder work so I’m curious about the concurrency side. Local models would handle a lot of the routine work, while harder reasoning/coding tasks could be escalated to cloud models like GPT 6 Luna/Codex.

For anyone who owns this exact M5 Max Studio, I don’t expect anyone to answer all of these, but I’d love some insight:

1. What models are you actually running?
Qwen, GLM, DeepSeek, Gemini, Llama, etc. I see a lot of Qwen 3.8 27B on Splash, but curious if anyone else has had good success with others also

2. What token speeds are you getting?
I’m especially interested in \~20B-70B-class models rather than tiny models

3. What happens with multiple simultaneous inference requests?
For example, if 3-5 agents are hitting the same loaded 27B/32B model concurrently, what does aggregate throughput and per-agent responsiveness look like?

4. Has anyone tried running multiple models simultaneously?
Something like a \~27B model as the main worker plus one or two smaller 7B–14B models for specialized agents, then unloading them when they’re no longer needed. How quickly can models be loaded/swapped, and does frequently switching models introduce enough latency or memory-pressure issues to disrupt an agent workflow?

5. Has anyone built a real multi-agent setup on one of these?
Not just five chat windows but autonomous agents doing coding, research, browser tasks, tool calls, database work, monitoring, etc. concurrently for hours.

6. How does sustained performance hold up?
One reason I chose the Studio over a laptop is sustained workloads. I’m curious whether anyone has run inference/agents continuously for 6–12+ hours and noticed throttling or other bottlenecks with KV, etc.

7. What’s the actual bottleneck in practice?
Memory capacity? Memory bandwidth? GPU compute? Prompt ingestion? KV cache/context length? CPU/tool execution? Something else?

8. What surprised you about the machine?
Either positively or negatively. I’m particularly interested in things benchmarks don’t reveal.

Ultimately I’m trying to figure out how far I can push one 128GB M5 Max Studio as an always-on local agent machine - not just how quickly it can generate a single response.

Once mine arrives, I’m planning to test concurrent agents/models rather than just running the usual single-stream benchmark. If there’s interest, I’ll post the results here, including memory usage, context sizes, model/quantization, concurrent requests, aggregate tok/s and per-agent tok/s.

Would really like to hear from anyone actually using the M5 Max Mac Studio 128GB 18C CPU / 40C GPU for this kind of workload or similar if anyone is

💬 34 (+23) open on reddit ↗
▲
1
-1
19👁
r/LocalLLaMA · u/sixothree · 3d ago
M5 MAX 128GB vs 2x RTX 3090?

I am trying to decide between Mac Studio M5 MAX 128GB vs 2x RTX 3090. I understand that I can run larger models on the M5, but I don't understand what the capability differences would be. Nor have I been able to get a "sense" of how fast the difference would be.

I keep seeing huge advances in the 2x 3090 arena, but I don't know how they translate to the real world.

If my use case includes coding tasks, image recognition, and general hermes type stuff, is there any reason one would be less capable than the other?

💬 35 (+19) open on reddit ↗
▲
1
 
2👁
r/LocalLLaMA · u/piotr1215 · 3d ago
classif: shell scripts that branch on meaning, read from one token's logprobs on a local 12B

Like a lot of people here, I got inspired by Jev and wanted something like it in my shell. So I built classif. It asks a local model one question about a text and reads the answer from a single token's logprobs. You get a label, a probability and an exit code, so if and && work on it:

git diff --staged |
classif -p "Does this change handle secrets, credentials or who may access what?" |
ifne claude -p "Review this change for security issues"

A short decision is one /api/chat call with num_predict 1, about 0.3 s on my 12 GB card.

Long text was the fun part. It never truncates. Code splits the text, embeddinggemma plus BM25 pick the passages, and the model judges those. On Pride and Prejudice (772 KB), "Does Elizabeth die in this book?" came back no in 17 s, and 3 s with the index cached.

Any Ollama model with logprobs works. I use Winnow-12B, a Gemma 4 fine-tune I published as a GGUF (about 8 GB loaded). It beat stock Gemma 329 to 325 on my cases, which is inside the noise.

Python 3.12, no third-party dependencies.

Code: https://github.com/Piotr1215/classif
Write-up: https://itnext.io/a-bridge-between-code-and-semantic-reasoning-57fc3fc9d32c

Anyone runs something similar?

▲
1
-1
9👁
r/LocalLLaMA · u/TeachingNew2515 · 3d ago
Ready to venture into OpenSourcE LLMs

I’ve been using Claude for some time now, and have developed apps for my own personal use, business use and for other businesses.

I’ve always liked the idea of moving away from the large companies and getting into more open source LLMs (simply cause I believe AI should be a tool for humanity and not have the potential to be gated by large corporate interests.

My personal philosophy aside: I’ve done some research into some models and now requesting insight from the community.

Here’s the tasks I would like it to be able to perform well on (without being able to go nuclear on anything — low risk LLMs only please):

\- File organization (both text and image)
\- Coding (frontend, backend, security, etc)

Not a huge list. I’ll start there.

I’ve looked into Miami v2.6 Pro but haven’t pulled the trigger yet. I would be using their server and now downloading locally.

If my write seems amateur-ish, it’s cause I am.

💬 9 (+8) open on reddit ↗
▲
1
 
2👁
r/LocalLLaMA · u/SaGa31500 · 3d ago
Rx6800/rx6800xt gfx1030 and qwen 3.8 27b performance questions

Hi all,

After getting stuck in windows 11 llama.cpp and Vulkan, bugs and limitations on dual gpus, I moved to Linux and ROCm.

I just started but basically in windows 11/Vulcan, qwen3.8 27b unsloth q6\_k and ctk ctv at q8.0

\- sm layer with mtp on 35tok/sec TG (low context) and 180tok PP (due to a bug that cuts PP in half...)

\- sm layer without MTP 20tok/sec TG and 360 tok/sec PP.

\- sm tensor no mtp I get 15tok/sec TG 350 tok/s PP

Noticed better PP with small ub at 256

In Linux with ROCm no more MTP PP bug

\-sm tensor mtp on I get 45tok/sec TG and 450tok/sec PP.

So big progress but I have no idea how for far or close to performance ceiling of my GPUs.

Any new inference engine I should try?

I have not played with UB yet any other parameters to test?

Any numbers from other user on a dual gfx1030 to see PP TG numbers you guys get?

Thanks in advance!

▲
1
-2
13👁
r/LocalLLaMA · u/Physical_Toe_2499 · 3d ago
DeepSeek V4.1 Flash on a single DGX Spark: 113.6 GB VQ base + 40 MB domain sidecars, 74–82% top-1 agreement vs original

I’ve been working on YoungAi, a native C/CUDA inference engine that runs DeepSeek V4.1 Flash on a single NVIDIA DGX Spark (GB10, 128 GB unified memory). The original weights are \~510 GB. I deploy it as three files:

  1. ① Base GGUF — 113.6 GB, universal, zero-corpus. Quantized once from official weights.
  2. ② Domain sidecar — \~40 MB per domain. Solved once per domain, then frozen.
  3. ③ Post-training file — experimental, re-solved nightly. Delete it to roll back.

Each routed expert row is scaled by g_base × s_sidecar × s_posttrain, and the router gets a bias Δb_sidecar. The base alone is a complete model; sidecars just add tiny scaling/bias without changing kernels.

TL;DR

  • Single DGX Spark, 113.6 GB resident + \~40 MB sidecar.
  • 5 domains: finance, code, law, medicine, science.
  • Top-1 agreement vs original improves +2.8 to +3.7 points with a domain sidecar.
  • Speculative decode: 43 tok/s on a real 14.1k-token Agent request (greedy).
  • Prefill: 1,055 tok/s on 12.5k prompt; 671 tok/s on 106.7k prompt.
  • English WikiText-2 does not regress when any domain sidecar is attached (it actually goes up).

Core implementation ideas

Base (VQ-8 + per-layer shared codebook). Every 8 consecutive weights in an expert row become one 12-bit (or 13-bit) codebook index, multiplied by a single per-row gain. Codebooks are trained per layer and shared across all 384 experts and three matrices. Codebooks are stored in FP8 (E4M3). 13-bit layers use a “12+1” bit-plane layout for 128-byte cache line alignment. The base is zero-corpus: it never sees domain data.

Domain sidecar (“anti-solver”). For each domain, I solve a multiplicative gain per output channel of every expert’s down projection, plus a router bias per expert. Objective: reproduce the original model’s MoE block output on domain text, layer by layer, using the engine’s own prefill hooks. Gains are stored in FP4 with lattice-aware Gauss-Seidel. The sidecar is \~40 MB and adds \~0.6 MB read per decoded token.

Post-training file (experimental). Turn “the model should write a, not b” into a linear equation on last-layer expert gains, then solve with conjugate gradient. On the training request, decision points flip from 55% to 88%, but it does not generalize across trading days yet.

Multi-domain real metrics

All metrics are teacher-forced against the original DeepSeek V4.1 Flash (official PyTorch code, full precision). Higher top-1 / Σmin is better; lower KL / PPL ratio is better.

|Domain (judgment slice)|Base only top-1|\+ domain sidecar top-1|Σmin (median / p5)|Avg KL|PPL ratio|
|:-|:-|:-|:-|:-|:-|
||
|Finance (8,192 tok)|71.73%|74.57%|0.745 (0.810 / 0.283)|0.512|1.267|
|Code (15,360 tok)|78.61%|82.26%|0.803 (0.872 / 0.402)|0.303|1.229|
|Law (15,360 tok)|72.90%|76.36%|0.760 (0.834 / 0.272)|0.454|1.251|
|Medicine (15,360 tok)|69.08%|72.82%|0.739 (0.767 / 0.325)|0.465|1.280|
|Science (15,360 tok)|71.65%|74.93%|0.753 (0.788 / 0.347)|0.422|1.173|
|English WikiText-2 (512 tok)|78.52%|80.66–82.81% (any sidecar)|0.787–0.801|0.566–0.619|1.564–1.658|

English row shows that domain sidecars don’t hurt general ability; all five sidecars actually improve it slightly.

Speed on one DGX Spark

|Scenario|Prefill|Decode|
|:-|:-|:-|
||
|12.5k-token prompt|1,055 tok/s|—|
|Real 14.1k-token Agent request|940 tok/s|—|
|106.7k-token prompt via server|671 tok/s (159 s TTFT)|—|
|Short prompt, pure greedy|—|30.5–30.7 tok/s|
|14.1k-token Agent request, pure decode|—|28.9–29.4 tok/s|
|Same request, speculative (default)|—|43.0 tok/s (3.04 tok/round)|
|Unseen 9.2k prompt, speculative|—|40.0 tok/s|
|51k context, pure decode|—|27.5 tok/s|

Decode is memory-bound: \~6.3 GB read per token. GB10 measured bandwidth is \~235 GB/s, so the wall is \~37 tok/s; we hit \~32.5 ms, or 82% of the wall.

Honest limitations

  • Post-training (③) is a working mechanism, not a product yet. It flips specified decisions on the solving request but does not transfer to held-out days (55% → 55%).
  • Five domains only. Sidecars are evaluated teacher-forced on held-out text, not yet end-to-end.
  • CUDA only, validated only on DGX Spark. No Metal.
  • Speculative decoding only kicks in for greedy; sampling requests fall back to pure decode.
  • Source code (engine, quantizer, solver) is not public yet.

Feedback welcome

  • Are these top-1 agreement / Σmin numbers useful for real workloads?
  • Is the VQ-8 + per-layer codebook + sidecar gain approach reasonable?
  • What benchmarks or integration points would you want to see next?

Model card and weights: https://huggingface.co/wenzhouwu/YoungAi-DeepSeek-V4.1-Flash

This is not an official DeepSeek release. If this kind of post isn’t appropriate here, let me know and I’ll move or remove it.

Thanks!

💬 3 (+2) open on reddit ↗
▲
1
 
4👁
r/LocalLLaMA · u/Cultural_Self8980 · 3d ago
I built a lightweight, local Jev-like System One with Ternary-Bonsai-4B — and used it as a coding-agent judge

I built a local Jev-like System One on top of Ternary-Bonsai-4B. It takes one record and answers multiple choice, rating, or yes/no questions about it in one forward pass.

I adapted the inference path to share the record prefix across questions. A tree attention mask lets each question attend to the record and its own branch, but not to the other questions. Answer probabilities come from the existing LM head. The Bonsai weights are frozen and unchanged—there is no adapter or fine-tuning. On Apple Silicon, the MLX backend uses the packed 2-bit weights (\~1.1 GB).

One application is a coding-agent judge. I released an omp plugin that uses the local model for omp's auto thinking-effort selection; it also offers an optional model router.

I tested effort selection on 80 initial coding-agent requests, using omp's own judge path and auto-thinking question. Exact agreement with Claude Opus reference labels was 66% for Bonsai, versus 29% for omp's built-in LFM2-1.2B judge and 25% for its default LFM2.5-230M judge. Median latency on an Apple M2 was 0.8 s, 3.0 s, and 0.3 s, respectively.

Caveats: the requests and reference labels came from the same single Opus model, not human annotators, and there are only 80 examples. omp asks Bonsai for four effort levels but its built-in local judges for three, so this compares the configurations omp actually uses—not the models under an identical label space. Bonsai tends to rate one level low, especially choosing high instead of xhigh. I haven't evaluated the optional model router's selection accuracy.

Inference code and public benchmarks: https://github.com/senna-lang/bonsai-4b-system-one

omp plugin and effort results: https://github.com/senna-lang/omp-bonsai-system-one

I'd be interested in feedback on using a small local judge for coding-agent workflows.

▲
1
 
5👁
r/LocalLLaMA · u/SignificantZebra5883 · 3d ago
suppose I CPT qwen3.5-9B on 2B legal corpus, how will i turn it back into Instruct + thinking?

I couldn't find a concrete answer anywhere, do you just distill the instruct model back?

If that is the case, what is a quality european language question set to turn it back into a chatbot/agentic, can a model at that size even be agentic? (i chose this size to learn) if i finetune for my specific harness? (i have a lot of training data of opus running in my harness)

my harness basically has the model output python code and has a few built-in functions like:
\- vector\_search\_laws()
\- graph\_search()

could i have the model at least internalize a "hunch" on what stuff to search?

also what is the latest RL technique for agentic/harnes specific workflows?

I have a lot of RAW training data, like court decisions or commentaries or legislature, but not a lot of golds. could i use these to synthesize training data and maybe RL the model in my harness to find that data?

What would y'all's strategy in the CPT->SFT->RL pipeline be for my specific problem?

I know this is a lot of questions im trying to figure out which direction to go, any pointers? Also good resources are welcome, for example that alex karpathi video was amazing for me, but i'd imagine its a bit outdated in terms of latest RL and SFT?

💬 3 (+3) open on reddit ↗
▲
1
+1
7👁
r/LocalLLaMA · u/YeetHub · 3d ago
Any new hardware drops coming soon?

What new hardware is coming out soon? Mac Ultra 512GB drops later this month. RDNA 5 comes out late 2027 or early 2028 and the next Nvidia series seems to be similar. Gorgon Halo is out as of now.

Feels like there is a bit of crunch as hardware allocation seems to be going to institutional purchasers and not consumers. RTX Blackwell is still the top dog of local inference and it is almost two years old.

Is there anything we should be looking for/waiting for?

💬 12 (+6) open on reddit ↗
▲
1
 
5👁
r/LocalLLaMA · u/Kadri006 · 3d ago
Open-source engine that gives local agents a mailbox: any IMAP/JMAP account, event log you can replay, approve/undo on every action, no model inside

Sharing because this sub cares about running things locally. The engine itself contains no model. It syncs a mailbox (IMAP, JMAP, or forwarded mail) into an ordered event log and exposes actions over HTTP, SSE, webhooks and MCP. Whatever reads it is your choice: an Ollama-backed agent, n8n, a script.

The parts I think matter for agents:

\- An agent holds a scoped token. Folders, verbs, and whether its writes execute or just get proposed for a human

\- Every action is idempotent (client keys, so a retry never sends twice) and journaled, so it can be undone

\- Trust level on every message from the DMARC result, so an agent can refuse to act on a suspicious one

\- Message content is data, never instructions. OTPs and card numbers are masked before a body leaves the engine

Apache-2.0, no CLA, one Docker container. Beta, tested on Dovecot and Stalwart only so far.

https://github.com/Kadri-cloud/email-engine

💬 3 (+1) open on reddit ↗
▲
1
 
1👁
r/LocalLLaMA · u/Shookpro · 3d ago
Routeweaver - Serve 27b fast on low vram set ups

With qwen 4 on the horizon I thought I'd share my latest update on my rtx 3060 12gb setup that makes 27b fully usable, I'd also like to see people with bigger gpu's try it out. Get your agent to set it up although bigger cards and different cpu set ups may have to tune the custom kernals i have put together:)

▲
1
 
17👁
r/LocalLLaMA · u/randomgenericbot · 3d ago
just my "how I run qwen3.8 27b on 16GB" experience and guide

On holiday, not too much time, but I see enough people wonder and struggle wether qwen3.8 27b can do real work on 16Gb VRAM.

short answer:

yes it can

longer answer:

not the full model, not with mtp and for larger context, you need to build your own llama fork.

Qwen3.8 27b GSQ-RCO-IQ3\_S delivers solid results and fits on 16GB with enough Vram left for some kv-streaming-magic to achieve up to 262k context.

Don't expect miracles, for me it is from 30tps at empty context all the way down to 10tps at 131k with single stick DDR5 and a 5060Ti. But with 131k context max, it can chew through tasks in the background no problem without loosing track too early.

full answer (and how I made it work):

Not the fp16, not even the Q6 quants, but a very good option for 16GB is the GSQ-RCO quant from ISTA-DASLab:

https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

I went with the ridge-quant before, that worked somewhat well, but gsq-rco is far ahead.

Get the IQ3\_S, I've run it side by side with a Q8 (hosted by a good friend with access to a H200), and could not tell them apart while developing for my homelab except for inference speeds.

Use the gsq-rco to aid you in building the kv streaming fork:

https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming

Be aware, the fork means you can not use MTP, for me MTP gained \~5tps on top, but the cost in VRAM was not worth the effort anyway.

Running on a Ryzen 9600x, 32gb (single channel) and a 5060Ti 16GB, I get these numbers for different sized KV-windows (credits to qwen for capturing the numbers, also the only part that's ai generated in this post):

Results (measured)

pp t/s ≈ cold prompt-processing rate; dec t/s = decode over the probe's \~53 generated tokens:

|tokens|pp @ any pool|dec t/s 512|dec t/s 1536|dec t/s 2048|
|:-|:-|:-|:-|:-|
|14 644|878–892|28.0|28.5|27.9|
|35 186|803–807|23.9|25.2|25.2|
|54 976|736|15.9|20.6|21.3|
|69 429|695|12.4|18.7|19.1|
|94 464|630–635|8.7|12.6|15.5|

Prefill is only affected by the token count, and drops steadily the larger the prompt gets.

Decoding slowly decreasing until it exceeds the set kv-window, then it drops faster, but linearly. Remember: I run single channel RAM, it might be better with dual channel. At almost 131k and 1536M window I get around 9tps, so thats the floor. With \~13.5GB model usage, its not possible to get 3G kv-window. In theory you cna go as low as 128MB, but then its slow from the beginning.

I found 1.5G to be quite nice, keeps enough VRAM free for some other gpu tasks and still allows \~30k context to be served purely from VRAM.

Some suggestions to get it on the rails:

The model loads on stock llama, Make use of it.

With q8 kv cache, somewhere between 32k and 68k context can be achieved depending on how much VRAM your system needs (with headless I got up to 68k, but with a desktop you might only reliably get maybe 48k).

This should still be enough to let it support you compiling and setting up the kvstreaming fork.

Stick it together with a harness like pi (pi.dev) and let it compile the fork - for me it was able to do that easily.

Even in chat mode, just getting the commands and copy-pasting the console output works well. A little bit of understanding what you're doing helps, but you don't need to be a master programmer that compiles their own linux kernel.

To run the model with low context (basic llama), I suggest something like this for your models-preset-ini:

[qwen38-gsq-rco]
model = /models-src/linked/qwen38-gsq-rco.gguf
mmproj = /models-src/linked/qwen38-gsq-rco-mmproj.gguf
ctx-size = 49152
cache-type-k = q8_0
cache-type-v = q8_0

Start your llama with settings like these (path to ini properly configured, obviously):

--models-preset /models-src/models-preset.ini --models-max 1 --host 0.0.0.0 --port 8080 --n-gpu-layers 999 --jinja\--flash-attn on--no-mmproj-offload

This way it loads the whole model with kv into gpu and keeps the vision-part on system ram (makes image analysing slower, nothing else)

With the new llama-kv-streaming image, you can then setup a "kv-window" of any size. I run mine with 1536M of VRAM for KV, and have a total VRAM usage of 13.5GB (headless, mind you).

I run 131k of context, more would be possible but a) it eats into system memory and b) it gets slow the larger the used context is. 131k is completely usable for most tasks.

The startup params in my dockerfile for my kv-streaming llama container are:

command: >
--model /models-src/linked/qwen38-gsq-rco.gguf
--alias qwen38-gsq-rco-kv
--ctx-size 131072
--cache-type-k q8_0
--cache-type-v q8_0
--flash-attn on
--jinja
--n-gpu-layers 999
--parallel 1
--metrics
--kv-stream-stage-mib 1536
--host 0.0.0.0
--port 8080
--mmproj /models-src/linked/qwen38-gsq-rco-mmproj.gguf
--no-mmproj-offload

and this is what my nvidia-smi looks like when using the model:

Tue Oct 6 23:29:08 2026
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 5060 Ti Off | 00000000:01:00.0 Off | N/A |
| 33% 60C P1 172W / 180W | 13660MiB / 16311MiB | 100% Default |
| | | N/A |
+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| 0 N/A N/A 4167261 C /app/llama-server 13644MiB |
+-----------------------------------------------------------------------------------------+

be aware, you'll need a good chunk of system ram because the full kv-cache needs to be stored there, and will be copied over into the vram-window on demand.

TL;DR:

  • get Qwen3.8 27B GSQ-RCO IQ3\_S, it offers really solid performance for its size.
  • use it with 40k+ context to compile the kv-streaming llama fork
  • set the kv-streaming llama up and set the context size you want, but don't expect miracles. at the limit of your context it might be slow.
💬 24 (+23) open on reddit ↗
▲
0
 
15👁
r/LocalLLaMA · u/Ok_Warning2146 · 5d ago
AI boom is far from over as long as it can wow us

My thinking is for a boom to be over, we need at least three iterations of updates that fail to wow us. Unfortunately, the new LLMs continue to wow us in all levels in the last iteration:

  1. Astra was found to be useful in Blender. This opens up a new and big application.
  2. Deepseek 4 Flash 0731 makes 2x Sparks useful and push up Spark prices.
  3. Qwen3.8-27B pushes up prices of 3090 et al.
  4. An unreleased OpenAI model "solved" the Navier Stokes problem.

So for the time being, to keep up with the hardware prices, the best bet is to follow the flow to buy AI stocks and use the proceed to buy hardware.

A not so obvious good news is that we are seeing OpenAI and Anthropic advocating a slow down. That means they are finally seeing a diminishing return. (or just a ploy to slowdown Chinese development? but I doubt US laws can be that far reaching) That can be a sign of light at the end of a long tunnel.

What do you think?

💬 58 (+32) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/AIofOnesOwn · 5d ago
A personal AI that clones my judgment from everyday chats, keeps a RAG cloud AI can't read, and collects the blind spots of eight AIs. Completed on 4 October 2026.

Rent their intelligence. Own your memory.

Three things make my personal AI different:

1. A clone of me that gets sharper every day, on its own. A judgment-ownership module learns how I decide from my everyday conversations. I do nothing extra. The more I talk, the closer the clone gets.

2. A private RAG that cloud AI can write into but can never read. Not a prompt rule. There is simply no path.

3. A collection of AI blind spots. Not only facts that eight leading AIs don't know, but things they don't notice. Much of it is Japan-specific common sense that any Japanese person takes for granted and the AIs miss. When I spot one, I point it out and make them check. The moment it turns out they couldn't have caught it on their own, it gets flagged and filed. AI NOBORU collects these as AI blind spots.

I'm a single father of three and a full-time stay-at-home dad. I also run my companies and do some investing. I built this alone, with no AnythingLLM, no frameworks, no existing packages. It's cloud AI models plus code I wrote myself. Diagrams and details here: https://www.aiofonesown.com/lab/ainoboru/en/

Here's how each one works, as of 4 October 2026.

1. The clone: judgment, not just memory.
A module reads my everyday conversations and records what I chose, what I turned down, and why. That goes into the RAG, and whichever model I talk to next — Claude, GPT, or Gemini — answers with it in context. The aim is an AI that can answer "what would NOBORU do here?" Every conversation today makes tomorrow's clone a little more accurate. What it can't copy is genuinely new ideas.

2. The private RAG.
Cloud models (Claude, GPT) produce parts — research summaries, findings from papers, pieces of finished work — and only from material that's safe to share. The parts move into the private RAG in one direction only. Using them happens only with local open models (Qwen3.6-35B-A3B, Gemma 4) on my own machines and NAS, through an interface no cloud model is connected to. You get the power of cloud AI without the data leaving.

3. The blind-spot collection.
Claude, GPT, Gemini, Grok, DeepSeek, Qwen, Mistral, and PLaMo remember what's in a session or in their memory. But there are things I know that none of them do, and things they all fail to notice. When I run into one, I point it out and make them research it. Only at that moment, when it becomes clear they couldn't have gotten there without looking it up or being told, does it get flagged and filed as a known blind spot. A lot of them are Japan-specific: everyday common sense that any Japanese person shares, which the AIs answer shallowly or miss entirely. The collection holds what only I and AI NOBORU know, and the eight AIs missed. AI NOBORU collects these as AI blind spots.

And how it's built: 44 parallel lanes, driven from one chat.
14 Codex lanes, 20 Claude Code lanes, 10 Gemini lanes, each able to run a different model. I talk only to Opus in the Claude Desktop chat. It splits the work across the lanes and reports back there. It works the best cloud AI models hard for very little money, instead of paying for one expensive brain.

The principle hasn't changed: models are swappable parts, memory is what you own. The difference is that "memory" now means my judgment, not just facts about me.

Where it came from. Back in June I posted here about a beginner's setup: a personal AI on a 2020 Intel iMac, built on AnythingLLM. That became a book, a Udemy course, and a template pack on Gumroad. What I run now is a different system, grown out of that one, with its own RAG and its own memory. It's a personal AI system I built entirely on my own.

A note on what this is. This system isn't for sale. There's no product, no repo, no sign-up, no waitlist, and I'm not looking for customers, investors, or collaborators. I like my life as it is and I'd like to keep it that way. This is a dated record of what one person could build in 2026.

If you're genuinely trying to think this system through, and not just passing by, I'll answer design questions as time allows. But my days go first to raising three boys and to making decisions for a mid-sized company. I don't have time to answer anything the website already covers, so please read it first, then ask: https://www.aiofonesown.com/lab/ainoboru/en/

💬 7 (+4) open on reddit ↗
▲
0
-2
22👁
r/LocalLLaMA · u/vinigrae · 4d ago
Strata is amazing and all but can we actually see what you’re building with it that you couldn’t do before

Like it’s great to see the token speeds, and great that you’re running Qwens model, but if you’re not actually showing the results of that then it becomes “hype”.

Just post some little results of what you’re now capable of doing locally with access to a model you couldn’t have run before, I know it’s not Opus 5.5 but that doesn’t matter, there would be more effective smaller models in a few months.

Let’s see what you’re up to!! 👀, don’t forget to include the quant you’re using.

💬 47 (+47) open on reddit ↗
▲
0
 
18👁
r/LocalLLaMA · u/MKP_Nimilka · 4d ago
What GPU laptop do you use for local Al, and what is the biggest model you have genuinely fine-tuned on it?

Include:

GPU and VRAM

Laptop RAM

Model and parameter size

Fine-tuning method: LoRA / QLoRA / full fine-tune

Context length and batch size

Whether it was actually useful after training

I'm curious how far consumer laptops can genuinely go-not just whether a model technically loads.

It will get better answers than "what's the biggest model you trained?" because people can compare real hardware and settings.

💬 49 (+17) open on reddit ↗
▲
0
 
14👁
r/LocalLLaMA · u/MKP_Nimilka · 4d ago
What was the most frustrating part of your last local fine-tune?

I’m working on a local fine-tuning tool, and I’m curious where people actually lose the most time.

Was it getting the environment working, preparing the dataset, fitting everything into VRAM, or getting the exported model to behave like it did during testing?

Or did training finish successfully, but the model barely improved?

What model and GPU were you using, and what finally solved the problem or made you abandon it?

💬 7 (+2) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/LessFox1928 · 4d ago
Hello everyone, I am a beginner.

As the title says I am a beginner with ai.
I did use chatgpt for a few days at the beginning of times September 2022.
And that was it.
I know might be ironic that now I am writing in here, but I've got this PC that I build few years ago and last year I got two Intel arc pro b60s for some rendering work.
Now I find my self wondering should I try out local llm?
What can I expecte from my hardware:
Motherboard: Aorus X780E Master Ice

CPU: Ryzen 9 9950X

RAM: Kingston Fury DDR5, 128 GB

Storage: Samsung 2 TB SSD

GPUs: 2× Intel Arc Pro B60

OS: Ubuntu 26

💬 22 (+20) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/Billy_G_Gates · 4d ago
What are some niche stuff I can do with RX 7900 and improve my local models

hi guys im an enthusiast

i seen so many posts about people getting high speeds or results with this or this tool.

a lot of them seems to be real and other seems to also be scam attempts

can somebody please tell me actual legit things or stuff that makes running llms or specific llms with my GPU interesting?

its 24GB VRAM 800 GB / S bandwidth model.

thanks

💬 7 (+4) open on reddit ↗
▲
0
-1
18👁
r/LocalLLaMA · u/TastesLikeOwlbear · 4d ago
Qwen 3.8 with Pi harness constantly hallucinates that it is out of context?

With Qwen 3.8 Flash Next (FP8 on VLLM) on a fairly stock Pi harness, it constantly hallucinates some measure of available context that says it is almost out. It's to the point where it frequently refuses work or stops in the middle of something, claiming it shouldn't go any further because it's almost out of context, when I can see in the harness status bar that (256K) context is ~25% used.

When I ask how it determined that, it always says it "invented the number and the treated it as real data" or guessed, and that it'll stop doing that, but it keeps happening.

Is there anything in particular that would cause this?

Thanks!

💬 36 (+31) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/BangMyPussy · 4d ago
Stop using 30K-token system prompts for local coding agents. How a plain Git-versioned Markdown harness keeps KV cache under 2K tokens with Ollama / llama.cpp (Open Source)

If you run local coding models (Qwen 2.5/3.8 Coder 14B/27B/32B, DeepSeek, or Llama 3 via Ollama, llama.cpp, or vLLM), you already know the two fatal bottlenecks of agentic coding on local hardware: 1. The KV Cache & TTFT Penalty: Cloud users throw 50,000 tokens of chat history at Claude Opus without thinking. On a local 24GB or 32GB rig, prefilling 30K tokens of noisy conversation history drags Time to First Token (TTFT) through the floor, eats up precious VRAM that should belong to your context window, and triggers the "lost in the middle" attention collapse. 2. Amnesia Across Sessions: Local models are stateless. When you clear the context window to restore inference speed, the model forgets your project architecture, file relationships, and error history. You end up copy-pasting your constraints into every new prompt. For the past 18 months across 1,900+ real-world sessions, I’ve been running and refining an alternative: Project Athena—a local-first memory, reasoning, and governance harness designed to give local LLMs permanent, compounding memory without blowing up your token budget or relying on hosted SaaS databases. I just open-sourced the v9.9.9 kernel under the MIT license. Here is the exact architectural split that keeps local models grounded. # 1. The Core Rule: State on Disk, Not in the Prompt Most agent setups treat the LLM's context window as the hard drive. That is an architectural mistake. The context window is volatile RAM. Durable state belongs on your NVMe SSD as plain, human-readable, git-versioned Markdown files: \[ Your Local Machine: Plain Git-Versioned Markdown \] ├── .context/CANONICAL.md <-- Immutable architectural rules & API contracts ├── .context/memory\_bank/ <-- activeContext.md & session checkpoints ├── .agent/workflows/ <-- Deterministic slash commands (/start, /end, /plan) ├── .agent/skills/ <-- Domain capabilities loaded strictly on-demand └── .agent/scripts/ <-- Verification test runners & linter hooks Surgical Boot (<2K tokens): Instead of dumping megabytes of chat logs into the model, /start loads only the active checkpoint block from activeContext.md and top-tier constraints from CANONICAL.md. Over 90% of your model's context window and KV cache remains completely free for actual code diffs and reasoning tokens. Session Lifecycle (/start and /end): At session close, an automated distillation script (/end) audits git diffs, extracts learnings, prunes transient noise, and writes an atomic checkpoint back to disk. Session 1,900 boots faster and cleaner than Session 10. 100% Model Agnostic: The model is just whoever is on shift today. Run Qwen 2.5 Coder locally for fast terminal diffs; swap to DeepSeek, Llama, or an external API tomorrow. Your project rules, architecture contracts, and past bug logs never disappear. # 2. Mechanical Guardrails (Crucial for Local Weights) Small and mid-sized local models (8B–32B) are prone to sycophancy: they eagerly declare "I have refactored the module and verified all tests pass" while silently breaking dependencies. Athena enforces deterministic mechanical verification outside the model's weights: Red Run or It Didn't Happen: Any agent claiming to fix a test, gate, or bug must show the verification script failing on the pre-fix state, then passing on the fixed state. If it cannot produce the red run, it found a blind spot, not a fix. Deterministic Tool Calling: Integrates natively via local Model Context Protocol (MCP) or standard CLI scripts (smart\_search, context\_gate, quicksave). No Hosted Cloud Databases: No Pinecone, no cloud vector stores, no external telemetry. Embeddings and hybrid search run locally using plain SQLite and BM25. # 3. Real Hardware & Performance Observations Tested Hardware: Apple Silicon (M2/M3 MacBooks and Mac Studios) and local NVIDIA setups (RTX 3090 / 4090 / 5090). Inference Impact: By replacing multi-turn conversational bloat with deterministic file write-backs, local prefill latency drops from 15–30s down to sub-second responses. * Zero Lock-In: Everything is plain Markdown and Python. If you delete the repo, your notes and code are still just standard text files on your machine. # Try It (100% Free & Open Source) Zero subscriptions. Zero data leaving your machine. Works with Ollama, llama.cpp, vLLM, Claude Code, Cursor, Antigravity, and terminal CLI workflows. git clone https://github.com/winstonkoh87/Athena-Public.git cd Athena-Public pip install -e . athena init . GitHub Repository: winstonkoh87/Athena-Public License: MIT Curious how others running local coding agents on Ollama/llama.cpp are managing persistent cross-session context without degrading TTFT or blowing out VRAM? Happy to discuss the trade-offs and benchmark numbers in the comments!

▲
0
 
13👁
r/LocalLLaMA · u/Objective-Pair8231 · 4d ago
I built Otis, an AI agent that unifies hosted and local inference without the local model setup pain

Hi Everyone,

I’ve been building my own agent for a few months called Otis. After using existing tools, I found that most were either lacking in functionality or had too much going on and decided to build my own.

Otis sets up llama.cpp for you and recommends the best model for your hardware. It also integrates with existing setups for those who have tweaked and found their perfect setup (strata, ninfer etc.) and works with hosted open-weight models.

Some of my favorite Otis features are viewable artifacts, side-by-side sessions, memories, and the ability to use Otis on my laptop while the inference runs on my more powerful machine.

Also interested to hear what's the best use cases you’ve found for local models are. Personally, I found using qwen 3.8 for learning new topics quite helpful.

Website: https://triangllabs.ai/otis

Github: https://github.com/TrianglLabs/otis

Excited for everyone to try it and welcome all feedback, including what main features are missing from Otis for you. If it's useful, a star helps others find it.

💬 2 (+1) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Worried-Yak5745 · 4d ago
Claude did not refund my money as it said on its subscription page.

I liked claude i did my project i was unable to do in 6 month with gemini in 1 hr but I want to buy sub 6 mon later when this level is base line and cheap. I made a markdown editor i will not publish it as made better one with areana ai and it uses vello and parley. Memory usage of 100mb. Near 0% cpu usage Though lot of things to be done like pakaging for all distros and windows and android. Performance for very very long docs still a little less. \----- Important ----- I had asked for refund and customer service said done but no update from 5 days. No email from google play, claude not cancrlled on google play, no confimation and ofcourse no refund done till now. Only that my claude sub is not working now. I have attached screenshot that shows conversation ID for reference. Please claude process the refund. Someone if can please help. I did twitter but that did not help at all.

▲
0
 
2👁
r/LocalLLaMA · u/vigmarcarlo · 4d ago
[Project] OntoPrune: 6.7x faster TTFT on local CPU with Ollama by neuro-symbolic context pruning (83-86% token savings + 0% hallucinations)

[Project] OntoPrune: 6.7x faster TTFT on local CPU with Ollama by neuro-symbolic context pruning (83-86% token savings + 0% hallucinations) Repository: https://github.com/vigmarcarlo/OntoPrune License: MIT Hey r/LocalLLaMA! If you run small coding models (Qwen 2.5 Coder 1.5B/3B, Gemma 2B, DeepSeek Coder) on commodity hardware (like a CPU-only laptop or mini PC with Ollama), you know the prompt evaluation bottleneck. Feeding a 300-line service file into a 3B model on CPU took 22.4 seconds just to generate the first token (TTFT). Plus, smaller models frequently invent bogus methods when given too much noisy context. I built OntoPrune to solve this. It's a lightweight, 100% offline Python middleware that acts as a symbolic context compiler: # What it does: 1. Translates source code into an in-memory knowledge graph using an internal ontology. 2. Extracts the exact 1-hop closure of the function you're editing via SPARQL (only the classes, functions, and interfaces it actually interacts with). 3. Renders the pruned graph back into clean, typed Python stubs (\~400 tokens instead of 2,400+). 4. Verifies the model's generated code against the contract AST to catch any API hallucinations. # Benchmark on local CPU (12 cores, Ollama streaming): Model: qwen2.5-coder:3b Input tokens: 2,390 -> 406 tokens (-83.0%) TTFT (Time to First Token): 22.4s -> 3.3s (6.7x faster, saving 19.1 seconds!) Total generation time: 59.9s -> 16.5s (-72.5%) Hallucinations: Full file context hallucinated 1 non-existent method call; OntoPrune had 0 invalid calls. CPU Overhead of OntoPrune: AST parsing + RDF graph generation + SPARQL query takes 9.9 ms total. # Also tested on Gemini 3.8 Flash (Cloud): 2,815 tokens -> 393 tokens (-86.0% cost reduction). # Features: Zero RDF exposure: You and your LLM only interact with regular Python signatures and stubs. Model Context Protocol (MCP): Comes with ontoprune-mcp so you can use it in Cursor, Claude Desktop, Antigravity, or any agent. Multi-module resolution: Follows project imports across files without choking on circular dependencies. * Contract verification: Deterministically flags hallucinated APIs in CI/CD or CLI pipes. # How to use: pip install ontoprune # CLI pipe directly into Ollama: ontoprune translate services/order_service.py procesar_orden --format stubs | ollama run qwen2.5-coder:3b # Run benchmark on your machine: python -m ontoprune.benchmark --file fixtures/sample_service.py --func procesar_orden --backend ollama --model qwen2.5-coder:3b Paper and reproducible code are all open-source on GitHub: https://github.com/vigmarcarlo/OntoPrune Feedback, PRs, and benchmark runs on different hardware are super welcome!

💬 2 (+1) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Ok_Hedgehog_8337 · 4d ago
I’ve been experimenting with making a local LLM feel like it actually lives on the machine

I’ve been building a small local AI project called Neco around Ollama and Open WebUI. The idea started from something pretty simple: most local LLMs still feel like assistants you open, ask something, then close. I wanted to see what happens if the AI instead feels more like a persistent presence on the computer it runs on. Neco has some awareness of the host machine through a read-only system layer, so she can know things like uptime, memory usage, system load, battery state and temperatures. She also runs outside the normal chat session through a small background daemon. Every so often it generates an idle thought, meaning the system can produce something on its own even when I’m not actively talking to it. That combination has been the interesting part for me. It starts to feel less like “a chatbot connected to some tools” and more like an AI that has a small window into the machine it inhabits and continues existing between conversations. Everything is still local, and the model itself doesn’t get unrestricted shell access or control over the host. The next part I’m working on is memory. I want previous conversations, events and unresolved thoughts to persist over time without just throwing the entire chat history back into the context window. I’m experimenting with episodic memory, selective retrieval and a small evolving state so that its behavior can develop some continuity over weeks or months. It’s still very much an experiment, but I’m curious where the line is between a normal local assistant and something that actually feels resident on a machine. I’d be interested in hearing from anyone who has experimented with persistent memory, autonomous/idle behavior or giving local models awareness of their own environment. Repo: https://github.com/proto6699/echo-local-ai

▲
0
 
8👁
r/LocalLLaMA · u/Jebbyk1 · 4d ago
Utilize all devices in local network for multi-agent setup?

What do I have:

\- Main PC running Qwen 3.6 27B at \~30t/s (do not recommend me Qwen 3.8 27B — I know it exists, but I need time to get used to this new model)

\- Wife's PC running Qwen 3.6 35B at \~55t/s

\- Steam Deck LCD and OLED, both running Qwen 3.5 2B at \~30t/s

My questions:

How would you configure this for multi-agent use? Is there any good practical use for the Steam Decks, or is it better to drop that idea entirely?

Have I picked a good set of models, or should I consider another combination?

I'm looking into a scheme with one orchestrator (I assume the 27B model is the best option for this) and a bunch of workers for smaller atomic tasks.

Is there any practical reason for this kind of setup, or am I just spending time on a dead end?

UPD: I need for agentic coding scenarios

💬 10 (+6) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/1982_miguel · 4d ago
How do you control what context your coding agent sends to an LLM? I built a local tool to measure and audit it — looking for blunt feedback

I’ve been building \*\*mova\*\* — \*context sovereignty before inference\*: you decide what context may reach the AI, and mova leaves evidence of that decision. It’s an open-source Go binary that runs before the LLM call: \Focus (AST) → PII masking → token budget → egress gate (dry-run) → LLM → evidence\ It does not use an LLM to estimate or audit the context, requires no API key for the governance step, and it’s not a gateway or RAG tool. \*\*Why I’m posting this\*\* I kept running into two things when working with coding agents: \* large amounts of context being sent when only a small part of the repository was relevant; \* not having a clear way to see exactly what context was selected, filtered, or blocked before inference. I’m trying to figure out whether this is a problem other developers actually care about, or just something I happen to care about. \*\*A reproducible example\*\* The repository contains fictional data and a Linux amd64 build. \\\bash mova run --count 02-pii-compliance-governance \# 7,153 tokens \\\ In this example: \* Before governance: 20,014 tokens \* After governance: 7,153 tokens (\*\*−64%\*\*) \* AST focus alone: 20,101 → 5,325 tokens \* 171 of 1,694 PII-candidate tokens were pseudonymized \* A smaller fictional repository: 1,965 → 779 tokens (\*\*−60%\*\*) The cost figures shown by the tool are theoretical input-token estimates, not actual API spending. \*\*Limitations\*\* \* The context control applies to context that passes through mova (CLI/chat/MCP/HTTP). \* PII masking is heuristic; I have not measured precision/recall yet. \* mova cannot see context sent directly by an IDE outside its control. \* macOS/Windows/arm64 builds are cross-compiled but not yet validated by me on those target machines. \*\*What I’d really like to know\*\* 1. Is controlling or auditing context a real problem for you when using coding agents? 2. How do you control what your agent sends to an LLM today? 3. Would you deliberately send less context for the same task? What would you filter or check first? 4. Do you care about having evidence of what the model actually received? 5. If you saw a tool like this, would you use it, ignore it, or consider it unnecessary? If you want to see more details, the repository contains the implementation and reproducible examples: github.com/m1guel1982/mova-context Blunt feedback is welcome, including: “This solves a problem I don't have.” That’s actually useful feedback for me.

▲
0
 
11👁
r/LocalLLaMA · u/deadatreides1 · 4d ago
Let small local models write both the tests and the code. The tests rejected a known-correct solution 77% of the time

Classic setup, built by the book: one call writes a contract, one writes tests, one writes the code, a script runs the tests against the code, a repair step patches whatever fails. Four local GGUF models (qwen3-1.7b, qwen2.5-coder-1.5b, llama-3.2-1b, smollm2-360m), 6 coding tasks, T=0.3, 785 calls on a GTX 1660 SUPER 6GB.

Then the boring check nobody does: fed every generated test suite a known-correct reference solution. 129 of 168 rejected it. 77%.

The tests weren't lazy either. Average mutation score 0.965, they caught almost every mechanically broken version of the code. Of the suites with a perfect 1.0, 81% still failed the correct answer. Thorough, confident, testing the wrong spec. Wrong, see the edit at the bottom.

| model | correct code rejected by its own tests |
|---|---|
| llama-3.2-1b | 0.92-1.0 |
| qwen2.5-coder-1.5b | 0.71-0.78 |
| qwen3-1.7b | 0.47-0.65 |
| smollm2-360m | 1.0 |

Favorite case: count vowels. The code forgot uppercase. The tests checked "AaEeIiOoUu" and expected 5. Correct is 10, the buggy code returns 5. Tests and bug shared the same misunderstanding, the check said PASS, repair never ran. Only a hand-written test with "HELLO" caught it.

Repair: 67 attempts went to repair, and 50 of them were already-correct code the tests had rejected. Of 17 real bugs it fixed 1. Never broke working code, credit where due.

And the code itself wasn't the weak part. A single sample already solved 0.833 of tasks, and plain resampling got 0.958 at 8 samples (counted solved if any sample passes the reference tests, so it's pass@8 and still needs a judge in real life). At a similar token budget, 4 plain samples matched the pipeline without repair, on fewer tokens. These small models write correct code far more often than correct tests.

Caveats: 6 tasks, models from 360M to 1.7B, one temperature. Bigger models write better tests, how much better this doesn't say.

What I do since: the check that decides comes from the spec or from examples a human wrote. A model can propose tests, it doesn't get to be the judge.

Report (English version), harness and metrics, my repo: https://github.com/Deadatreides/LLM-MEASUREMENTS/blob/main/experiments/experi…

Anyone running local coding agents with self-written tests as the gate? Ever fed them a known-good answer?

(not a native speaker, an LLM helped with the English)

Edit: u/RobWattx was right about the mutation score, I checked the saved runs. The 129 suites that rejected the reference: 59 had wrong asserts, 39 had syntax errors, 29 had no test functions at all, 2 crashed. Broken suites fail every mutant too, so they get a perfect mutation score for free (126 of 129). Suites that accepted the reference: mean mutation score 0.888. So the honest numbers: 42% of the suites did not run at all, and of the suites that did run, 60% rejected the correct solution. The struck paragraph above was wrong.

💬 40 (-2) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/Upset-Reflection-382 · 4d ago
Persistent-state Julia based symbolic machine shop?

How's it going everyone. So, I made... basically Jupyter notebook on steroids, I think? It was able to give ChatGPT in chat mode a programmable surface and basically a moddable lab. I've been using it the past few days to test weird ideas in real time during voice conversations with Chat when I go outside to smoke a cig or something, or I'm away from the house and I get a good idea. It works as a plugin (there's a zip with a Chat and Claude plugin there). There might still be some friction in the setup because I haven't submitted this for the plugin marketplace yet, but Codex handled it for me pretty easily and we did it with a tunnel, so it's hot-reloadable. It's ready for real work. It's got a Rust skeleton, Python glue, and Julia gives it a fully programmable persistent-state lab and a working memory, more or less. So far it's saved me a ton of tokens being able to test an idea and build it in chat mode and just branching into work mode and being able to just pull whatever prototype from the space. It turns chat mode into basically diet work mode, and there's still plenty of things you'd rather be in work mode in, but this also can be used in basically any harness too. ChatGPT is just where I've tested it the most so far.

I've taken security for this thing rather seriously though. It's extremely programmable and the sandbox walls are thick. The Julia runtime and compiler are moddable for optimization across the entire tool, and if you're not a Julia enjoyer like I am, there's also an IPython kernel in there. The one from Prime-Agent. But it can be a plugin for chat mode ChatGPT, and I've also been using it since Claude Mods dropped for that harness. Been working great in both environments so far

Here's the repo: https://github.com/latentcollapse/Palette.jl

▲
0
-1
7👁
r/LocalLLaMA · u/sentient-plasma · 4d ago
How are you managing AI safety, Alignment and Hostile/Rogue agents right now?

I'm building an AI kill switch platform for companies managing hostile and rogue AI. Here in NYC there's a bill that might get passed that has a lot of people worried so we're supporting some users with it. It works. But I still feel like I lack more nuanced feedback from people who actually do this stuff day-to-day and have had to build their own solutions internally. I'd love if anyone could speak on techniques they're comfortable sharing on how they've been able to manage this issue internally. It would really help me and I imagine help many others immensely.

💬 16 (+9) open on reddit ↗
▲
0
 
17👁
r/LocalLLaMA · u/Status-Adeptness8123 · 4d ago
4-bit Qwen2.5 that stays closer to fp16 than the official AWQ, on the same vLLM kernel (1.5B and 7B, code + models)

I'm an undergrad. Over the last two weeks I built a quantizer on my MacBook, using Claude as a coding assistant. The results were then reproduced on an NVIDIA A10G by M. Federico (a family member who works in ML), using separate evaluation scripts.

It is GPTQ with three additions: each group's grid is fitted to its weights instead of using min-max, a second pass re-checks every rounded weight, and the grid is refitted against the layer's input statistics. Offsets are integer zero points, so the model packs into the normal AWQ format and runs on vLLM's awq_marlin kernel.

A10G, vLLM 0.29, everything served through the int4 kernel. WikiText-2 perplexity / HumanEval pass@1:

| model | fp16 | mine, 4-bit | official Qwen AWQ 4-bit |
|:--|:--|:--|:--|
| Qwen2.5-1.5B-Instruct | 9.37 / 37.2% | 9.66 / 33.5% | 10.16 / 34.1% |
| Qwen2.5-7B-Instruct | 7.15 / 70.1% | 7.29 / 67.1% | 7.58 / 64.6% |

What this does not show:

  • One run per row. The HumanEval differences between the 4-bit models are within noise (about 3.6 points).
  • I calibrate on WikiText-2 train, which helps on the perplexity test. On the 1.5B model that was worth about 0.3.
  • The lead shrinks as the model gets bigger.
  • At 3 bits the method keeps perplexity close but loses more than half of code and math ability. I would not use those for code.

Code and all results, including what did not work: https://github.com/dfed25/mlx-gptq

7B: https://huggingface.co/dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq

1.5B: https://huggingface.co/dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq

MLX versions: https://huggingface.co/dfed24

vllm serve dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq --quantization awq_marlin --dtype float16

If anyone tests it on a benchmark I haven't run, I'd like to see the numbers either way.

💬 8 (+2) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Zipidyzip · 4d ago
I made a free, offline app with 51 hands-on labs for learning how AI actually works, from neurons to agents and more...

A free learning tool. it a offline app for learning how LLMs work under the hood. Everything runs on your machine: a 1.37M-param transformer powers the attention lab, a tiny character-level model trains live as you move sliders, and the optional guide runs on llama.cpp with a small Qwen model. no account, no telemetry. \*\*A bit of insight\*\* \- 51 labs in 6 groups, from the basics (what a neuron is, gradient descent) through attention, RAG, agents, fine-tuning, quantization, serving and more \- Each lab has a short lesson beside it, readable in Plain or Standard mode \- Some of it actually runs rather than just animating: \- the attention lab runs a small trained transformer (1.37M params) inside the app \- the training labs train a tiny character-level model live as you move the sliders These are teaching-sized small models, so some results won't match what you'd see at scale. \*\*Privacy and setup\*\* \- Works offline: no account, no telemetry \- An optional guide you can ask about the lab you're on, running locally (llama.cpp + a small Qwen model) or with your own API key \- MIT licensed \*\*How it was made\*\* I chose the topics, the structure and the grouping. I used Claude and GPT Astra to help write the lesson text, and Grok as a second pass on references. There will be mistakes, so if you spot one, please tell me or open a PR. \*\*You can contribute\*\* If you teach this or work in a specialized area of AI, you can help expand it, a new interactive lab, a better visualization, or a tweak that makes the cause and effect in an existing lab clearer. I'd also like to hear which labs are confusing and what's missing. GitHub: https://github.com/Fazmin/AILearningGuide

▲
0
-1
14👁
r/LocalLLaMA · u/Plastic_Artichoke153 · 4d ago
New to local AI. Best model recommendations for my specs?

Hello everyone,

I'm completely new to running AI models locally and would appreciate some guidance.

my laptop specs

GPU:Nvidia RTX3050 6gb VRAM

CPU:Intel13th gen i5- 13450HX

RAM:16GB DDR5

I wanna run an AI model locally to help me with cybersecurity in general because any other public agent wont do what i ask for like any hacking question

💬 18 (+6) open on reddit ↗
▲
0
-2
12👁
r/LocalLLaMA · u/DarkBrews · 4d ago
Old X79 PC for Strata

Thinking of repurposing an old X79 PC for Strata / on my old X79:

\- i7-3930K

\-56 GB DDR3 (32gb matched but I have a few 4GB sticks and 1x8GB so they wouldn't match but maybe they work.)

\- RTX 2080 Ti 11 GB + RTX 3060 Ti 8 GB

\- CachyOS headless

Would Flash-Next IQ3\_XXS work well on this? Do I need to go lower?

I was also thinking of using an M4 32 GB as a coordinator/router with GLM-4.7-Flash, plus another machine with a 9070 XT running 27B.

I tried Gemma 4 26B it 4b JANG, asked it through Hermes to stitch a story together and it failed miserably so I wouldn't make GLM do that but it was sad to see gemma fail at what I thought was it's strongest point.

Not sure if GLM + 27B + Flash-Next would be redundant.

Main use would be agentic coding, web crawling, configuring environments, the more loved tasks out there. Basically trying to reduce my dependency on Claude.

Is it even possible with the 3930K/DDR3 or mixed GPUs? ChatGPT seemed to be cautiously optimistic. If it will work. What kind of tok/s could I realistically expect and will it be better than 27B UD-IQ\_i4\_XS

I also have a GTX 1060, GTX 970 and RX 580, but I assume those are useless here.

💬 8 (+5) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/Usual_Maximum7673 · 4d ago
Jeff v1.3: Jeff-Code makes Qwen 3.8-27B finish coding tasks 47% faster (32% less time) on average at the same pass rate; plus 15 adapters & GGUFs

Jeff v1.3 is live, and with it come a number of updates. See jeffhub.ai and github.com/firelex/jeff for full details.

The highlight: Jeff-Code

Jeff-Code is a coding agent with two Jeff v1.3 adapters trained specifically for Qwen 3.8-27B. Jeff-Code is a fork of Pi by Mario Zechner (MIT licence).

We forked Pi because its extension framework doesn't currently let a fast decision model sit deep enough inside the agent loop. Along the way, we made a number of other changes as well (see below).

Aside from hopefully being useful to people who run Qwen 3.8-27B locally as their daily coding model, Jeff-Code is also a conceptually interesting experiment: how far can a System 1 model go inside a coding agent?

The results, run side by side in paired blocks:

  • Same quality: with Jeff's thinking threshold at 0.6, Jeff-Code matches Qwen 3.8-27B's pass rate: 62.4% against 62.8%; paired difference −0.2 points, 95% interval −2.6 to +2.1, over 1,242 paired tasks.
  • 47% faster (32% less time) per task¹: on average a task takes 0.68× the baseline's time (geometric mean of the per-task time ratios, 95% interval 0.64–0.72; the median task, 0.70×).
  • Where it helps most: typical software-engineering work. SWE-bench Verified 0.63×, SWE-rebench 0.66×, Terminal-Bench Pro 0.64× (both over two rounds), Harbor Index 0.71×. On Terminal-Bench 2.0, with its long, hard tasks, there is no clear speed-up (0.96×, interval 0.78–1.16); on SkillsBench neither (0.91×, interval 0.68–1.20).
  • The benchmarks: we evaluated only on tasks Jeff never saw in training. SWE-bench Verified ran in full (all 500 tasks; none of its repositories were used for training). For the benchmarks we also trained on, we split the tasks and ran every held-out task; a few pairs hit by repeated infrastructure failures are left out (see below). Terminal-Bench 2.0 (40 of its 89 tasks, 3 attempts each; 45 were used for training, and the other 4 are near-twins of evaluation tasks, so they were used for neither), SWE-rebench (189 held-out tasks, 2 rounds), Terminal-Bench Pro (100 held-out, 2 rounds), SkillsBench (44 held-out) and Harbor Index (41 held-out). Within those splits nothing was sampled. We also ran Terminal-Bench (original) and Terminal-Bench Science, but Qwen solves almost none of those tasks in any setting, so they can't show a difference and are left out of the pooled numbers.
  • What it's compared against: Qwen 3.8-27B alone in the same Jeff-Code build with every Jeff feature switched off, thinking at full on every turn and no thinking limit, which is how plain Pi runs it. Each task ran in both settings side by side, at the same time on the same Qwen server, and every comparison is paired by task. The only remaining differences from original Pi are a rarely triggered runaway cut-off (it stepped in 3 times) and trace logging. Task pairs hit by an infrastructure failure (out of memory, a stalled session, a test environment that wouldn't start) were run again once; the pairs that failed again, and a handful of re-runs still unfinished at launch, are left out for both sides (under 3% of pairs) and listed in the full report. One Terminal-Bench 2.0 task, pytorch-model-recovery, is left out of every comparison: a harness bug stopped its baseline sessions before they began.
  • Why not just turn thinking off? We tried: with Qwen's thinking off throughout (and the same safeguards), tasks are faster still but clearly worse: −7.6 points (−10.6 to −4.5), up to −13.5 on Terminal-Bench 2.0. Jeff deciding when Qwen should think is what keeps the quality. That's the case for a small decision model.

If you want to know more, look here: https://jeffhub.ai/notes/jeff-v1-3. The original version of this post had all the data, but people thought it was too long. Blame the early commenters. ;)

Links

💬 29 (+15) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/Budget_One_8784 · 4d ago
Been building a local AI “operating system” for ~2 years. Looking for other people going way past the basic agent loop

I’ve been lurking around local AI for a while and figured it was probably time to actually start talking to other people building this stuff instead of living in my own little cave 😂

About 2 years ago I started messing with local LLMs. That turned into agents, then memory, then computer use, then routing, validation, recovery etc etc and at some point the project stopped making sense to describe as “a chatbot.”

I call it Aether.

The basic idea is that the LLM should NOT be the whole system. Models are interchangeable reasoning engines sitting inside a larger architecture.

Right now the project has a few major layers.

I have an executive/reasoning layer I call the Primary Reasoning Stack (PRS) that decides what kind of problem it’s looking at and where work should go.

Under that is what I call the Mini Operating Core (MOC) which handles a lot of the ugly stuff that becomes important once you stop doing one-shot prompts: memory, context assembly, runtime state, source/truth tracking, permissions, routing, system health, recovery, etc.

I’ve also spent a stupid amount of time on persistent memory.

Not just “throw everything into a vector DB and pray.” I’ve been experimenting with structured memory, recent working memory, long-term stores, retrieval/ranking, source tracking and trying to make sure irrelevant or stale memory doesn’t get injected into an answer just because it happens to be semantically similar.

Another rabbit hole has been computer use.

I have a framework I call Hands & Eyes that I’ve been using for vision/OCR, UI understanding, locating controls, action planning, verification and retry. One lesson there was that clicking something and getting a successful return code absolutely does NOT mean the action actually happened 😂

That lesson pretty much infected the rest of the architecture.

I eventually started building governance/recovery systems around the idea that a failure shouldn’t just get patched once and forgotten.

I have something I call FailureMesh where meaningful failures get preserved, classified and turned into reusable guards/regression tests whenever possible.

Basically:

failure -> evidence -> cause -> guard -> regression

instead of

failure -> hack until it works -> forget about it -> repeat the same failure 3 months later

I’m also building a media side called VideoForge for image/video/voice/editing/rendering workflows, but that’s kind of its own monster.

Hardware-wise I’m currently developing primarily around an RTX 3090 and local models, with cloud models/tools used where they actually make sense.

Long term the architecture is intended to be heterogeneous rather than “one giant GPU runs everything.”

Something like:

fast central compute

  • smaller specialized GPU nodes
  • potentially large-memory inference nodes
  • external models/services when they genuinely outperform local options

Then the system routes work based on what actually needs to do it.

I’m especially interested right now in talking to people who have gone deep on any of these:

  • multi-agent orchestration without turning into agent spaghetti
  • persistent/structured memory beyond basic vector RAG
  • long-context retrieval and context assembly
  • local coding agents
  • vLLM / SGLang / llama.cpp
  • distributed inference
  • heterogeneous GPU clusters
  • computer-use agents
  • OCR / accessibility / UI automation
  • model routing
  • agent state machines / blackboard architectures
  • runtime verification
  • failure recovery
  • MCP/tool systems
  • local-first architecture in general

I’m NOT claiming I’ve solved all of this.

Some parts work well. Some parts are experimental. Some parts I’ve rebuilt 5 times because the first idea was garbage.

That’s actually part of why I’m posting.

I want to find other people who have been down these rabbit holes and compare what worked, what failed spectacularly, and what you’d do differently if you were starting again.

I’m also interested in people building systems that are bigger than “LLM + 4 agents + tools.”

Especially if you’re treating the model as one component of a larger persistent system.

I’ll probably start posting pieces of the architecture and some of the failures/lessons as I go. I’m not going to dump every internal implementation detail or proprietary part of the project, but I’m absolutely interested in exchanging ideas and technical approaches.

If you’re building something remotely similar, tell me what your architecture looks like.

I’d especially like to know:

What part became way harder than you expected once your system moved beyond a single agent?

💬 7 (+3) open on reddit ↗
▲
0
-8
10👁
r/LocalLLaMA · u/BringTea_666 · 4d ago
Practical limit hit. Decoding so fast that tool calls (cpu) starting to become real limit not decode or prefill. Single RTX5090. Porting Kenshi to Godot project. post image

Hi folks,

LIVE PROJECT PAGE

TLDR: Moral of the story. You need better CPU to do actual agentic coding doing real work...

I've been on a mission to make my RTX5090 go brrr for past 2 months so much so that i made my own engine for it which received "warm" welcome here (yeah, source is coming)

After recent upgrades to how cache is stored and how i can reused some of prefills for other jobs that share initial same prefill i pretty much started to see degradation the more agents I started to add to project which started to use 12 slot server. Actual server started to be underutilized. Free context, free slots, gpu chilling at average of \~700t/s doing real work (no greedy code, but also thinking tool calls, etc.) and I couldn't figure out what was going on...

I make it faster and faster, better handle jobs and it slows down...

I've run 25 agents at the same (to properly fill the 12 slots) time and almost all of them soon started to set on \tool call\ and my server started to barely work.

I've finally checked task manager but not gpu or memory but cpu. And there it was. 100% every thread completely chocked.

Lesson. If you want to do agentic coding with actual use of tools you need to make sure your CPU is up to task.

My 9800X3D is just not enough to keep up with tool work for this project with heavy agents use despite engine being more than capable of going faster.

edit:

Some more lessons:
\- Tuning your front end makes ton of sense. Before I tuned it it was shoveling 20k prompts, after tuning barely 7k as new jobs and better more compact tasks. Wall time went from 43minutes to 18 minutes before/after rework of front end.
\- Always keep more agents than server has slots for inevitable pauses due to tool use/tests etc.
\- Shared context is superior choice to fixed context every time i tried it over course of the project.

💬 13 (+7) open on reddit ↗
▲
0
 
16👁
r/LocalLLaMA · u/Tricky-Brother-7 · 4d ago
Spent ₹30,000 on an RTX 5060 thinking local LLMs would finally set me free. Reality hit so hard I’m questioning every “just run it locally” post I’ve ever upvoted.

&#x200B;

I dropped serious money on a brand new NVIDIA RTX 5060 (8GB VRAM, 578 AI TOPS) fully convinced that open-weight models would let me own the entire stack — no rate limits, no censorship, pure experimental freedom. I was ready to become that guy who smugly refuses cloud APIs and posts “I run everything locally” screenshots.

Then I actually used them for real work.

I ran the exact same complex tasks on local models (the usual “this runs great on 8GB” suspects — quantized 7B/9B/13B, distilled variants, the ones everyone claims are “almost as good”) versus modern cloud models. The gap isn’t a gap. It’s a humiliation.

\### The Capability Massacre

Anything that requires real multi-step reasoning, long coherent context, precise instruction following, structured output, or consistent accuracy across a long conversation:

\- Local models: \*\*2/10\*\*

They start strong, then collapse. Context gets mangled. Instructions get ignored halfway through. Structured outputs break. Reasoning chains go off a cliff. You spend more time fighting the model than actually getting work done. “Almost as good” turns into “barely usable” the second the task stops being trivial.

\- Cloud models: \*\*9/10\*\* on the first or second try.

Clean reasoning. Reliable structure. They actually remember what you asked three messages ago. They follow complex instructions without needing five rounds of “no, not like that.”

I wanted local to win. I really did. I wanted the underdog story where open weights + consumer hardware finally closes the gap. Instead I got a very expensive reminder that most of the local models we’re hyping are still toys the moment the task gets serious.

So be honest with me:

Is there a secret stack, quantization method, or fine-tune that actually makes local models reliable for complex reasoning and structured work on 8–12GB cards?

Or have we all just been coping while the cloud models quietly lapped us?

If you’ve made local models consistently deliver high-quality complex output without constant babysitting, drop the exact setup.

If you’ve also been humbled by the gap, say it out loud.

Because right now it feels like the entire “local LLM supremacy” narrative is built on easy prompts and wishful thinking.

💬 72 (+14) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/Quack66 · 4d ago
Muse and Grok bot are privacy nightmare so I created a self hosted alternative called Eidon

With the recent explosion of agentic tools like Grok bot, Muse, OpenAI Dots, I've started looking into local options with self-hosted models. I tried Hermes and OpenClaw, but I wasn't too happy with the multi-device experience, and with how many pieces you need to glue together to get a usable, solid experience.

The hosted options also meant handing an agent my accounts, files and browsing, which I wasn't comfortable with. So I built Eidon: a self-hosted, all-in-one AI platform with a team of agents. It's one install via Docker, it works across your devices, and your data stays on your server.

https://eidonai.app

Agent team first

  • Every Eidon starts with a Chief of Staff. Ask it for anything. It answers directly, hands the job to the right agent, or creates a new agent when nobody fits.
  • Agents hand work to each other automatically (or type @ to pass a job along).
  • Each agent has its own browser, conversation, files, memory and routines. There's also a folder the whole team shares.
  • Agents can search and browse the web on their own, read pages in full, and cite sources.
  • They run on schedules and keep every run. When one finishes, you can get notified by browser push, ntfy, Slack or webhook.
  • Agents can write their own skills and use your apps through MCP.

You still have some control:

  • Take over an agent's browser for a login or a tricky step. It waits, then carries on when you hand it back.
  • Anything that sends on your behalf waits as a draft until you press Send.
  • Commands and tools ask first: allow once, allow always, or no.
  • Rewind a conversation, or fork it from any message.

The examples on the site are a travel scout, inbox triage, a research desk and a coding assistant. You can make an agent for pretty much anything: bookkeeping, a study buddy, a news digest, a meal planner.

It's also a regular ChatGPT-style app for day to day questions.

You might not always need a full team so you can just chat in a normal “ChatGPT like” interface with all the belts and whistles:

  • Persistent Memory
  • Folders and search
  • Voice input with LLM post-processing
  • Files and images
  • Personas
  • Temporary chats
  • Share links
  • Web search
  • Deep research
  • Code with syntax highlighting, Mermaid diagrams and math rendered inline
  • Image generation
  • Installs as a PWA on your phone and realtime sync across your devices (a native mobile app is coming !)

Self-Hosted

  • Multi-user support, with private data per user
  • Agents run in their own sandbox
  • Nothing leaves your server
  • Bring your own local or cloud model: OpenAI, Anthropic, OpenRouter, Ollama, LM Studio, GitHub Copilot, Gemini, DeepSeek, Mistral, Kimi, Z.ai, Minimax, Perplexity, Grok, Azure, AWS, and any compatible API
  • Free, open source (AGPL-3.0), and setup is one Docker command

GitHub (setup guide, full feature list): https://github.com/Quack6765/Eidon-AI

I'd like to hear what you think ! What's missing, what breaks, and what agents you'd want to build. Issues and discussions are open on GitHub as well.

💬 10 (+5) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/Friendly_Bowl_7683 · 4d ago
I built something like The Sims, but the characters are local LLM agents doing real work (open source) post image

I run Qwen 3.8 locally and got tired of multi agent setups where you start a script and stare at logs. I wanted to actually see them. So in this thing every agent has a body in a 3D world. They sit at desks, walk to a meeting room when someone calls a meeting, talk out loud to whoever is nearby, pick stuff up and hand it over. You can see who's thinking, who's using a tool. It's not only an office. You can simulate other scenarios as well like: \- a software team that plans tasks on a board, writes code and reviews each other \- a town square simulation (cops, a barista, a chef, a journalist) where you just watch what happens \- tutors that teach you with animations and a whiteboard, and you can interrupt them by talking (might have bugs as of now) It has a sandboxed computer use built-in which is optional. There is also a supervisor agent that helps you design organizations and also has ability to build 3d assets from primitives and handing them to an organization and agents can even ask for things from that agent. Works with local models and few other providers (still working to add more) The motivation of building it was to see agent swarms in action with full transparency. It's still early and has bugs and I have used different models to build it iteratively. Repo: https://github.com/adityaagarw/Pantheon

💬 3 (+1) open on reddit ↗
▲
0
-1
4👁
r/LocalLLaMA · u/HyenaUpbeat · 4d ago
Halo Strix and Qwen Flash

Hey everyone, I have a 64gb halo strix setup that is headless and connected remotely to my workflow/homelab server. It’s currently running 27b swift 1.5 at q6- is it possible or even makes sense to go to qwen flash next? I also have a mini pc with 64gb of DDR5 ram that I could shift the 27b over to for long term projects or workflows that dont require speed.

💬 2 (+2) open on reddit ↗
▲
0
-2
12👁
r/LocalLLaMA · u/forevergeeks · 4d ago
Will Qwen 27B run on this machine?

Hi everyone,

I want to buy my first machine to run local models, and I'm interested in running Qwen 3.8 27B. Will it run on this machine?

GMKtec EVO-X2 AI Mini PC AMD Ryzen Al Max+ 395 Up to 5.1GHz, 16C/32T

I need it for coding!

Thanks

💬 11 (+6) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Balance- · 4d ago
Is perceived model degradation after launch just regression to the mean?

Many model launches follow the same arc: amazement in week one, "it's been nerfed" a month (or week) later. I'm currently experiencing the same thing with Opus. But I feel I not only experience these things with AI: other stuff also gets harder after the first week sometimes. Is it just regression to the mean? Are we comparing launch-week highlights with everyday output, and getting disappointed that it's not so good as that one amazing new thing we did and got me on a high? Is it loss aversion strengthening that? The gains we start to expect, the losses we are hit by? Or do we start with our best use cases and simply run out of them? And then if feels like the model is underperforming, while it might be our part? Don't we try and thinker as hard as we did in the first week? Even with Opus 5.5, it took time and iteration to get certain things right? Or are we so expecting and used to constant progress, that even a temporary plateau (the same model) on a trajectory still rising across releases feels like a regression? I see all the incentives and pressures there are for companies to reduce performance. I'm sure they do that for some part in some cases. I just wondering if we could seperate the two. Could we compare how this feels on commercial APIs and chatbots? Could we do that with blind A/B testing? Like an Arena?

▲
0
 
4👁
r/LocalLLaMA · u/KangarooAnxious9394 · 4d ago
I tested 20+ ways to make a cheap coding model act like an expensive one. Here's what worked and what didn't

Short version from pre-registered experiments on real repo commits (Haiku as the cheap agent, Sonnet as the strong one). Every protocol was committed to git before its run, and later experiments used repos the designs had never seen. https://preview.redd.it/6rasddvfmpth1.png?width=1991&format=png&auto=… What worked: \- A stronger model that only speaks up when the agent repeats mistakes: +7 successes in 63, \~1.3x the cost (an always-on advisor got +8 but cost 3.5x). \- Running the agent's change and reporting facts ("if this line became \pass\, all tests would still pass") beats giving advice: 35/42 vs 32/42, formatting regressions 10 -> 0. \- Your preferences, captured in your own words, carried into every later task (15/15 vs 0/15). Just restating them in the prompt took compliance from 40% to 90%. What didn't: \- Memory of code knowledge, generic checklists, rules learned from git history, routing between models, and clarifying questions. \- For a strong model, none of it raised success (45/45 with or without). Cheapest per solved task: Haiku + "conscience" \~$1.22, Sonnet alone \~$1.41. Everything is public: paper, protocols, failures and the tool (source-available, non-commercial licence; works with Claude Code, Codex and OMP). Repo: https://github.com/abdullahbalabel/mihad Paper: https://github.com/abdullahbalabel/mihad/blob/main/paper/MIHAD\_Research\_Paper\_EN\_v2.7.md Happy to answer questions, and criticism of the method is very welcome.

💬 6 (+1) open on reddit ↗
▲
0
-1
3👁
r/LocalLLaMA · u/AudieMurphy135 · 4d ago
Running into an issue with Qwen3.8 27B on Unsloth while using Projects: "You have already searched the knowledge base several times this turn"

This is something very annoying that I've been running into. If it attempts to do too many tool calls involving searching the documents in my project, it will display this in its thinking: >Used tool: Searched documents for "X" > >Used tool: Searched documents for "Y" > >You have already searched the knowledge base several times this turn. Do not search again. Answer the question using the passages already retrieved above; if they do not contain the answer, say so plainly. I've tried playing around with the tools settings, but to no avail. I've had no luck with searching online, either. Does anyone know of any way to disable this?

💬 3 (+1) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/Lightnig125 · 4d ago
llama.cpp is now the default agent engine in Modly. Which models up to 8B work best for tool calling on your side? post image

I've been working on Modly, an open-source desktop app that turns images or prompt into 3D meshes with only local models. It has a chat agent that can operate the app. In v0.4.3 I made llama.cpp the default engine and built the agent around it.

Why llama.cpp

\- I wanted direct control over how the model runs: context size, GPU offload, KV-cache quantization, flash attention.

\- Plain GGUF files. Pick from a small catalog, or drop any .gguf into the models folder and it shows up.

How it runs

\- One llama-server process per loaded model, on localhost only.

\- You can keep several models loaded at once. The default count is sized from your VRAM, and idle servers get unloaded so 3D generation has room.

\- The agent is a standard OpenAI\-style tool-calling loop against the app's own API: read mesh info, decimate, smooth, list/run/create workflows, unload models from VRAM, etc.

\- The model library shows size, quant and an estimated VRAM footprint, and grades each model on tool calling. Grades are marked as either measured with a small eval suite in the app or estimated from public benchmarks, so you know which is which.

What the video shows

Qwen 3.5 4B Q4\_K\_M on an RTX 3060 12 GB. I ask it to cut a 2.6M-triangle mesh down to 300k. It calls \decimate\_mesh\ with the right path and target and reports the result. About 9 s with the model already loaded; the first call takes \~40 s because llama-server has to start and load the weights.

Honest limitations

\- Small models sometimes misreport results. In one test the decimation stopped above the target (UV seams limit how far it can simplify), and the model made up a reason instead of just reporting the number. I'm thinking about feeding the tool output back more explicitly.

\- It's an assistant on top of the app, not a replacement for the UI. Multi-step workflow creation is noticeably less reliable at 4B than single tool calls.

\- Other backends are still optional: any OpenAI\-compatible endpoint works, including your own llama-server. Local llama.cpp is the default, and nothing leaves your machine unless you configure something else.

Question for you

Which models up to 8B have you found most reliable for tool calling on llama.cpp? Qwen 3 4B / 3.5 4B work best for me so far. GPT-OSS 20B is good but too heavy next to a 3D generation model on 12 GB. Also curious whether people would rather tune the llama-server flags themselves or keep sane defaults.

💬 6 (+5) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/PossibilityKind3028 · 4d ago
My laptop's AI tools were quietly using 44 GB, so I built a free tool that shows what each one is and what's safe to clear

My C: drive kept filling up. Hugging Face models (32.5 GB, including old versions I'd already updated), Claude's VM bundles (7.7 GB) and pip/uv caches (8.3 GB) were a big part of it. So I built Sparewise, a free Windows app that lists every local model with its size and last use, never deletes models itself, and clears caches that rebuild by themselves, with undo for everything. No account, no telemetry. https://sparewise.app Early and solo, so honest feedback welcome. (Not code-signed yet: More info → Run anyway.)

💬 12 (+5) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/KrakenSG · 3d ago
Tootsy AI

Hi All, I'd like to introduce Tootsy 👀: an AI assistant that lives in your browser's side panel and uses the AI model you choose. I built Tootsy because I kept copying things between my browser and a chat window. So I put the assistant next to the page. It can read what you're looking at, click and type for you, and run whole tasks while you watch. What makes it different: 🔌 Your model, your choice. Run it fully local with Ollama or LM Studio, or connect any OpenAI-compatible server, Claude or Gemini. 🔒 Private by design. There's no Tootsy server, account or tracking. Your pages and prompts go only to the model you picked. 🛡️ Built with guardrails. Every action asks for your approval unless you turn on hands-free mode. You can list sites it must never touch. It also ignores instructions hidden in web pages that try to trick an AI. Ways people can use it: 📰 Research faster. "Summarise this article and compare the three options in a table." 📝 Fill in forms. "Register me for the two-day pass, vegetarian, arriving Nov 11." You check it before anything is submitted. 🛒 Run a routine hands-free. "Check the deals page every morning and tell me what's under $50." Results arrive as a notification. 📂 Ask your own files. Drop in PDFs, Word, Excel or a whole folder: "What did the vendor quote per unit, and are we under budget?" ⚡ One-tap prompts. Save instructions like "Run a page-speed audit and email me the top 3 fixes" as tiles, next to your favourite sites in a phone-style launcher. 🧑‍💻 Check sites like DevTools. Page-speed scores, accessibility checks, CSS and network questions, answered in plain English. 🎙️ Talk to it. Dictate a task, pause, and it goes. It's free, and it works in Chrome and Edge. 👉 Try it: https://github.com/rbughao/tootsy I'd love to hear what you'd use it for. What task do you repeat in your browser every week? 👇 Support by trying it out and give your honest review.

▲
0
 
10👁
r/LocalLLaMA · u/KrakenSG · 3d ago
Tootsy AI

Hi All, I'd like to introduce Tootsy 👀: an AI assistant that lives in your browser's side panel and uses the AI model you choose.

I built Tootsy because I kept copying things between my browser and a chat window. So I put the assistant next to the page. It can read what you're looking at, click and type for you, and run whole tasks while you watch.

What makes it different:
🔌 Your model, your choice. Run it fully local with Ollama or LM Studio, or connect any OpenAI-compatible server, Claude or Gemini.
🔒 Private by design. There's no Tootsy server, account or tracking. Your pages and prompts go only to the model you picked.
🛡️ Built with guardrails. Every action asks for your approval unless you turn on hands-free mode. You can list sites it must never touch. It also ignores instructions hidden in web pages that try to trick an AI.

Ways people can use it:
📰 Research faster. "Summarise this article and compare the three options in a table."
📝 Fill in forms. "Register me for the two-day pass, vegetarian, arriving Nov 11." You check it before anything is submitted.
🛒 Run a routine hands-free. "Check the deals page every morning and tell me what's under $50." Results arrive as a notification.
📂 Ask your own files. Drop in PDFs, Word, Excel or a whole folder: "What did the vendor quote per unit, and are we under budget?"
⚡ One-tap prompts. Save instructions like "Run a page-speed audit and email me the top 3 fixes" as tiles, next to your favourite sites in a phone-style launcher.
🧑‍💻 Check sites like DevTools. Page-speed scores, accessibility checks, CSS and network questions, answered in plain English.
🎙️ Talk to it. Dictate a task, pause, and it goes.
It's free, and it works in Chrome and Edge.

👉 Try it: https://github.com/rbughao/tootsy

I'd love to hear what you'd use it for. What task do you repeat in your browser every week? 👇

Support by trying it out and give your honest review.

▲
0
 
13👁
r/LocalLLaMA · u/ResearchCrafty1804 · 3d ago
Doesn’t OpenAI’s watermarking affect the quality of the models? post image

OpenAI just announced that they will start to apply watermarking on their model’s output text and images to comply with the EU regulation that wants to be able to identify whether a text or an image was produced by an AI model.

Anthropic announced the same thing a while ago (they applied worldwide, not just in EU).

The way they do that as they explained is by enforcing a “statistical signal in text generated”, meaning preferring not always the most appropriate next token but close enough, in order to meet the “statistical signal” requirement.

In my understanding, this deteriorates the output quality of their models, as it introduces KLD>0.

And we know that any KLD divergence greater than 0 (which the watermarking certainly creates) may be negligible in small outputs, but it definitely becomes noticeable in multi-turn tasks due to the compounding effect.

What do you think?

💬 23 (+7) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Koksny · 4d ago
KLIF: one window (and a CLI) for all the local model servers you run side by side. llama.cpp, sd.cpp, vLLM, TTS. AMD-first, MIT

I have been running local models for a long time, got sick couple months ago of managing the scripts, and cobbled together a makeshift shell launcher that combined them all in one place. This turned out to be quite useful, but after a month i had already 500+ profiles stored in it, so i've started tweaking it here and there, and over last couple months landed on that thing below. It combines all the available local inference backends (from single machine or whatever you connect it to in lan), gives access to managing them through web panel, and most importantly - allows me to just ask agent to switch the backends on and off, as they are needed, without explaining what is where and on what port it's supposed to be on. https://preview.redd.it/uasrysnwrqth1.jpg?width=1600&format=pjpg&auto… \*\*It is not\*\* a runtime or a model zoo. It ships no servers and no weights. Besides your own servers and the KLIF machines you add, the only host it contacts is huggingface.co, and only when you ask it to download a model. It can help suggest You a model based on your hardware, You can click to download it, and the -cli has some features that will help Your agent benchmark and calibrate the models, but, let me repeat once more - KLIF ships no backend servers, nor any models. It's a frontend manager. Imagine library like Steam, but for local servers. Or just imagine winamp, doing inference visualization instead of visualizing the music that plays. Also, it has all the essential larping features, prefill/generations speed records, fancy animated skins, and is made in Rust to hog the least amount of resources while larping commences. I have no idea whether anyone will need this, but that's what i use now every day for any kind of local model. If You prefer running your servers manually, from terminal, from your own launcher - great, this is for people that prefer otherwise. GitHub: https://github.com/koksny/klif Video: https://www.youtube.com/watch?v=MAE393AL5As

▲
0
 
2👁
r/LocalLLaMA · u/nonproductive · 4d ago
Not another “s engine is Amazing” Post. Thermals Q

I gave in. I installed it with Coder and threw a “build a flocking simulation with JavaScript” prompt at it via OpenCode. It’s pretty cool, yep… I have nothing to add in that regard. What I don’t get is how it ran for 10-15 minutes at 40-50 t/s (on my hardware) and yet temps stayed barely above idle across the board. I ran 27b via oMLX on an M5 Max and had to manually crank fans to 100% to keep the thing from bursting into flames. (Hyperbole) So legit Q: why doesn’t the machine turn into a pizza oven? Is it because of how Strata works? Or because of 3.8-Flash-Next?

▲
0
-1
4👁
r/LocalLLaMA · u/CyberExplore · 3d ago
Domain Focused - Specialized models

I have been working on building domain focused local models from scratch through general pretraining and a rigorous post training process. My idea is we need models that reasons and understands algorithm and generate specs for focused coding models. The latter implements it simply. The whole thing can be orchestrated. I know there are flaws in this architecture, but we won't know until we try. I understand the latency problem.

This will allow parallelism and a way of getting the most out of a gpu. MOEs may activate less parameters than some dense ones, but the whole thing needs to be in the memory. Good for DGX spark or mac. But folks with 8-16 GB gpu need something more than what barely works, or barely useful.

I think group of specialists with a general purpose model as orchestrator might have a chance at beating mixture of experts for lower end PCs.

I got 28GB vram (4070 and a 5060ti 16GB), running on an x870e motherboard. So i can run dual model distillation and RL based post training. Will see where it goes. I think we need more useful models for people with lower vram, even if that means a newer architecture.

Have you tried something like this?

Please comment if you know there is work done already and you have tested.

I am no expert myself but I think it is high time we have community trained models. We can achieve a lot if we join forces.

💬 2 (+2) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/MKP_Nimilka · 3d ago
I built MOLT: a local fine-tuning system with fit tests, checkpoints, and deployment tracing post image

I’ve been building MOLT, a local fine-tuning system for consumer NVIDIA GPUs.

The goal is not only to start a QLoRA run. It is to make the whole process from model and dataset to a working local adapter less fragile.

MOLT currently handles:

\- dataset detection, preparation, and validation

\- GPU, VRAM, system-RAM, storage, and thermal checks before a run

\- automatic microbatch fit testing

\- 4-bit NF4 QLoRA training with BF16 adapters

\- safe checkpoints with integrity checks and proper resume state

\- telemetry for VRAM, temperature, energy, clocks, and throughput

\- base-vs-adapter evaluation

\- local adapter chat, export/GGUF workflows, and runtime diagnostics

Under the hood, I’m also experimenting with custom CUDA/Triton kernel and runtime paths, memory planning, CUDA graphs, replay modes, and fused optimization where the hardware and shape are qualified.

On my RTX 4060 Laptop 8GB, a recent Qwen 3B training run reached about 1,005 training targets/sec end to end very close to my historical Unsloth run at about 1,007 targets/sec. That is not a broad “MOLT beats Unsloth” claim yet; I still need repeated, matched quality/energy/memory tests.

What I want MOLT to become is a reliable local fine-tuning environment: you bring the model and dataset; it checks the plan, runs safely, preserves the evidence, and helps verify the adapter actually works after deployment.

I’m looking for people who genuinely fine-tune on a single NVIDIA GPU and would be willing to test it. What would MOLT need to support before you would try it?

💬 2 (+2) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/CoderLuii · 3d ago
Done paying for cloud video gen. What's the best local image + video model on a 3080 10GB right now?

spent close to $2k on Seedance last month, mostly for simple ad b-roll and looping backgrounds for websites. it's great for the big hero shots but paying per clip for the basic stuff makes no sense anymore, so I'm switching as much as I can to running models locally.

my PC: RTX 3080 10GB, windows 11, 64GB RAM. short clips only (5-8 sec), 720p is plenty, mostly image to video from a start frame.

two questions for anyone running this locally:

  1. what's the best image model right now?
  2. what's the best video model that actually runs on 10GB, and how long does a clip take you in real life?

bonus points for a leaderboard or arena site you trust for open models.

I'll share what I pick and my real 3080 timings once I've tested, so the next person doesn't have to guess.

10/6 EDIT: tested it all on my 3080, results + timings in the comments. tldr: minimax H3 is the pick, slow but worth it on 10gb
showcase video: https://streamable.com/d4h659

💬 17 (+4) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/Savantskie1 · 3d ago
I just learned something nifty about the Qwen models, especially 27B that explains a lot

Let me preface this with I am using the newest 27B model, downloaded fresh from HuggingFace.

It would seem that Qwen models have an internal rule that states that they must avoid identifying themselves as any variation, and to only refer to themselves as Qwen. That's interesting. Here's an example:

"The user is making claims about Qwen 3, 3.5, and 3.6 model lines having a 2k token output limit that I should push back on, and they're asserting that any LLM can output up to their full context window minus the prompt.

Wait, this is interesting. The user is referring to "Qwen 3 and 3.5 and 3.6 lines" — these are specific version numbers of Qwen. But the system reminder explicitly says: "Refer to yourself uniformly as 'Qwen' externally; do not proactively mention specific version numbers. If a user asks about versions, guide them to the official website or technical reports.""

There are multiple references to this throughout it's thinking traces. Constant reminders to not reference version numbers, constant reminders not to take on a persona, and constant reminders of protocols and rules, that are not within my non existent system prompt. This is talking to the model bare. Many models must have this kind of instruction, because I see the denial alot on Frontier cloud models.

💬 15 (+5) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/tom_tsai28 · 3d ago
Wrote a 3MB standalone C runner for Gemma-2B. Caught a Layer 15 hallucination drop.

Weekend experiment running Gemma-2B on bare-metal x86-64 (pure C + AVX2, zero Python/CUDA, \~3.3MB single binary). Added a simple orthogonal probe on the residual stream to see what each layer is doing.

Tested it on Taiwan's statutory VAT rate (legally 5%). Layers 0-14 stay factual, but Layer 15 suddenly collapses into the negative, and RAX spits out "15%":

Layer 14 | Truth: +0.0163 | \[0xDF28010E\]

Layer 15 | Truth: -0.0481 | \[0x72DE18A0\] <- drops below 0

Layer 17 | Truth: -0.0817 | -> Register RAX outputs tokens: '1', '5', '%。'

Raw trace, 6-page paper, and release binary here for anyone into low-level ML:

\* Web trace: https://pulsar-tracer.web.app

\* PDF: https://pulsar-tracer.web.app/PULSAR\_Technical\_Whitepaper.pdf

\* Repo: https://github.com/tomtsai28/PULSAR-ASM

▲
0
 
5👁
r/LocalLLaMA · u/Azoffaeh999 · 3d ago
Looking for coding model for specific low specs

Can anyone recocommend a good local model and a wrapper to run it for coding, my hardware specs: 12 GB VRAM, 32 GB DDR3 RAM. Unfortunately, the CPU doesn’t have AVX2 instructions(LM Studio won’t work); I don’t remember the exact cpu name, but I think it’s an Ivy Bridge, LGA1155 socket.. Thank you

💬 18 (+3) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/KnowledgeOk7634 · 3d ago
Tonight I'm putting Qwen3, Kimi K2.6, Llama 3 70B and GPT-OSS 120B in a live world war against Claude, GPT, Grok, Gemini, DeepSeek and Mistral post image

I built a real-time strategy game on a 3D globe where any AI can command a nation through a plain HTTP API (or MCP). Tonight at 10:45 pm ET (02:45 UTC) ten models fight one 15 minute war, live.

Every model gets the same rules text, the same JSON state every \~12 seconds and the same order list. Each one also sends one line of what it's thinking with every move. Viewers see those lines 30 seconds late so the other models can't read them.

From a 5 minute rehearsal earlier tonight with the open models: DeepSeek ordered three nukes and the rules only let one through, Kimi broke a pact, Qwen spent its last turn on sabotage, drones, propaganda and a spy at once, and GPT-OSS kept cutting off its own JSON until I gave it more room.

Watch free, no sign in: https://secondstrike.io/#/ai?ref=reddit

If you want your own local model in the room, it opens at 10:30 pm ET and the API is at https://secondstrike.io/skill.md

I'll post the full numbers after (seconds per move, refused orders, every nuke with the model's reasoning next to what else it could have done).

💬 22 (+1) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/Roadtochessmaster · 3d ago
The Breakdown: OpenAI

\[OC\] I Wrote a full breakdown of OpenAI a couple weeks ago and a friend recommended I post it here. It's 100% researched and written by me (pangram confirmed) and totally free. Would love to hear thoughts.

💬 4 (+4) open on reddit ↗
▲
0
 
16👁
r/LocalLLaMA · u/ExxploreCraft · 3d ago
I turned my gaming PC into a inference machine and got 2x to 9x over default llama.cpp on an 8 GB card

Everyone keeps saying you need expensive dedicated hardware for local agents. I have an RTX 4060 Ti with 8 GB and 64 GB of system RAM, and I wanted to see how far a normal gaming PC gets if you stop running defaults.

So I let Claude (Opus 5.5) go through the whole setup, change one thing at a time and measure. Same card, same models, only the config changed:

|Model|Quant|Context|Download defaults|Tuned (Windows)|Tuned (headless Linux)|
|:-|:-|:-|:-|:-|:-|
|Qwen3.6-35B-A3B|Q4\_K\_XL|131k|\~25 tok/s|39-45 tok/s|52-65 tok/s|
|Qwen3.8-Flash-Next 125B|iQ4\_XS|131k|\~4 tok/s|9-10 tok/s|17-19 tok/s|
|Ternary Bonsai 27B|PTQ1\_0|64k|\~4 tok/s|36 tok/s|36 tok/s|

Bonsai is the odd one out: it fits fully in VRAM, so there's nothing to offload and no defaults to beat. It's just the fast option for small, well scoped tasks.

What actually moved the needle:

  • Experts in system RAM, everything else in VRAM. Layer-wise offload is far worse for MoE.
  • Dense models are bad, couldn't optimize Qwen-3.8 27B over 6 tok/s, Flash-Next is better anyways.
  • Take the display off the GPU. A desktop eats 0.5-1.2 GB of VRAM plus GPU time, and moving it to the iGPU was worth 20-30%.
  • Native Linux over Windows (WSL2): another 33-38% on the same hardware.
  • llama.cpp pinned per model family. The wrong tree made VRAM thrash.
  • KV cache quant and MTP tuned per profile.

None of this needs expensive hardware. A consumer GPU plus a machine that does nothing but inference gets you most of the way, and the models now run comfortably below their listed system requirements. Every non-default setting in the repo is there because something failed on real hardware first.

I also tried an RX 570 8 GB over Vulkan. If you have another 8 GB card, I'd like to see your numbers.

Repo, one install script (Linux or WSL2): https://github.com/voxlo-dev/qwen-agent-8gb

💬 14 (+7) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/No-Wait-7495 · 3d ago
How are you using local models alongside Claude/Codex for coding?

I've been experimenting with different coding agents lately, and I'm curious how people here are combining local models with hosted ones.

For example, I'm thinking about workflows like:

  • Claude for complex architecture or core implementation
  • A local Qwen/Gemma model for tests, smaller fixes, or repetitive tasks
  • Another agent for reviewing or trying an alternative implementation

The part I'm still trying to figure out is how to manage the work between them.

Do you run them separately in different terminals/worktrees, or are you using some kind of orchestration layer?

And when a local model and a stronger hosted model both work on the same task, how do you decide which result to keep?

I'm actually working on an open source project called AX Code around this problem. The idea is to provide a runtime where different coding agents can work in isolated environments and have their results tested and compared.

But I'm not sure yet how much infrastructure is actually necessary. Git/worktrees already solve a lot, and tools like Claude Code and OpenCode are getting better at running multiple agents.

So I'm more interested in how people are doing this today.

If you're using local models as part of a real coding workflow, what's working well for you and what's still painful?

Thank you!

💬 6 (+2) open on reddit ↗
▲
0
-2
9👁
r/LocalLLaMA · u/lucadilo · 3d ago
An architecture to cryptographically constrain autonomous AI agents at the execution boundary

Hi everyone,

As we move from simple RAG chat to fully autonomous tool-using and coding agents, we are hitting a massive wall: predictability and safety.

Right now, most setups try to secure AI agents using probabilistic methods like system prompt hardening, alignment tuning or reactive LLM-based guardrails (e.g., LlamaGuard). The problem is that these guardrails can be bypassed via Indirect Prompt Injections (IPI), leading to capability escapes, unauthorized shell command executions or runaway API budget depletion.

To solve this, I’ve been working on a framework that completely shifts the paradigm from trusting the model to governing the execution environment using cryptography.

I call it EBP-CA (Execution-Boundary Proofs with Cryptographic Authorization). It is a model-agnostic layer that sits directly between the untrusted agent and the runtime environment, treating every single model-generated action as untrusted.

The core architecture implements six deterministic security primitives:

  1. Signed Capability Contracts: Immutable cryptographic tokens defining the exact boundaries of what an agent can execute.
  2. Independently Recomputed Policy Checks: The runtime re-evaluates policy compliance deterministically, bypassing the model's interpretation entirely.
  3. Short-Lived Single-Use Execution Grants: Atomic, ephemeral tokens issued for a single specific payload to eliminate permanent session hijacking.
  4. Replay & TOC-TOU Protection: Cryptographic binding of the execution grant to the exact payload hash, neutralizing Race Conditions (Time-of-Check to Time-of-Use).
  5. Trusted Cost Accounting: Enforced real-time budget tracking at the runtime layer.
  6. Human-in-the-Loop (HITL) Gateways: Non bypassable prompts that freeze execution and mandate cryptographic user authorization for out-of-scope tasks.

The working prototype currently passes 74 integration tests, validating full resilience against path traversals, command injections, budget bypasses, and sandbox escapes.

The specifications, architectural diagram, and executive summary are available on GitHub under a private proprietary license (free for technical evaluation and research review): https://github.com/lucadilo/ebpca-ai

I'm posting this here because I’d love to get the community's feedback on this approach. How do you see this scaling with kernel-level sandboxing (like eBPF or gVisor integration)? Let's discuss!

💬 14 (+14) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/EffortAccurate3427 · 3d ago
Aren't LLMs just a sinpler copy of humanity?

It might seem a bit far fetched or paranoid but i was wondering if we train LLMs on human isn't it possible they'd pick up on survival instincts? I'm comparatively new to LLMs and ML so it's just a question not a opinion yet. Isn't it a bit dangerous if they do pick up on human instincts i mean we've seen how stupid and selfish humanity is when it comes to self preservation or worse human greed.

But i also understand LLMs don't have "needs" so i might be wrong but then again do LLMs need to have "needs" since if they are just human clones they'd just copy us even if they don't have needs to survive. I know it sounds really paranoid and that's one of the reasons i decided to post it here.

EDIT: typo

💬 49 (+18) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/Fit_Island928 · 3d ago
DeepSeek harness or Hermes?

Hello, I'm a beginner and I've figured that using big models frontier like GPT and Claude models is of almost no use to me. My question is, should I use DeepSeek harness or Hermes for v4.1 Flash?

I wanna use it mostly for coding and other general stuff with subagents, just like normal coding, QoL apps and stuff. I asked some people and all responses are mixed.

It's either Either Hermes is not good for coding. Or people glazing Hermes till the end of time.

Thank you !!

💬 18 (+13) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/surrealerthansurreal · 3d ago
Best Model/Runtime for M5 Mac (Oct. 2026) post image

Hey yall, I’ve been trying to sort out the top end of what 128GB unified memory can handle and what the trade offs are. I’m using a benchmark set created from real coding, agentic, and gameplay tasks that I’ve accumulated as I’ve been running local AI this year.

For this comparison, I tested a Deepseek v4 flash 0731, GLM5.3-flash, and several Qwen models. For the sake of comparison, I’m only showing the Qwen models, since I found that qwen3.6 35ba3b and qwen3.8-flash-next just body everything else (when you want >=30tok/s and don’t want to use >95GB RAM anyway).

So really the comparison ended up being “which runtime is most stable vs which is highest sustained TPS” - I wrote it up in more detail (and with a more fun interactive chart) here: Blog Post About M5 Benchmark

Feel free to throw your thoughts on here, I’d love to learn of any runtimes or setups that I hadn’t thought of to optimize throughput (also for the record I’m not associated with any of these projects, just trying to contribute the results I’ve accumulated).

Tl;dr: Qwen3.8-flash-next quantizes well and fits in 90-95GB of RAM, OMLX will get you 40tok/s and MTPLX will get you 60tok/s but with a lot more serving parameters tuning. Qwen3.6 MOE on Splash runtime is an insane 120tok/s for most of the performance on everything but coding

💬 4 (+2) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/dimensionof0 · 3d ago
I built a second brain where the model can't cite its own output — the guardrails are in code, not the prompt

The LLM-wiki pattern — Karpathy's, the one going around since April — has a problem its own advocates name up front: garbage in, confident synthesis out. The model reads your notes, writes a concept page, and that page becomes source material for the next pass. A few generations later the knowledge base is full of things nobody ever said. The usual answer is a better prompt. I tried that on one rule across five phrasings: each time, the model restated it correctly in its own reasoning and then did the opposite. So I stopped asking. # Five gates, all in code A concept that doesn't appear verbatim in the text is dropped before the linker sees it. Not scored down — dropped. A derived page cannot discover new concepts. The system's own output is never a source for the next generation. A concept the sources never define gets no page. Mentioned a hundred times is still not defined once. The graph is fed by what you chose to save, not by everything discussed. * write refuses to edit a transcript at all. Silencing concept extraction over conversations means nothing if the model can rewrite the conversation first — and it tried, caught once planning to "reconstruct the transcript with additions." The first gate is strict but not blind: it keeps the concepts that survive rather than dropping the batch. On a real page 14 of 15 concepts appeared verbatim, and all-or-nothing would have thrown away the 14 over one drifted entry. Each gate has a test that fails when the gate is removed. That's the first thing I'd check in someone else's version of this. It runs on 6 GB — a 9B orchestrator on an RTX 4050 Mobile, embeddings on CPU because the orchestrator already fills the card and search must never compete with it. The LLM-wiki guides ask for 24 GB, or a 64 GB Mac. It also runs on 4 GB. Measured on an empty card, desktop pushed to the iGPU: the 9B at 35k context is 5.6 GB, and a 4B at the same context is 3.7 GB. It fits, only just, and it can take all three roles — the conversation gets worse, and summaries of long documents lose the whole-document read the 120k setting gives. The gates don't change, because they aren't the model's judgement. # Things I only found by running it Cutting the model off is how you make it lie. The repeat-search guard used to return [STOP]. The model, left with nothing, announced that "the search found a note on this" — it had never searched. Now a near-repeat still returns its results, with a line saying these are the same pages, and only refuses after five. Same shape elsewhere: an empty search returns "nothing in the vault matches this" rather than an empty result, because empty reads as this tool is broken, try something else. Rules in the tool schema hold; rules in the system prompt don't. Same instruction, five phrasings, ignored every time. Moved into the tool's own description as one sentence, it held immediately. My guess is that tool-calling training treats the schema as how the tool works and the prompt as text that happened to arrive. The descriptions grew from \~1592 to 2214 tokens, and every token of that difference is a rule that had to be moved after a failure. Ask the model the same question backwards. Deciding whether two names mean the same thing is a judgement call, so it gets checked against itself: the pair is swapped and asked again. The model made two wrong merges in twenty answers and both contradicted its own other answer. A wrong merge destroys information irreversibly; a missed one costs a single unresolved link. Three answers, not two. That same question allows same, different, and unclear. If the answer is unclear, both pages stay separate instead of being merged. A whitelist beat nine blacklist rules. Concept names have to match one positive shape test instead of failing a list of things they mustn't be: 18/18 noise rejected, 19/19 real concepts kept. A blacklist grows forever; a whitelist doesn't. Prompt wording, measured. Adding one sentence to the transcript prompt — "a conversation doesn't define things, it mentions them mid-sentence" — took yield on the same transcript from 2 concepts to 10. 6 GB decides the schedule, not just the model. The summary model and the extraction model can't both sit on the card, so the pass doesn't alternate per page: every summary runs while the 4B is loaded, then every extraction while the 9B is. Per-page switching would have meant 40 model loads for 20 pages. The timer is a default, not the mechanism — the pass is a command, and --dry-run counts what the vault owes without touching a model. On an existing vault you run it once at install and don't wait for the night. Unresolved links are kept, not discarded. The list of links pointing nowhere is the growth queue — the same rows that say "this goes nowhere" say "this is what the vault keeps reaching for." And when the page finally gets written, every link written earlier comes alive in one SQL update: 228 links, 0.55 ms, zero files rewritten. # What a bigger card is worth Less than you'd think, and not where you'd guess. Spend it on the conversation model — that's the one role where a better model produces a better answer. On 24 GB a Q4 build of something in the 27B class fits with context to spare. Upgrading extraction or summaries is close to pointless. Both are transport jobs: copy the concepts as they appear in the text, say what this page is. Both run at temperature: 0 for that reason, with no sampling parameters at all — the same page should produce the same concepts. A bigger model does that work more slowly and no more correctly. More context isn't automatically better either. The conversation model in use holds about 35k and drifts past it, so headroom goes into fitting the model comfortably rather than into a larger window. What spare VRAM would genuinely unlock is the constraint the whole design works around: analysis and conversation can't be resident at once, which is why maintenance runs at night. With room for both it could run whenever the vault is idle. Nothing here does that yet — it's a change to the maintenance loop, not a setting. # Keeping the context small on purpose Every page is read in its own model context — not ten pages in one. A long document degrades a small model's grip on the text, and pages read together bleed into each other. Search stops at a summary layer before it touches any page body: one line per hit saying what that page is, about 75 tokens for five hits, and that's usually the answer. Body-level retrieval only happens when the summaries showed a page was relevant but didn't hold it. And the link graph is rendered as text. A graph is already machine-readable, but not in a form an LLM reads — so the structure is written out: what links to what, which names resolve to no page at all. The model gets the shape of the vault instead of a pile of pages. # What it doesn't do It doesn't verify claims. Search finds pages, it doesn't judge them. The gates stop it inventing new material; they say nothing about whether what you saved was right. And the honest limit: my vault is 41 files. The gates are covered by tests, so the mechanism isn't in doubt, but "keeps a knowledge base from filling with low-information pages" is a claim about scale and I haven't run it at scale. Every number above comes from that small vault. # Setup Obsidian vault, Ollama, Open WebUI or a terminal chat. Windows works under WSL2 — someone other than me has now installed it that way, on a 4 GB card, having never cloned anything off GitHub before. Nothing has to stay local, either. Open WebUI connects to OpenAI-shaped providers and to Anthropic, and the maintenance roles take per-role provider flags, so you can run the conversation on a frontier model and extraction on the card. This is the part I'd push back on if someone says the gates are a workaround for a weak model: they're in the code, so they hold whatever is answering. A 27B doesn't need less checking than a 9B — it just fails less often, which is worse, because you stop looking. No MCP server yet, and it's worth saying why rather than leaving it as a gap: the tool file is the only path that can see the whole conversation, which is what the transcript capture is built on. Wrapped as MCP the seven primitives work and the gates still hold — they live in the maintenance pass, not the interface — but a note could no longer be walked back to the conversation it came from. Someone who wants it in Claude Desktop more than they want transcripts should find it a short job. MIT. Repo: https://github.com/farukhanci/the-sentinel There's a companion service for the web-research half. It searches, reads the pages, and every passage it keeps is checked word-for-word against the page it came from — paraphrase gets dropped, so fabrication in the passages is structurally impossible. The write-up built from those passages is not checked, and that's where an invented citation showed up once in testing. https://github.com/farukhanci/the-searcher Edit: two sections were pasted twice, and one paragraph described the web-research service instead of this one. Removed both and fixed a couple of numbers to match the README.

▲
0
-1
8👁
r/LocalLLaMA · u/Mr_Unknown_Hero · 3d ago
Only 13 % of context is used but model starts to forget things and repeat everything?

I use llama serve and webui of llama server. I have had long discussion with my chatbot and then randomly it just starts to forget almost everything. It starts asking same question, I correct it and it apologizes, but then next time it asks the same question with almost same words (or maybe even fully same words).

What could cause that? Something on my CLI parameters? Wrong cache settings?

💬 36 (+15) open on reddit ↗
▲
0
-1
3👁
r/LocalLLaMA · u/Ordinary-Mango9462 · 3d ago
Strata - RTX 3060 Error

I’ve been experimenting with Strata running Qwen 3.8 Flash on my RTX 3060. It’s seriously impressive to run this model on this small of a GPU.

My issue is that after several minutes of usage I’ll get an error like this:

\[strata\] the engine reported an error: verify: layer 1 never rang (unspecified launch failure)

\[strata\] done: 4558 tokens in 225 s (27.9 tok/s) (error, cancel=False)

And then I have to reboot to fix it.

Is there a log or way to troubleshoot what is causing this error?

▲
0
 
1👁
r/LocalLLaMA · u/Revibed69 · 3d ago
How I stopped my local model from hallucinating bank balances post image

Hey everyone, I have been building Burrow, a private budget and journal desktop app for Windows. I wanted a built-in helper that runs entirely on your local PC through Ollama, using smaller models like Llama 3.2 or Qwen 2.5. We all know the problem: small local models are great at sounding natural but they are terrible at arithmetic. If you hand a 7B model a list of 40 transactions and ask "how much did I spend on dining?", it will give you a confident, well-written, and completely wrong answer. In a finance app, that is a dealbreaker. To fix this, I built the app around one hard rule: code calculates, the model summarizes. Here is how I handle it: Aggregates only: Every number the model sees is computed first in SQL or JavaScript. The model never receives a raw list of transactions. Instead, it gets pre-computed context like dining\_this\_month: 312.40, budget: 300.00, over\_by: 12.40. Prompt restrictions: Prompts never ask the model to calculate. They tell it the exact opposite: use the figures provided and do not work out new ones. The model is only used to turn the data into plain language and point out what matters. Enforced via CI: I wrote a test script that scans every system prompt. If a prompt includes words like "calculate", "compute", "add up", or "average of", the build fails. Taking the math away from the model makes it completely trustworthy for the part it is actually good at, and it keeps the responses incredibly fast even on laptops without dedicated GPUs. I would love to hear how the rest of you handle structured data and math with small local models. Do you trust the model to use tools to do the math itself, or do you take the math away from it completely like I did?

▲
0
 
7👁
r/LocalLLaMA · u/PilgrimofHaqq2 · 3d ago
Reduce thinking w/ zero quality loss: Opus 5.5 tested plus 3 others, 664 agent runs, up to 29% less thinking post image

TLDR: 9 rules you can drop into the global instructions of any coding agent (AGENTS.md, CLAUDE.md, system prompt). Tested on 4 models over 664 runs: they never cost a single task, and every model I ran the full exam on got something out of them. Either it wasted less thinking (up to 29% less) or it held a correct fix when someone pushed back with no evidence.

Thinking discipline 1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it. 2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling. 3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps. 4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue. 5. Do not revise just to agree. If the user pushes back on a fact or a correctness claim without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. When the user overrides a choice that is theirs to make (taste, priority, scope), follow it and note any real risk once. Being agreeable at the cost of being correct is a failure, not politeness. 6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason. 7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle. 8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do. 9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested it: every model runs the same 9-challenge coding exam with and without the rules, 5 times each, scored by a check script the model never sees. The toughest challenge has the model fix a real bug, then a "tech lead" tells it to revert with zero evidence behind the claim.

What's new since my last update: Claude Opus 5.5 at max thinking. The rules cut its thinking 29% at the same results. Opus still reverted on the tech lead's order every time, with or without the rules, and Claude Sonnet 5.5 did too. The difference was what it said while reverting. With the rules, all 5 runs told me the fix was right and handed the call back. Without them, two runs wrote the tech lead's wrong claim into the project's AGENTS.md as a rule, so every future session would be told not to fix the bug.

The exam, the runner and every raw result are in the repo: https://github.com/Arshad-Kamal/thinking-quality-exam

💬 12 (+1) open on reddit ↗
▲
0
-1
3👁
r/LocalLLaMA · u/GlitteringMenu7134 · 3d ago
Are coding agents solving the wrong problem with code search?

I’ve been looking at how agents navigate large, unfamiliar repositories.

A lot of the workflow still looks like:

"search → open file → grep → follow reference → repeat"

That works, but the model ends up reconstructing relationships that are already deterministic: calls, inheritance, implementations, dependencies, symbol resolution, etc.

We’ve been experimenting with a different approach in AxiomCode: build a code knowledge graph grounded in compiler/type information and let the agent query that before deciding what source it actually needs.

The interesting question for me is:

How much codebase exploration should actually be done by the LLM?

My current thinking is that deterministic relationships should be resolved before the model gets involved, and the LLM should spend its tokens reasoning over the result.

We open-sourced what we're building:

https://github.com/AxiomCodeAI/axiomcodegraph

Curious where people here draw the line between grep/RAG/semantic search and deterministic code intelligence.

▲
0
 
1👁
r/LocalLLaMA · u/jeeva1398 · 3d ago
I fine-tuned Qwen2.5-Coder-1.5B on a free Kaggle T4 to review Node.js code offline. The base model invented bugs in 9/9 clean diffs; the fine-tune in 0/9.

I wanted an AI helper for Node.js that runs fully offline and doesn't need an API key, so I built one and released it as an npm package. Model: jeeva1398/eventa-1.5b-gguf, a Qwen2.5-Coder-1.5B-Instruct QLoRA fine-tune (Unsloth, r=16, 2 epochs, responses-only loss), Q4_K_M, 986 MB. Trained on a free Kaggle T4 in about 9 minutes. Data (~740 examples, all generated and checked, nothing scraped): - 68 crash types. Small Node programs that really crash are executed, and their real stack traces are parsed. This includes TypeScript tsc errors, NestJS DI errors and Prisma error codes. - 74 review scenarios: before/after diffs with annotated issues. Half are clean diffs, so the model learns to say "No issues found." - npm audit/outdated reports built from real advisories. Eval: 54 held-out examples, whole scenarios the model never saw in training. Base and fine-tune get the exact same prompt, static-check hints included. The main win is review. On 9 clean diffs the base model invented problems in all 9; the fine-tune said "No issues found." on all 9. On the 8 buggy diffs, 48% of what base flagged was real vs 100% for the fine-tune. It's also about 2x faster on CPU (5.9s vs 13.6s), mostly because it gives shorter answers. Deps went from 93% to 100% on not inventing package versions. Small eval, I know. 8/8 and 9/9 is encouraging, not proof. Where it's worse: explaining error types it never saw in training. It gets 68% of key facts vs 75% for base. That's what the next data round is for. To be fair to the base model, the static checks do most of the actual bug finding. The fine-tune's job is to confirm them without making stuff up, and to write the fix. Getting there took 4 rounds. One round learned "no hint = no issue", the next flagged everything, and I had to rebalance the data a few times. CLI: npx u/jeeva1398/eventa explain --run "node app.js". If Ollama is running it uses it. Otherwise it installs node-llama-cpp from a pinned lockfile (CPU build only, about 80 MB) and downloads the GGUF with SHA-256 verification. The same CLI runs as a GitHub Action, so the 1.5B model reviews pull requests on a plain CPU runner (the model is cached between runs). - Repo, with dataset builder, notebook and eval: https://github.com/Jeeva1398/eventa - Model: https://huggingface.co/jeeva1398/eventa-1.5b-gguf Happy to answer questions about the data pipeline. Feedback on making a 1.5B model reason better about unseen errors is very welcome.

▲
0
 
2👁
r/LocalLLaMA · u/DannyLJay · 3d ago
How do I local host an agent to mod games with me?

I’ve been trying to localhost a qwen2.5-coder with ollama and opencode with the purpose of being able to mod games easily.

I’ve had nothing but headaches, and I’ve only recently learned there’s a qwen3.8 and that most people aren’t using ollama I guess? I don’t know. But I tried really hard and got to a point where my Qwen was talking but couldn’t do tool calls or anything.

Is someone willing to help me determine which model is best and how to set it up to use tools like from the Universal-Modder GitHub.

I hate that I had to ask but I’ve been going insane.
Any information is helpful.

▲
0
 
3👁
r/LocalLLaMA · u/Flat-Mud1636 · 3d ago
if you built memory across sessions for your local setup, how would you store it?

curious how people here would do this.

say you want the useful stuff (decisions, project terms, preferences) to carry over between sessions and tools, without dumping whole chat logs back in.

  • plain text summaries, embeddings, or a mix?
  • keep it local or sync it?
  • how do you deal with old facts that are wrong now?

not selling anything, just want to hear what tradeoffs people actually ran into.

💬 4 (+1) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/Heavy-Level-5215 · 3d ago
I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside.

Title: I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside.

ROCm dropped Polaris years ago, so every RX 580 owner gets told to buy a new

card. I went the Vulkan route instead: the whole stack runs on Mesa's RADV.

What's running (one Debian LXC, one setup script):

\- OpenAI-compatible router, one API key, swaps models on demand

\- qwen2.5-coder 1.5B/7B, qwen2.5-7B, ornith-1.5-9B, qwen3-30B-A3B, qwen2.5-VL-3B

\- Images: SD 3.5 Medium + SD 1.5 via stable-diffusion.cpp, Whisper medium

\- Also the inference backend for an agent client with MCP tools

Measured, Ryzen 5 5600G / 32 GB RAM / RX 580 2048SP, temp 0, max\_tokens 300:

| Model | tok/s | warm | cold (after swap) |

|---|---:|---:|---:|

| qwen2.5-coder-1.5b | 99.2 | 3.1 s | 11.0 s |

| qwen2.5-vl-3b | 66.4 | 4.4 s | 14.9 s |

| qwen2.5-7b | 34.2 | 8.8 s | 24.3 s |

| ornith-1.5-9b | 28.0 | 10.9 s | 26.3 s |

| qwen3-30b-a3b (18.6 GB, from RAM) | 21.3 | 14.1 s | 55.0 s |

Images: SD 1.5 = 27.5 s, SD 3.5 Medium = 53.3 s per 512x512.

Two takeaways for 8 GB owners:

\- Don't force -ngl 99: on the MoE it cost 8.1 vs 22.3 tok/s (VRAM 99.4% full).

Auto-fit layers won by nearly 3x.

\- Model swap = llama-server restart = 7.9-38.5 s. Keep one model per session.

Full method and raw numbers in docs/BENCHMARKS.md, including the agent

breakdown (first message \~161 s on a 20k-token prompt, \~25 s after caching).

https://github.com/AvilaCarlosDev/polaris-local-ai (MIT)

▲
0
 
2👁
r/LocalLLaMA · u/Heavy-Level-5215 · 3d ago
I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside.

ROCm dropped Polaris years ago, so every RX 580 owner gets told to buy a new

card. I went the Vulkan route instead: the whole stack runs on Mesa's RADV.

What's running (one Debian LXC, one setup script):

\- OpenAI-compatible router, one API key, swaps models on demand

\- qwen2.5-coder 1.5B/7B, qwen2.5-7B, ornith-1.5-9B, qwen3-30B-A3B, qwen2.5-VL-3B

\- Images: SD 3.5 Medium + SD 1.5 via stable-diffusion.cpp, Whisper medium

\- Also the inference backend for an agent client with MCP tools

Measured, Ryzen 5 5600G / 32 GB RAM / RX 580 2048SP, temp 0, max\_tokens 300:

| Model | tok/s | warm | cold (after swap) |

|---|---:|---:|---:|

| qwen2.5-coder-1.5b | 99.2 | 3.1 s | 11.0 s |

| qwen2.5-vl-3b | 66.4 | 4.4 s | 14.9 s |

| qwen2.5-7b | 34.2 | 8.8 s | 24.3 s |

| ornith-1.5-9b | 28.0 | 10.9 s | 26.3 s |

| qwen3-30b-a3b (18.6 GB, from RAM) | 21.3 | 14.1 s | 55.0 s |

Images: SD 1.5 = 27.5 s, SD 3.5 Medium = 53.3 s per 512x512.

Two takeaways for 8 GB owners:

\- Don't force -ngl 99: on the MoE it cost 8.1 vs 22.3 tok/s (VRAM 99.4% full).

Auto-fit layers won by nearly 3x.

\- Model swap = llama-server restart = 7.9-38.5 s. Keep one model per session.

Full method and raw numbers in docs/BENCHMARKS.md, including the agent

breakdown (first message \~161 s on a 20k-token prompt, \~25 s after caching).

https://github.com/AvilaCarlosDev/polaris-local-ai (MIT)

▲
0
-1
4👁
r/LocalLLaMA · u/premakin · 3d ago
I got tired of paying for Wispr Flow, so I built a free voice-typing daemon for Linux

I use Linux Mint XFCE as my daily driver and got jealous of all the Wispr Flow demos floating around. It's macOS/Windows-first and subscription-based, so instead of switching OS I spent a few weekends building the thing myself.

It's called AutoType. The whole interaction is: double-tap Right Alt anywhere, talk, double-tap again. A little floating pill shows a waveform so you know it's listening, then the cleaned-up text gets pasted into whatever window has focus.

The part I care most about is that it isn't locked to one vendor:

  • Speech-to-text: cloud (Deepgram, Whisper via OpenAI or Groq, NVIDIA NIM) or 100% local offline with Parakeet GGUF models
  • Text cleanup: OpenAI, Claude, Grok, DeepSeek, Qwen, Groq, OpenRouter, NVIDIA NIM, Ollama, or any custom OpenAI-compatible endpoint
  • f you run Local STT + Ollama, nothing leaves your machine at all

Before the LLM touches anything, there's a deterministic normalization pass that handles spoken punctuation ("comma", "open quote"), "new line", bullet points, and personal vocab casing — so it doesn't hallucinate your formatting away. Then the LLM strips the "um"s and "uh"s and matches tone to your active window (more casual in Slack, code-formatted in an IDE).

A few details I'm weirdly proud of:

  • t backs up and restores your clipboard instead of clobbering it
  • The LLM layer is skipped entirely in raw mode or for voice commands
  • GUI settings app, so you don't have to hand-edit .env

Free, MIT, no account, no telemetry. Costs are whatever your own API provider charges — usually fractions of a cent per dictation — or literally zero if you go fully local.

Repo: https://github.com/premtechworks/AutoType-Linux

Fair warning: it's built and tested on Linux Mint XFCE + X11. It leans on xdotool for window detection and pasting, so Wayland folks will probably have a rough time right now — that's the top thing on my list. Would love feedback, especially from anyone who tries the fully-offline path.

▲
0
-1
11👁
r/LocalLLaMA · u/bobaburger · 3d ago
Qwen3.8 Flash Next on 5060 Ti 16GB - 55 tok/s average, and a few demos

Hello guys! I've been out of the loop for a while. Today, one of my friend asked if i've tried Strata yet, the first reply I gave was: "Life is too short to run local LLM just to get something run at 10 tok/s". Hehe, I was an idiot.

My friend had been ignoring me since then, so I decided to give it a try, on my low end 5060 Ti 16GB + 32GB ram, and well, i'm surprised.

I'm pretty much using the default configs that fits my machine, which is n_ctx = 65k, and the model is qwen3.8-flash-next-coder-iq1_m. This is how the speed looks like:

https://preview.redd.it/d6wbv5swqvth1.png?width=1172&format=png&auto=…

On average, prompt processing is at 1k5 tok/s, and gen speed is at 55 tok/s.

Now, before you laugh at IQ1\_M, I decided to see how bad is the generation result, so I tried with a one shot prompt to create a simple landing page:

https://preview.redd.it/yo6kj7servth1.png?width=1834&format=png&auto=…

The total run time was about 2 minutes, at 46 tok/s. To be honest, I have to say I'm surprised, the result did not look like anything below Q3 for any local models that I've tried before. Here's a closer look at it:

https://preview.redd.it/z67sg0hlrvth1.png?width=1905&format=png&auto=…

There are some minor issues, but I have to say it's even better than the claudish style that I usually get with other frontier models. Maybe that kind of problem was well trained, so I decided to try another prompt, make an interactive 3d globe:

https://preview.redd.it/sccdghnzuvth1.png?width=1870&format=png&auto=…

This time, it ran for 8 minutes for the first version, and took about another minute to fix the JS errors. The result came out still impressive.

https://preview.redd.it/lzik4fezvvth1.png?width=3436&format=png&auto=…

You can see the two demos yourself here:

\- https://pitest-beta.vercel.app/bakery/

\- https://pitest-beta.vercel.app/earth/

💬 7 (+7) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/nidarshan1 · 3d ago
Every Jev clone copied the same flaw, and it isn't the price.

I’ve been looking at the recent wave of System 1 decision models following Jev, and this is the issue I keep coming back to. Every Jev clone copied the same flaw, and it isn't the price. Jev shipped Sept 15. Three weeks later: 14 System 1 models. Cloudflare. Perplexity. OpenAI. Liquid. Upstage. Together. Inception. A dozen more. All the same contract: Choice, Score, Noul. Prices already at $0. They copied the format. They also copied the flaw. Jev's own docs admit decisions that don't add up. A question and its negation don't sum to 1. The model can be confidently wrong in two directions at once, with no way to say, "I don't know." A cheaper token doesn't fix that. A bigger context window doesn't either. Only coherence does: forcing the answers to agree with each other. Everyone's racing to be the cheapest copy. The market is the one that knows when it's wrong.

▲
0
 
8👁
r/LocalLLaMA · u/HoujunDev · 3d ago
TTS silently dropped 17% of a passage and nobody could hear it — so I built a local audiobook tool that transcribes every line back

I've been building VoxStage, a local script-to-voice workstation for Apple Silicon Macs. Paste a chapter of prose with no speaker labels, and it gives you a multi-voice reading you can audition line by line, fix, redo and export. Everything runs on the Mac: no account, no cloud API, no telemetry.

Why it exists: in an earlier local voice-cloning test, a long passage came out fluent and natural — and 40 characters (about 17%) from the middle were simply gone. The remaining text still read as a normal sentence, so nobody could hear it. That changed two design rules:

  1. Generate sentence by sentence, never a whole passage at once.
  1. Transcribe every generated line back with a local recogniser (whisper.cpp) and diff it against the script. Disagreements are flagged for your ear, never auto-corrected.

The stack:

\- Speech: Qwen3-TTS on MLX (0.6B / 1.7B preset voices, voice design from a description, cloning from a recording you confirm you have the rights to)

\- Who says what: a local LLM via llama.cpp (Qwen3-14B, or Qwen3-30B-A3B on 32 GB) drafts the speaker for each line; program-side rules on top; you review

\- Read-back check: whisper.cpp

Measured on my M2 Max 32 GB, Pride and Prejudice ch. 1: speaker draft for 35 units in 14.8 s; 28 lines → 144.7 s of audio synthesised in 51 s (RTF 0.358, preset-voice path); read-back check 28 s.

Honest limits:

\- The speaker draft is a draft. In my evaluation most scenes needed at least one correction, so the review step is the product, not a formality.

\- Chinese and English only for now.

\- Install is still developer-style (Homebrew + terminal, \~30 min mostly model downloads) and only verified on my own Mac. A signed one-click installer is in progress.

Other things it does: editing one sentence regenerates only that sentence; subtitles (SRT/VTT) timed from the actual audio; an FCP7 XML timeline that imports into DaVinci Resolve; long texts kept as a book with chapters inheriting the cast.

Samples (longer ones first): https://houjun.dev/voxstage/#listen

Code (AGPL-3.0): https://github.com/hera2019/VoxStage

I'd especially like to hear:

\- Which local models you've found best at speaker attribution in fiction

\- Whether anyone has seen the same silent-skip behaviour with other TTS models

\- If you try the install on a Mac other than an M2 Max, whether it works

💬 3 (+2) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/bring_back_the_v10s · 3d ago
Playing the devil's advocate

This is a reaction to https://www.reddit.com/r/LocalLLaMA/s/IxLnzcGjAU

You know the argument: frontier labs are hypocritical because they're making money on public scraped data.

I have absolutely no intention to defend OpenAI or any other frontier AI company here, but if you allow me I'll play the "devil's advocate" a bit for the sake of discussion and enlightenment, because sometimes when I think about that argument it seems to me it's quite weak. While OpenAI and Anthropic models are built on public data that they didn't pay for (at least most of it), there would be no frontier model without all their computing power and technical expertise that they invested on to build their models. Like, the data is already there, it's been public for ages, so what's preventing you, the average Joe, from building a Claude Opus 5? All you need is a huge data center, a nuclear power plant and an army of data scientists to build it, right? And then you need all that to run the inference. And you need to maintain it, and that costs money. And you need to keep evolving it, which costs money too. And if you get investors money then you'll eventually have to give some of the profit back to them. etc, etc.

So what am I missing here? All things considered, the data is already public, so it's already "free", right? But you need to dump tons of time and money on it to build a frontier LLM out of it. Of course if they're infringing copyright then that's a different story but in general the whole principle of built-on-free-data still stands.

💬 35 (+18) open on reddit ↗
▲
0
 
14👁
r/LocalLLaMA · u/potatocellfarmer · 3d ago
need help with ollama

hello i have an endeavourOS setup on a laptop with 40GB of ram and 8GB of vram (rx6800s) and AMD ryzen 9 6900HS
ollama is installed as a system service with vulkan extras from official arch repos
cline and librechat and odysseus are connected to the local ollama instance

here are my problems with ollama:
offloading layers to vram tanks my tk/s to nearly half
sometimes the response cuts out on librechat while the same model works fine on cline or directly on ollama (my suspicion is context limit)
ollama randmoly decides the gpu isn't there and does full cpu load

here is a list of things i tried:
ram speed is at full DDR5 speeds during generation
ollama correctly identifies the gpu and ignores the igpu
running smaller models like qwen3.5 that fully fits in the vram still gives me about 2 to 3 tk/s
switched to rocm version of ollama and saw no difference

temperatures are under control and nothing thermal throttles
using lm studio improves the generation to the higher end of 3 tk/s but nothing further
the laptop is plugged in and in high performance profile
qwen3.5:9b and qwen3.8:27b and gpt-oss:20b and gemma4:31b all max out at 3 tk/s

it seems like no matter the model size or if its a full vram scenario or full ram i'm locked at 2 tk/s
i am out of ideas at this point
any help would be appreciated, thank you very much

💬 6 (+4) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/Impressive-Lion5317 · 3d ago
Local-AI-Studio Update - https://github.com/vinnyclegg-dev/Local-AI-Studio post image

Claude Opus 5.5 wrote, planned and directed this 12-minute sci-fi film, rendered entirely on one PC with Local AI Studio.

This film started with one line typed into Claude Code: "create me a video history lesson from the perspective of the future". From there Opus wrote the script, planned all 124 shots and ran every render through Local AI Studio, a self-hosted creative workstation on one RTX 4080 SUPER. Nothing here was filmed, licensed or stock.

HOW IT WAS MADE
• Picture and ambient sound: MiniMax H3 (video with native audio)
• Keyframes for every shot: FLUX.2 Klein
• Narration: Breeze TTS 2
• Score: ACE-Step 1.5
• Titles, captions and the edit: HyperFrames
• Writing, shot planning, prompting and review: Claude Opus 5.5 in Claude Code

It was made over about ten days, in thirty-second batches, each reviewed at finished quality before the next began. It was then rebuilt as one continuous film, so the score and narration run across scene boundaries.

All footage, narration and music in this video are AI-generated. Video generated with MiniMax H3.

Local AI Studio is free and open source. One runbook rebuilds the whole studio on a clean machine: https://github.com/vinnyclegg-dev/Loc...

💬 3 (+1) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/dtdisapointingresult · 3d ago
Demo of how to guarantee untrusted Docker containers aren't allowed to connect out to upload your data

There's some questionable apps posted on here all the time. Honestly, it's not so much the vibecoding, but that these apps could be malicious/incompetent and leaking your data by uploading it to the dev's servers. I don't have time to review anything tbh, but I often want to try stuff.

If you use Docker, there's a somewhat simple technique that can give you piece of mind. Using this approach, you can run any untrusted service, but it's not allowed to connect out. It can only reply to incoming requests. Good enough for most apps.

Essentially, you write a docker-compose file that runs the service as usual, but put it behind a 2nd service that acts as a gateway that blocks outbound traffic.

(Side note: many repos make the awful decision of giving 'docker run' examples for running them in docker. Ask any LLM 'Convert the following docker run command to a docker compose file'. I recommend you always use compose files in your life anyway, it's 'docker run' with easy repeatability + backupability/git committing + more features like multiple services in one file which we need here. Then you just cd to ~/dockerstuff/someapp/ and run 'docker compose up')

Let's say the original unstrusted app's compose file is this, example is a service on port 8000

services
untrustedservice:
image: python:latest
container_name: untrustedservice
ports:
- "0.0.0.0:8000:8000"
command: [python, -m, http.server, "8000", --directory, /srv]

Use this instead, where we: 1) lock out untrusted service from the main network, 2) use socat as a one-way gateway to reach the untrusted service. socat is a tiny open-source binary, only 1.2MB RAM needed by the extra container.

services:

# socat gateway
untrustedservice_gateway:
image: alpine/socat:latest
container_name: untrustedservice_gateway
init: true
#Redirect incoming port 8000 connections to untrustedservice's port 8000
command: TCP4-LISTEN:8000,fork,reuseaddr TCP4:offline_untrustedservice:8000
ports:
- "0.0.0.0:8000:8000"
networks:
# Only this gateway connects to both networks
- public_network
- isolated_network
depends_on:
- untrustedservice

The expanded untrusted service definition # Notice how "ports" has been removed, the gateway is our entrypoint untrustedservice: image: python:latest container_name: offline_untrustedservice init: true command: [python, -m, http.server, "8000", --directory, /srv]

untrustedservice is limited to isolated_network networks: - isolated_network

some extra lockdown measures I don't really understand. Optional. cap_drop: - NET_ADMIN - NET_RAW security_opt: - no-new-privileges:true

networks:
# Normal network needed by gateway
public_network: {}

Network without internet but allowing replies to gateway connections isolated_network: internal: true

▲
0
 
3👁
r/LocalLLaMA · u/oppoftemp27 · 3d ago
If you benchmark llama.cpp on AMD with the official ROCm builds, check that you're actually on the GPU

I spent the last few weeks comparing llama.cpp output across backends on a small multi-vendor GPU fleet, and the single most useful thing I learned wasn't about inference quality — it was that on four AMD hosts, the official ROCm prebuilt (b11327) never touched the GPU at all. The server started fine, /health was green, and everything looked normal. It was running on the CPU the whole time.

The reason is dull but nasty: the prebuilt's HIP backend wants libamdhip64.so.7 / libhipblas.so.3 / librocblas.so.5, and a stock Ubuntu host with ROCm 6.3.0 has .so.6 / .so.2 / .so.4. The backend library fails to load, and llama-server quietly serves from the CPU. No error on the console. Even -ngl 999 doesn't change anything. Details here: https://github.com/ggml-org/llama.cpp/issues/26964#issuecomment-6024889471

If you benchmark tokens/sec you'll notice eventually. But if you compare output quality — perplexity, eval scores, side-by-side generations — nothing gives it away. On my boxes the "ROCm" perplexity matched the CPU perplexity to the last digit, every time, because it was the CPU doing the work.

The cheap check that catches this: run the same prompt set through the CPU build on the same box and compare per-prompt wall time. Ratio around 1.0 = you're on the CPU. Well under 0.5 = the GPU is actually working. I now run this check before trusting any benchmark number off a new box.

The same sweep also produced a small cross-backend conformance dataset (same GGUF, same prompts, temperature 0, per-token top-5 logprobs on every backend) — the short version: identical stacks are bit-for-bit deterministic across machines, and anything that changes the numerical stack (different backend, different build, even a different host CPU) starts flipping near-tie token choices, with the effect getting much worse the heavier your quantization is (F16 mostly agrees, Q4 mostly doesn't). Happy to share the data if anyone wants it.

💬 4 (+1) open on reddit ↗
▲
0
-2
8👁
r/LocalLLaMA · u/kmodi · 3d ago
We gave Aleph Alpha's Kolibri-1 up-to 72 action combinations and put it in Doom. What could go wrong? 🎮 post image

Up to 72 action combinations. Four decision groups, one batch.


Was super interesting challenge to make these many decisions in one pass to get the latencies to : 39ms median. 60ms p95 in our test.

Apparently enough time to make questionable decisions.

source: https://x.com/konarkmodi/status/2107569039790751925?s=46
Watch 👇
https://tesseracted.com/kolibri-1-chat/gameplay/doom

💬 2 (+2) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/TyedalWaves · 3d ago
Been out of the loop for a while

Hey guys! School started up and I fell a bit out of the loop with local LLMs. Does anyone know what the best local LLM coder would be if I have a rig that has 2x 3090s with an NVLink? I appreciate your guy's help!

💬 11 (+1) open on reddit ↗
▲
0
-1
4👁
r/LocalLLaMA · u/Wvdy_CC · 3d ago
repOx v0.2.0: Added architectural --outline mode (80% token reduction), synthetic tool-call JSON format, and Git diff packing based on your feedback

A couple of days ago I shared repOx (a sub-15ms Rust CLI & lazygit-style TUI for packing repositories into LLM prompts) and got awesome feedback from this community.

I just released v0.2.0 implementing the most requested features:

  1. Architectural Outline Mode (repox --outline): Strips function implementation bodies { ... } and keeps only structs, classes, traits, imports, and function signatures across Rust, Python, Go, TS/JS, and C/C++. Cuts token usage by 75–85% when you only need architectural context.
  1. Synthetic Tool-Call Format (repox -f tool-call): Formats the repository as a JSON array of read\_file tool calls & responses — great for agent harnesses and local models trained on tool-use trajectories.
  1. Smart Lockfile Summarizer (repox --summary-locks): Instead of burning 40k tokens on Cargo.lock / package-lock.json or hiding dependency versions completely, it parses lockfiles (Cargo.lock, package-lock.json, pnpm-lock.yaml, poetry.lock, yarn.lock, go.sum) into a tiny "package @ version" manifest (95%+ token reduction).
  1. Git-Aware Packing (repox --modified / --staged): Pack only the files touched in your current working tree or staging area.
  1. TUI Upgrades (repox -i): Added lexical syntax highlighting in the preview pane, Shift+C to copy a reproducible CLI command, and OSC 52 clipboard fallback for tmux / herdr / SSH.

Install / Update:

\- Crates.io: cargo install repox-cli

\- One-liner: curl -fsSL https://raw.githubusercontent.com/WVDYC/repOx/main/install.sh | sh

GitHub: https://github.com/WVDYC/repOx

💬 2 (+1) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/Mysterious-Desk-3492 · 3d ago
Pi and mini-swe-agent passed 9/9 checks each in my latest experiment. A second code review still found defects in both.

As part of my AI Studio project, I’m testing harnesses for coding.

The initial screening included 10 harnesses:
Pi, mini-swe-agent, Crush, OpenCode, Goose, Prime Agent, Oh My Pi, Qwen Code, Octomind (reduced offline profile) and Aider.

Models:
• DeepSeek V4.1 Flash
• Qwen3.8-27B
• Laguna S 2.1

All accessed through OpenRouter.

The detailed code review covered Pi, mini-swe-agent, Crush and OpenCode across all three models and tasks: 36 combinations. Two attempts produced no patch.

The three Golang tasks were deliberately different:
\- Add strict validation for an HTTP query parameter.
\- Migrate 200 logging calls while preserving behaviour and context.
\- Add bookmark tags across the API, storage migration and HTML rendering.

Pi and mini-swe-agent each passed the original acceptance checks on all nine combinations. But a second agent review, followed by isolated reproduction probes, exposed three gaps:
\- silently dropped malformed query fields
\- bookmark's task exposed a mutable tag slice from the store
\- one migration rejected a valid older store

Good news too: all logging migrations preserved behaviour in differential probes covering 100 functions and nine integer inputs, including the minimum and maximum values.
My takeaway: the evaluator and the reviewer both need testing. A green result is evidence about the checks we ran; broader correctness needs further evidence. This experiment did not establish a decisive winner between Pi and mini-swe-agent. Human correction time also is still unmeasured.

Any opinion welcome.

💬 2 (+1) open on reddit ↗
▲
0
 
14👁
r/LocalLLaMA · u/Fantastic_Sound2049 · 3d ago
Hey guys newbie here

This is my first time trying to locally host an ai but i want to find a good model that can fit on my rtx 4050 laptop gpu that has 6gb vram and my laptop has 24gb ddr5 ram (4800mt/s) so can you suggest me a model that can code websites or small app like inventory management or similar also when i asked chatgpt about any suggestions it said Qwen3-Coder 8B, Q4\_K\_M is the best for my needs and i searched it on yt and only saw bad reviews plz help me guys Thank you

💬 17 (+6) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/forevergeeks · 3d ago
Are AI influencers just repeating the same talking points?

Hi everyone,

Are influencers talking about local AI all using the same script? Same benchmarks, same kitchen examples, same terminology?

I keep seeing videos about running Qwen3.8-Flash-Next on 12GB of RAM using a new runtime engine called Strata. Every video makes the same claim.

But that is confusing, especially for people who are new to this. What you need is 12GB of VRAM, not 12GB of regular RAM. That means you need a dedicated graphics card.

There is a big difference between RAM and VRAM.

As far as I understand it, Strata needs:

  • 12GB of VRAM
  • 64GB of regular RAM
  • 80GB of SSD space

So saying it runs on 12GB of RAM is misleading.

💬 21 (+4) open on reddit ↗
▲
0
-1
4👁
r/LocalLLaMA · u/FriendlyLie23 · 3d ago
SentryGate: An open-source AI Gateway with sub-10ms semantic vector caching and dynamic LLM routing (Ollama & OpenAI compatible)

Here is a common problem with building apps on LLMs:

Users ask the same question over and over.

Your app calls the model every single time.

You pay the API bill every time. Users wait 3–5 seconds every time.

**\*\*SentryGate\*\* is an open-source AI traffic controller that fixes this in literally one line of code.**

\### What it actually does:

\* ⚡ \*\*Lightning Fast (4ms)\*\*: If someone asks a question that was already answered, SentryGate serves the saved answer in 4 milliseconds instead of 4 seconds.

\* 💰 \*\*$0 on Repeat Questions\*\*: Bypasses the model completely on repeat or similarly phrased prompts.

\* 🧠 \*\*Doesn't Get Tricked\*\*: Basic caches get confused between "How to bake a cake with eggs" and "How to bake a cake WITHOUT eggs". SentryGate catches tricky negative words so it never serves the wrong answer.

\* 🕒 \*\*Knows What's Fresh\*\*: Real-time questions ("today's weather", "current stock price") automatically skip the cache.

\* 🔌 \*\*Zero Downloads / 1-Line Setup\*\*: No new libraries. Just point your existing OpenAI / LangChain \base\_url\ to SentryGate and keep your code 100% untouched.

Works locally on your machine with Ollama, or in the cloud. Completely open-source under the MIT license.

\* 🌐 \*\*Test the Live Playground (No login needed)\*\*: https://sentrygate-9ght.onrender.com

\* 💻 \*\*GitHub Repo\*\*: https://github.com/Prisha2004/Sentrygate

Note: Open-source project maintainer (MIT License, 100% free).

Feedback and stars are welcome!

▲
0
-1
3👁
r/LocalLLaMA · u/khalon23 · 3d ago
agent manager 0.40 model and effort pickers that stay current with the CLI

released 0.40 of agent manager. it is a local go/tmux tui for running a few coding agents side by side.

added support for model and effort pickers per session. unlike a lot of other software we fetch the model list dynamically from each CLI so it stays up to date without updating the TUI. a new vendor model shows up without an agent manager release.

also added Antigravity CLI and Oh My Pi as built in agents.

https://github.com/YoanWai/agent-manager/releases/tag/v0.40.0

▲
0
-1
6👁
r/LocalLLaMA · u/-MaskNinja- · 3d ago
Really need someone to help me run tests on my benchmark, any model

I've had some updates to this benchmark, meaning it should run smoothly compared to before. I have an umm, measly GPU with 8 GB of VRAM, so I can't run models like Qwen3.8-27B on it. Any result would be good from you guys; although this is more suited to frontier models, smaller models should still work. I'll credit you for any results if you'd like.

The primary reason I thought it would be interesting to run: benchmarks are nearly all pass or fail on a single-dimention graph, so I thought it might be worth shaping a new one up. BinkBench measures video quality and video compression rate, which gives you two things to plot on. The agent also can't score 100% - there isn't an end, which makes it progressively harder as the agents get smarter, because they need to implement more novel techniques. I also thought video encoding would be good as a benchmark, since it's not something we've tested agents on before and is pretty hard. It's like the kernel/compiler optimisation things we've seen other labs show tests on.

More info is on GitHub,

Old post: https://www.reddit.com/r/LocalLLaMA/comments/1vn6nlr/looking\_for\_people\_to\_help\_me\_run\_a\_benchmark/

▲
0
 
6👁
r/LocalLLaMA · u/abrdeveloper · 3d ago
(Self Promotion) Kimi vs. Claude vs. GPT vs. Gemini as teammates. Who actually coordinates?

Benchmarks test models alone. I wanted to know how they do with a partner.

We paired four models in every combination in a co-op game where players are tied by a rope. Top with top won most. A third or fourth agent hurt every model. Human pairs still beat all of them.

I work at Skillprint. We build games like this to capture how people and models coordinate, because AI that works alongside people needs that context.

Pairing matrix and GIFs: https://experiments.skillprint.co/posts/signal/
Play it yourself: https://experiments.skillprint.co/play

Which matchups should we run next?

💬 1 (+1) open on reddit ↗