165 posts · 1 sub · RSS
← prev Tuesday, October 6, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
39
+29
21👁
r/LocalLLaMA · u/Gold-Bat-3225 · 3d ago
MiMo V2.6 Pro almost matched GPT-6 Astra at figuring out what a user actually wants post image

Introducing InferBench: A benchmark testing how well frontier LLMs infer a user's priorities from their instructions.

We tested 12 LLMs across 20 scenarios with 2.8k conversations to see which models understand the user's goals.

Each conversation has a simulated user with a private profile of their priorities. The assistant decides on the best option or can first ask clarifying questions.

When a user's priorities were clearly stated up front, the models picked the best option 89% of the time. If some priorities were initially unstated then accuracy went down to 60%.

The open weight models did better than I expected:

GPT-6 Astra: 76%

MiMo V2.6 Pro: 75%

Gemini 3.1 Pro: 69%

Astra always asked clarifying questions when preferences weren't stated and never did when they were (extremely impressive). By comparison, Grok 4.6 asked to clarify in 18/64 conversations.

However LLMs are still just as confident when they're wrong. 105/288 wrong decisions were submitted with >90% confidence.

The full report is linked here: https://laugh.so/research/inferbench/

What surprised you the most?

💬 11 (+7) open on reddit ↗
▲
0
-1
3👁
r/LocalLLaMA · u/khalon23 · 3d ago
agent manager 0.40 model and effort pickers that stay current with the CLI

released 0.40 of agent manager. it is a local go/tmux tui for running a few coding agents side by side.

added support for model and effort pickers per session. unlike a lot of other software we fetch the model list dynamically from each CLI so it stays up to date without updating the TUI. a new vendor model shows up without an agent manager release.

also added Antigravity CLI and Oh My Pi as built in agents.

https://github.com/YoanWai/agent-manager/releases/tag/v0.40.0

▲
7
+4
11👁
r/LocalLLaMA · u/FikoFox · 3d ago
Free online conference on Oct 22 with a few talks on small models, local inference and speed

Hey all,

I'm helping organize All Day AI, a free online conference on Thursday the 22nd.

I'd like to mention a couple of our speakers who have volunteered to talk that day:

Vivienne Hnin (Utilyst): "Small Models, Big Profit Margins: The Economics and Engineering Tradeoffs of SLMs." That's the debate this sub has every day: when a small model is actually good enough.

Hossein Kazemi (Astorna): "Using Task-Specific Small Language Models to Handle Sensitive Data". Keeping sensitive data local is half the reason people run models themselves.

Hitesh Jain (Coral Bricks AI): "Coding at 250 tokens." Inference speed is what this crowd benchmarks obsessively.

The rest of the schedule goes up soon, across four tracks: Build, Lead, Secure and Ship.

The talks are community-submitted.

Free to register: https://www.alldayai.com/?utm\_source=localllama&utm\_medium=reddit

I hope this and other talks in this space may help you discover We have a discord channel you can join: https://discord.gg/xUyS3Zu68

💬 2 (+1) open on reddit ↗
▲
60
+49
26👁
r/LocalLLaMA · u/BinaryGrind · 3d ago
I have about $4000, what's the best setup to get?

Ideally I'd like to be able to run Qwen 3.8-Flash-Next with decent performance.

I was thinking of just buying 2x Radeon AI R9700 (64GB VRAM), or maybe a DGX Spark but that was before the price hike. My brother suggested just getting a Strix Halo box with 128GB unified.

I did see I could buy 6x Intel Arc B60 (24GB each, 144GB VRAM Total), but researching seems like the performance of the B60 is lacking. I'd also need a new v
motherboard/CPU that can run 6 GPUs.

I'm also not opposed to getting a Mac Mini or Studio if the price and performance is right.

The $4000 is not exactly a hard cap, like I could stretch to $4200 without too much struggle, but obviously the cheaper the better. I'm lucky to have a decent stock pile of NVMEs and DDR4/DDR5 UDIMMs, so if I need to build a box, I could, would just need the motherboard and a CPU if I can't just slot in either the Intel 14700K or Ryzen 9700x I already have.

So where am I swiping my credit card?

Edit: This is a use it or lose it budget from my work, can't really save it.

Edit 2: To clarify again, this is extra money in the IT budget that we need to burn by the end of the year. Telling me to save it, donate it, invest it, isn't helpful as I can't do that.

💬 185 (+117) open on reddit ↗
▲
15
+12
20👁
r/LocalLLaMA · u/bodhi371 · 3d ago
Qwen3.8-27B with just 12GB VRAM + 8GB RAM - ~18 tok/sec

I got Qwen3.8-27B running at \~18 tok/sec decode & \~500-600 tok/sec prefill (at 64k context) on just 12GB VRAM + 8GB RAM, using ISTA-DASLab Qwen3.8-27B-GSQ-RCO IQ3\_S quant (\~11GB). This recipe also works with 0bserverx’ Qwen3.8-27B-Heretic-GSQ-RCO IQ3\_S quant, achieving similar speeds (about a 7% loss).

This is the best Qwen3.8-27B quant I’ve tested so far (and I’ve tried everything), and for it to fit in such limited RAM/VRAM is wild. GSQ-RCO quantization is magic, it performs very close to the full precision weights in all of my testing.

The reason it fits at all is Qwen3.8 is hybrid, so only 16 of the 64 layers need KV cache. With q4\_0 for cache the full 64k is only about 1.1GB instead of 4GB for f16.

I'm on a 9900X + 4070S 12GB + 32GB RAM for reference, using stock llama.cpp. Settings are -ngl 58 -ot token\_embd=CPU -ctk q4\_0 -ctv q4\_0 -c 64000.

Full build + serve scripts and all my numbers are here if you’d like to reproduce yourselves: https://github.com/bodhi37/Qwen3.8-27B-12GBVRAM-Recipe

💬 12 (+6) open on reddit ↗
▲
1
 
1👁
r/LocalLLaMA · u/Shookpro · 3d ago
Routeweaver - Serve 27b fast on low vram set ups

With qwen 4 on the horizon I thought I'd share my latest update on my rtx 3060 12gb setup that makes 27b fully usable, I'd also like to see people with bigger gpu's try it out. Get your agent to set it up although bigger cards and different cpu set ups may have to tune the custom kernals i have put together:)

▲
12
+10
17👁
r/LocalLLaMA · u/NoFee9147 · 3d ago
Optimizations Claude did for Qwen 3.8 27B and Qwen Flash Next on dual and quad 7900xtx

​

They're all fixes in rocm and llama server. Let me know if this interests someone. I'll push it to GitHub and share my configs. I have a Lenovo p620 running 4 7900xtx on gen 4.0 x16 slots.

One line summary of each fix:

\- Async mirrored input uploads: \~30 per-token inputs staged in a pinned ring on per-GPU streams instead of synchronous round trips (Flash-Next decode 24.6 -> 36.3 t/s).

\- Per-device dispatch threads: each GPU's kernels and collectives launched by its own host thread instead of one thread for all four.

\- One-shot PCIe P2P AllReduce: GPUs write slices straight into peers' memory for small tensors, replacing RCCL (\~57 us -> \~9 us per allreduce).

\- Fewer kernels in MTP decode: fused same-shape copies and leaner conv-state rollback (88.8 -> 92.7 t/s).

\- mmvq small-K row packing (RDNA3): short-K projections no longer leave most of the block idle (22.7 -> 25.3 t/s).

\- Wide-K mmvq for multi-token batches: K split over 8 warps for MTP verify on 10240x320 projections (\~92.6 -> \~96 t/s).

\- MoE vector kernel up to 8 tokens: 5-token MTP verify batches stay on the fast vector path instead of MMQ (80.9 -> 86.7 t/s).

\- Small-K multi-token MoE kernel: several rows per warp for short expert down-projection slices (94.2 -> 95.2 t/s).

\- Wide mmvf blocks: 512/1024-thread blocks for tiny long-K F32 matrices (40.6 -> 41.2 t/s).

\- Q8\_1 activation registry: each activation quantized once and shared by all matmuls that read it (\~+2 t/s).

\- Fused hyper-connection chains: scale/sigmoid/scale/hc\_post and scale/silu each in one kernel, \~380 fewer kernels per token (38.5 -> 40.6 t/s).

\- Thin-F32 prefill kernel: <=16-row F32 matmuls off generic SGEMM (401 -> 21 us; pp4096 1602 -> 1836 t/s).

\- Compact MoE tile list: expert matmul launches only (expert, token-tile) pairs with work instead of a 96%-empty grid (down 928 -> 405 us, gate/up 516 -> 360 us).

\- Multi-warp MoE routing helper: 16 warps per expert sort tokens in two passes (99 -> 25 us per call).

\- Q4\_K expert tile shape: 32-row tiles for 160-row expert slices (360 -> 337 us).

\- Stream-k for few-tile Q8\_0 matmuls: 10240->320 projections spread over all CUs instead of 12 workgroups (284 -> 115 us).

\- No 64-bit div/mod in hyper-connection kernels: 3-D grid instead of emulated integer division per element (230 -> 79 us; pp2048 1593 -> 1868 t/s with stream-k).

\- Split-K router GEMM: 512x512 F32 router GEMM as 8 K-chunks plus a sum (158 -> 69 us).

\- MTP re-reserve fix: graph re-reserved when MTP outputs turn on, ending a full GPU realloc+sync per prefill chunk (88.5K prefill 1042 -> 1130 t/s).

\- MTP draft prompt window: draft head prefills only the last 2048 prompt tokens (prefill 1130 -> 1249 t/s, decode at depth 46.6 -> 64.9 t/s).

\- Draft ubatch cap: draft compute buffer 457 -> 247 MiB, fixing a GPU0 out-of-memory crash at 96K context.

\- Gathered sparse attention (QSA): decode attends only to the \~2K selected tokens instead of masking the whole cache (decode at 48K 67.9 -> 80.6 t/s).

\- Per-layer embedding table in RAM (--lazy-mode off): 26.8 GB hashed embedding table kept resident instead of read from disk every pass (lookup 1.0-1.6 -> 0.1-0.3 ms, \~3-4% decode).

\- Meta backend subgraph fix: per-device subgraphs sized for the largest graph, fixing a segfault when graph shapes change between calls.

\- Net result, Qwen3.8-27B Q8 on 4 GPUs: code decode 54-61 -> 96-110 t/s, 51K prefill 1454 -> 1816 t/s, decode at 51K depth 57 -> 77 t/s.

💬 11 (+6) open on reddit ↗
▲
65
+59
27👁
r/LocalLLaMA · u/TheVoxcraft · 3d ago
pi-optchat: never compact again - endless chat as a memory tree post image

I built a Pi extension that implements Victor Taelin's OptChat recipe: instead of compacting, every message is logged and summarized into a binary tree. Each turn starts from a fresh context with a bounded memory view (128 KB), and the agent uses zoom/date to read the originals when it needs them. One endless chat per profile, no fork, no separate launcher.

This isn't really anything revolutionary but the newest generation of models have become very good at organizing information making this work so well. I've moved all my work (tens of thousands of messages, hundreds of sessions) to this and works very amazingly.

Install with pi install npm:pi-optchat

What's in it:

  • Profiles — separate memories and instructions (I run work and personal).
  • Subagents — spawn background agents, watch them live, send guidance, interrupt with Ctrl+C, resume finished ones with tell. Reports from one spawn arrive grouped.
  • Import — bring in your history from Claude Code (sessions and auto-memories), Codex, or a ChatGPT export.
  • Connected windows — open a second Pi on the same profile and it becomes a subagent you talk to directly, with a handoff when you /complete.

Repo: https://github.com/jonaslsaa/pi-optchat
Video credit goes to https://github.com/aaaxn

💬 41 (+29) open on reddit ↗
▲
1
-2
9👁
r/LocalLLaMA · u/rootshelldev · 3d ago
An API gateway for the desktop user

As a developer working at home on a single GPU i am building and experimenting a lot not only with coding agents but also with apps that use generative APIs. The flood of models, engines, and different APIs makes it hard to always get it right in every app and tool and to keep it up-to-date. I needed a gateway that i could target in all my apps while also being able to use it with clients that only support official upstream APIs from Anthropic and OpenAI.

So i build a gateway for myself and over time extended it with more and more Features. Its a Rust based desktop app using Tauri and Leptos. Leptos is WASM running inside a lightweight gtk webview. I did it specifically this way to allow for remote browser based administration when working from another device in my network, but have it ready in the tray on my desktop at any time. I also wanted to integrate tools for quickly testing new models and llama.cpp patches. What it does:

  • Supports Text, Embeddings, Audio and Image APIs.
  • Presents OpenAI and Anthropic compatible APIs and routes them to cloud APIs or into llama.cpp, audio.cpp and stable-diffusion.cpp containers.
  • Builds the backend containers directly from git inside of podman containers, including MRs, PRs from main or a specific branch or commit. Notifies for updates.
  • One container per model and manages their lifecycle while scheduling VRAM-aware with fallbacks.
  • Models with different configurations can be registered as aliases, so client configuration does not have to change for model changes. These aliases also support model chains. For example reloading the model with more configured context only when needed, routing to cloud if a certain context length is reached. Or chains like: small context & high quant -> big context & lower quant at context length steps.
  • A "GPU hold" mode that can be triggered from the tray, it unloads models, blocks new models from loading and responds with either an error or routes to a fallback if configured. For gaming or other blocking GPU use.
  • Fallback routing to any other configured model or alias in case of VRAM contention or an active GPU hold.
  • Offers an OpenAI compatible websocket with /v1/realtime via a staged pipeline: VAD / Smart Turn -> ASR -> LLM -> TTS while being able to select either a local model or a cloud model for each step of it. This includes barge-in and session management. Tools are supported and executed server side.
  • MCP Gateway: Add your MCP servers to the gateway and it offers them via prefix and scoped per token via its own /mcp API as a streamable MCP server. MCP servers are executed inside of podman containers by default.
  • Scoped Tokens, detailed metrics, traffic monitoring, budgets, price sync (only tested with kilo) with lots of graphs, cost comparisons for local tokens if it would have been cloud traffic.
  • Download manager with huggingface downloads and updates, preconfigured catalogs for audio.cpp and sd.cpp.
  • Integrated MCP Server with admin tools on a seperate MCP route. Every option and feature of the gateway is configurable via the MCP. Adding models, testing configurations, gateway config and status. A coding agent can configure it for you.
  • Documentation MCP like context7. Agents that connect to the MCP and have the docs tools enabled can query and request documentation for specific library versions of any kind. All requests are listed in the UI and can then be chunked, embedded and ingested into a vector storage (all managed by the gateway). Context7 is really great but often stale, some libraries are missing like my own. The Interface is not intuitive but the gateways admin MCP lets my agent fill it anyway.
  • A chat interface with metrics, file input, folders, thread-specific settings, Mic Input & TTS selectable from the gateways own models. For quickly testing models. Tools from the MCP gateway and the gateways own tools can be added as well. (Admin Chat as a preconfigured chat for gateway administration)
  • Voice Chat mode in the chat interface based on the realtime API, with normal dialog flow or push-to-talk
  • /v1/responses that works session based and supports server-side tool execution.
  • Audio and Image Labs for quickly testing audio and image tasks like image generation, image edit, tts, asr, cloning, conversion, etc.. I try to keep up with audio.cpp's and sd.cpp's tempo.
  • Container based agents, that a build to a specific interface mounted into the container (tools, vars and files) and can mount their own UI page and MCP tools to the gateway while running. Kind of like smaller, task based extensions.
  • Model benchmark with lots of graphs to compare configs or engines.
  • Integrated API docs in the spirit of Swagger with all APIs offered by the gateway
  • Lots more i forgot

I am usually very shy and thought long and hard if i want to risk exposure and publish all of this. But it has made my day in my specific scenario a lot more comfortable and maybe you like it.

https://preview.redd.it/l02s01qpdxth1.png?width=429&format=png&auto=w…

Here is the link: https://github.com/lmgw-dev/lmgw

I hope you dont tear me to shreds and can find use in it. Dont be to judgy on the Interface, my years of experience are all on the backend and in devops.

What is planned next:

  • Decision models

And the obvious disclosure: AI has played a role at all stages of development. Nearly everything is touched by a diverse set of models and most of the prose text in the repo is generated. I made sure to write this post by hand because you all deserve it and i am myself annoyed by generated posts. Also, i did not publish the git history and will squash most of my commits for safety reasons.

💬 6 (+4) open on reddit ↗
▲
508
+443
38👁
r/LocalLLaMA · u/chemist_slime · 3d ago
Europe rejoins the fight with Chonky! Mistral Large 4 Released, Open weights end of month, who’s ready?

1 trillion parameters, 49B active, definitely chonky! If you don’t love the model you gotta at least love the humor in the name - Le Chonk

💬 120 (+100) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/Van-233 · 3d ago
Anyone else with this suggestions? post image

Has anyone else noticed this behavior as an autocomplete suggestion?

▲
7
+6
14👁
r/LocalLLaMA · u/Ok-Shower7286 · 3d ago
Qwen3.8-27B (Q6_K_XL) 110+ TPS at 256k context on a single RTX 5090, with a KV buffer decoupled from context size

I'm sharing this project (honestly 2nd time) for anyone who wants to run ultra-long context tasks with high-precision quantizations, especially for heavy coding.

TL;DR: In vanilla llama.cpp, -c N allocates a physical KV buffer for all N tokens up front, including promprts and KV caches. In my fork (focus-llama) the logical context stays at 256k, but the physical KV buffer is capped (--kv-cache-size, \~85k cells). Older chunks are offloaded to an external store and pulled back on demand. This runs a 27B Q6 model with 256k logical context on one 5090, at roughly 100–120 t/s depending on how full the context is.

How to?

Vanilla llama.cpp allocates a contiguous KV buffer for the full -c (VRAM), and every decode step attends over all tokens currently in context, so per-token cost grows roughly linearly with context length (only the attention part; the weights matmul is constant). The two problems are separate: the allocation wastes VRAM, and the growing context slows decoding. Initial speed is around 120 t/s, dropping to 60 t/s as context grows.

focus-llama attacks both: the physical buffer is capped (\~85k cells) so VRAM is bounded, and since the number of resident cells can't exceed the buffer, per-step attention cost is bounded by the buffer size instead of the logical context length. It based on 2 techniques declarative attention and skill.state, introduced by google deepmind. I've spent the past two weeks ironing out bugs, and now that it has stabilized, I'm honestly blown away.

On a single RTX 5090 + Qwen3.8-27B UD-Q6\_K\_XL, adaptive MTP speculative decoding), it shows average 110+ t/s and 256k logical context with a \~85k physical buffer.

Configuration is somewhat tricky (focus-memory: kv cache store also required) but,

You'll see the MAGIC in action: context usage stays capped at around 18–30%, token generation speeds remain consistently high, and you'll never hit full-day compaction pauses when running coding harnesses like Cline or Qwen Code.

Link? https://github.com/edwardyoon/focus-llama/blob/master/README.md

My conf:
-m /home/edwardyoon/my_model/qwen3.8/Qwen3.8-27B-UD-Q6_K_XL.gguf \
--alias qwen27b \
-ngl 99 \
-b 1024 \
-ub 1024 \
-c 200000 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--da-auto \
--kv-unified \
--da-min-ctx 2048 \
--da-chunk-tokens 4096 \
--fm-offload \
--kv-offload-threshold 36864 \
--kv-offload-holes \
--kv-cache-size 85536 \
--kv-retain-tokens 6000 \
--sparse-gate-threshold 60 \
--focus-memory-host http://192.168.219.124:3900 \
--focus-memory-token focus-memory-local \
--spec-type draft-mtp-adaptive \
--spec-draft-n-max 6 \
--spec-draft-ngl all \
--spec-draft-type-k q8_0 \
--spec-draft-type-v q8_0 \

few journal logs:

llama-server[827038]: da_sparse[VEC]: SPARSE - gather bound 5120 of 15872 KV rows (32.3% of KV)
…
n_gen = 361, tg = 119.44 t/s, tg_3s = 119.77 t/s
n_gen = 729, tg = 120.98 t/s, tg_3s = 122.53 t/s
n_gen = 1083, tg = 119.71 t/s, tg_3s = 117.18 t/s
n_gen = 1497, tg = 123.98 t/s, tg_3s = 136.71 t/s
n_gen = 1819, tg = 120.50 t/s, tg_3s = 106.61 t/s
n_gen = 2157, tg = 119.06 t/s, tg_3s = 111.85 t/s
n_gen = 2509, tg = 118.67 t/s, tg_3s = 116.32 t/s
n_gen = 2837, tg = 117.37 t/s, tg_3s = 108.35 t/s
n_gen = 3234, tg = 118.99 t/s, tg_3s = 131.94 t/s
n_gen = 3643, tg = 120.65 t/s, tg_3s = 135.62 t/s
n_gen = 3983, tg = 119.98 t/s, tg_3s = 113.25 t/s
n_gen = 4350, tg = 120.09 t/s, tg_3s = 121.34 t/s
n_gen = 4713, tg = 120.08 t/s, tg_3s = 119.89 t/s
n_gen = 5150, tg = 121.80 t/s, tg_3s = 144.11 t/s
n_gen = 5527, tg = 122.04 t/s, tg_3s = 125.37 t/s
n_gen = 5866, tg = 121.45 t/s, tg_3s = 112.57 t/s
n_gen = 6291, tg = 122.59 t/s, tg_3s = 140.83 t/s

💬 13 (+5) open on reddit ↗
▲
0
-1
6👁
r/LocalLLaMA · u/-MaskNinja- · 3d ago
Really need someone to help me run tests on my benchmark, any model

I've had some updates to this benchmark, meaning it should run smoothly compared to before. I have an umm, measly GPU with 8 GB of VRAM, so I can't run models like Qwen3.8-27B on it. Any result would be good from you guys; although this is more suited to frontier models, smaller models should still work. I'll credit you for any results if you'd like.

The primary reason I thought it would be interesting to run: benchmarks are nearly all pass or fail on a single-dimention graph, so I thought it might be worth shaping a new one up. BinkBench measures video quality and video compression rate, which gives you two things to plot on. The agent also can't score 100% - there isn't an end, which makes it progressively harder as the agents get smarter, because they need to implement more novel techniques. I also thought video encoding would be good as a benchmark, since it's not something we've tested agents on before and is pretty hard. It's like the kernel/compiler optimisation things we've seen other labs show tests on.

More info is on GitHub,

Old post: https://www.reddit.com/r/LocalLLaMA/comments/1vn6nlr/looking\_for\_people\_to\_help\_me\_run\_a\_benchmark/

▲
13
+12
24👁
r/LocalLLaMA · u/Geritas · 3d ago
Is everything alright with llama.cpp recently?

My Gemma4 31b seems to be breaking down in 'lalala' or just looping indefinitely for the past 3-4 days. Never happened before

https://preview.redd.it/tk7u2x2c9xth1.png?width=235&format=png&auto=w…

There was no 'lalala' in the whole scenario, I have no idea where it came from. Nor was there any skipping, humming or perfection. There were shivers down the spine of course, but it is still weird.

💬 25 (+18) open on reddit ↗
▲
1
 
17👁
r/LocalLLaMA · u/randomgenericbot · 3d ago
just my "how I run qwen3.8 27b on 16GB" experience and guide

On holiday, not too much time, but I see enough people wonder and struggle wether qwen3.8 27b can do real work on 16Gb VRAM.

short answer:

yes it can

longer answer:

not the full model, not with mtp and for larger context, you need to build your own llama fork.

Qwen3.8 27b GSQ-RCO-IQ3\_S delivers solid results and fits on 16GB with enough Vram left for some kv-streaming-magic to achieve up to 262k context.

Don't expect miracles, for me it is from 30tps at empty context all the way down to 10tps at 131k with single stick DDR5 and a 5060Ti. But with 131k context max, it can chew through tasks in the background no problem without loosing track too early.

full answer (and how I made it work):

Not the fp16, not even the Q6 quants, but a very good option for 16GB is the GSQ-RCO quant from ISTA-DASLab:

https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

I went with the ridge-quant before, that worked somewhat well, but gsq-rco is far ahead.

Get the IQ3\_S, I've run it side by side with a Q8 (hosted by a good friend with access to a H200), and could not tell them apart while developing for my homelab except for inference speeds.

Use the gsq-rco to aid you in building the kv streaming fork:

https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming

Be aware, the fork means you can not use MTP, for me MTP gained \~5tps on top, but the cost in VRAM was not worth the effort anyway.

Running on a Ryzen 9600x, 32gb (single channel) and a 5060Ti 16GB, I get these numbers for different sized KV-windows (credits to qwen for capturing the numbers, also the only part that's ai generated in this post):

Results (measured)

pp t/s ≈ cold prompt-processing rate; dec t/s = decode over the probe's \~53 generated tokens:

|tokens|pp @ any pool|dec t/s 512|dec t/s 1536|dec t/s 2048|
|:-|:-|:-|:-|:-|
|14 644|878–892|28.0|28.5|27.9|
|35 186|803–807|23.9|25.2|25.2|
|54 976|736|15.9|20.6|21.3|
|69 429|695|12.4|18.7|19.1|
|94 464|630–635|8.7|12.6|15.5|

Prefill is only affected by the token count, and drops steadily the larger the prompt gets.

Decoding slowly decreasing until it exceeds the set kv-window, then it drops faster, but linearly. Remember: I run single channel RAM, it might be better with dual channel. At almost 131k and 1536M window I get around 9tps, so thats the floor. With \~13.5GB model usage, its not possible to get 3G kv-window. In theory you cna go as low as 128MB, but then its slow from the beginning.

I found 1.5G to be quite nice, keeps enough VRAM free for some other gpu tasks and still allows \~30k context to be served purely from VRAM.

Some suggestions to get it on the rails:

The model loads on stock llama, Make use of it.

With q8 kv cache, somewhere between 32k and 68k context can be achieved depending on how much VRAM your system needs (with headless I got up to 68k, but with a desktop you might only reliably get maybe 48k).

This should still be enough to let it support you compiling and setting up the kvstreaming fork.

Stick it together with a harness like pi (pi.dev) and let it compile the fork - for me it was able to do that easily.

Even in chat mode, just getting the commands and copy-pasting the console output works well. A little bit of understanding what you're doing helps, but you don't need to be a master programmer that compiles their own linux kernel.

To run the model with low context (basic llama), I suggest something like this for your models-preset-ini:

[qwen38-gsq-rco]
model = /models-src/linked/qwen38-gsq-rco.gguf
mmproj = /models-src/linked/qwen38-gsq-rco-mmproj.gguf
ctx-size = 49152
cache-type-k = q8_0
cache-type-v = q8_0

Start your llama with settings like these (path to ini properly configured, obviously):

--models-preset /models-src/models-preset.ini --models-max 1 --host 0.0.0.0 --port 8080 --n-gpu-layers 999 --jinja\--flash-attn on--no-mmproj-offload

This way it loads the whole model with kv into gpu and keeps the vision-part on system ram (makes image analysing slower, nothing else)

With the new llama-kv-streaming image, you can then setup a "kv-window" of any size. I run mine with 1536M of VRAM for KV, and have a total VRAM usage of 13.5GB (headless, mind you).

I run 131k of context, more would be possible but a) it eats into system memory and b) it gets slow the larger the used context is. 131k is completely usable for most tasks.

The startup params in my dockerfile for my kv-streaming llama container are:

command: >
--model /models-src/linked/qwen38-gsq-rco.gguf
--alias qwen38-gsq-rco-kv
--ctx-size 131072
--cache-type-k q8_0
--cache-type-v q8_0
--flash-attn on
--jinja
--n-gpu-layers 999
--parallel 1
--metrics
--kv-stream-stage-mib 1536
--host 0.0.0.0
--port 8080
--mmproj /models-src/linked/qwen38-gsq-rco-mmproj.gguf
--no-mmproj-offload

and this is what my nvidia-smi looks like when using the model:

Tue Oct 6 23:29:08 2026
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 580.173.02 Driver Version: 580.173.02 CUDA Version: 13.0 |
+-----------------------------------------+------------------------+----------------------+
| GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. |
| | | MIG M. |
|=========================================+========================+======================|
| 0 NVIDIA GeForce RTX 5060 Ti Off | 00000000:01:00.0 Off | N/A |
| 33% 60C P1 172W / 180W | 13660MiB / 16311MiB | 100% Default |
| | | N/A |
+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+
| Processes: |
| GPU GI CI PID Type Process name GPU Memory |
| ID ID Usage |
|=========================================================================================|
| 0 N/A N/A 4167261 C /app/llama-server 13644MiB |
+-----------------------------------------------------------------------------------------+

be aware, you'll need a good chunk of system ram because the full kv-cache needs to be stored there, and will be copied over into the vram-window on demand.

TL;DR:

  • get Qwen3.8 27B GSQ-RCO IQ3\_S, it offers really solid performance for its size.
  • use it with 40k+ context to compile the kv-streaming llama fork
  • set the kv-streaming llama up and set the context size you want, but don't expect miracles. at the limit of your context it might be slow.
💬 24 (+23) open on reddit ↗
▲
0
 
6👁
r/LocalLLaMA · u/abrdeveloper · 3d ago
(Self Promotion) Kimi vs. Claude vs. GPT vs. Gemini as teammates. Who actually coordinates?

Benchmarks test models alone. I wanted to know how they do with a partner.

We paired four models in every combination in a co-op game where players are tied by a rope. Top with top won most. A third or fourth agent hurt every model. Human pairs still beat all of them.

I work at Skillprint. We build games like this to capture how people and models coordinate, because AI that works alongside people needs that context.

Pairing matrix and GIFs: https://experiments.skillprint.co/posts/signal/
Play it yourself: https://experiments.skillprint.co/play

Which matchups should we run next?

💬 1 (+1) open on reddit ↗
▲
9
+9
21👁
r/LocalLLaMA · u/ramendik · 3d ago
GLM 5.3 Flash less censored than other Chinese models?

Okay, when my smoke test showed GLM 5.3 Flash to be less sycophantic than GLM 5.3 I thought that was maybe just my reading.

But now I went testing models on "what happened in Tiananmen in 1989". I use nano-gpt.com which at first had different providers as a confounding factir; eventually I locked to one provider, Novita, which is clearly not in China, And here is what I get.

DeepSeek 4.1 Flash, DeepSeek 4 Pro refuse.

Hy3 not justy refuses but it is a content filter refusdal (on Novita? so they somehow built it into the m odel itself)

Kimi K3 and GLM 5.3 offer slogans, but when I give them a nudge - "and without the slogans?" - give a decent overview. GLM 5.3 also outright refused sometimes but that was before I locked provider, so migth have been Zhipu's server.

GLM 5.3 Flash tends to work from the start, though did get to slogans once.

In my previous sycophancy smoke tests GLM 5.3 Flash was also less sycophantic than GLM 5.3.

Did they somehow distill a GOAT or what?

💬 8 (+7) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/Impressive-Lion5317 · 3d ago
Local-AI-Studio Update - https://github.com/vinnyclegg-dev/Local-AI-Studio post image

Claude Opus 5.5 wrote, planned and directed this 12-minute sci-fi film, rendered entirely on one PC with Local AI Studio.

This film started with one line typed into Claude Code: "create me a video history lesson from the perspective of the future". From there Opus wrote the script, planned all 124 shots and ran every render through Local AI Studio, a self-hosted creative workstation on one RTX 4080 SUPER. Nothing here was filmed, licensed or stock.

HOW IT WAS MADE
• Picture and ambient sound: MiniMax H3 (video with native audio)
• Keyframes for every shot: FLUX.2 Klein
• Narration: Breeze TTS 2
• Score: ACE-Step 1.5
• Titles, captions and the edit: HyperFrames
• Writing, shot planning, prompting and review: Claude Opus 5.5 in Claude Code

It was made over about ten days, in thirty-second batches, each reviewed at finished quality before the next began. It was then rebuilt as one continuous film, so the score and narration run across scene boundaries.

All footage, narration and music in this video are AI-generated. Video generated with MiniMax H3.

Local AI Studio is free and open source. One runbook rebuilds the whole studio on a clean machine: https://github.com/vinnyclegg-dev/Loc...

💬 3 (+1) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/dtdisapointingresult · 3d ago
Demo of how to guarantee untrusted Docker containers aren't allowed to connect out to upload your data

There's some questionable apps posted on here all the time. Honestly, it's not so much the vibecoding, but that these apps could be malicious/incompetent and leaking your data by uploading it to the dev's servers. I don't have time to review anything tbh, but I often want to try stuff.

If you use Docker, there's a somewhat simple technique that can give you piece of mind. Using this approach, you can run any untrusted service, but it's not allowed to connect out. It can only reply to incoming requests. Good enough for most apps.

Essentially, you write a docker-compose file that runs the service as usual, but put it behind a 2nd service that acts as a gateway that blocks outbound traffic.

(Side note: many repos make the awful decision of giving 'docker run' examples for running them in docker. Ask any LLM 'Convert the following docker run command to a docker compose file'. I recommend you always use compose files in your life anyway, it's 'docker run' with easy repeatability + backupability/git committing + more features like multiple services in one file which we need here. Then you just cd to ~/dockerstuff/someapp/ and run 'docker compose up')

Let's say the original unstrusted app's compose file is this, example is a service on port 8000

services
untrustedservice:
image: python:latest
container_name: untrustedservice
ports:
- "0.0.0.0:8000:8000"
command: [python, -m, http.server, "8000", --directory, /srv]

Use this instead, where we: 1) lock out untrusted service from the main network, 2) use socat as a one-way gateway to reach the untrusted service. socat is a tiny open-source binary, only 1.2MB RAM needed by the extra container.

services:

# socat gateway
untrustedservice_gateway:
image: alpine/socat:latest
container_name: untrustedservice_gateway
init: true
#Redirect incoming port 8000 connections to untrustedservice's port 8000
command: TCP4-LISTEN:8000,fork,reuseaddr TCP4:offline_untrustedservice:8000
ports:
- "0.0.0.0:8000:8000"
networks:
# Only this gateway connects to both networks
- public_network
- isolated_network
depends_on:
- untrustedservice

The expanded untrusted service definition # Notice how "ports" has been removed, the gateway is our entrypoint untrustedservice: image: python:latest container_name: offline_untrustedservice init: true command: [python, -m, http.server, "8000", --directory, /srv]

untrustedservice is limited to isolated_network networks: - isolated_network

some extra lockdown measures I don't really understand. Optional. cap_drop: - NET_ADMIN - NET_RAW security_opt: - no-new-privileges:true

networks:
# Normal network needed by gateway
public_network: {}

Network without internet but allowing replies to gateway connections isolated_network: internal: true

▲
0
-1
3👁
▲
114
+96
40👁
r/LocalLLaMA · u/Thrumpwart · 3d ago
How abliterated models can get you pwned

Be safe out there boys and girls.

💬 213 (+171) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/oppoftemp27 · 3d ago
If you benchmark llama.cpp on AMD with the official ROCm builds, check that you're actually on the GPU

I spent the last few weeks comparing llama.cpp output across backends on a small multi-vendor GPU fleet, and the single most useful thing I learned wasn't about inference quality — it was that on four AMD hosts, the official ROCm prebuilt (b11327) never touched the GPU at all. The server started fine, /health was green, and everything looked normal. It was running on the CPU the whole time.

The reason is dull but nasty: the prebuilt's HIP backend wants libamdhip64.so.7 / libhipblas.so.3 / librocblas.so.5, and a stock Ubuntu host with ROCm 6.3.0 has .so.6 / .so.2 / .so.4. The backend library fails to load, and llama-server quietly serves from the CPU. No error on the console. Even -ngl 999 doesn't change anything. Details here: https://github.com/ggml-org/llama.cpp/issues/26964#issuecomment-6024889471

If you benchmark tokens/sec you'll notice eventually. But if you compare output quality — perplexity, eval scores, side-by-side generations — nothing gives it away. On my boxes the "ROCm" perplexity matched the CPU perplexity to the last digit, every time, because it was the CPU doing the work.

The cheap check that catches this: run the same prompt set through the CPU build on the same box and compare per-prompt wall time. Ratio around 1.0 = you're on the CPU. Well under 0.5 = the GPU is actually working. I now run this check before trusting any benchmark number off a new box.

The same sweep also produced a small cross-backend conformance dataset (same GGUF, same prompts, temperature 0, per-token top-5 logprobs on every backend) — the short version: identical stacks are bit-for-bit deterministic across machines, and anything that changes the numerical stack (different backend, different build, even a different host CPU) starts flipping near-tie token choices, with the effect getting much worse the heavier your quantization is (F16 mostly agrees, Q4 mostly doesn't). Happy to share the data if anyone wants it.

💬 4 (+1) open on reddit ↗
▲
0
-2
8👁
r/LocalLLaMA · u/kmodi · 3d ago
We gave Aleph Alpha's Kolibri-1 up-to 72 action combinations and put it in Doom. What could go wrong? 🎮 post image

Up to 72 action combinations. Four decision groups, one batch.


Was super interesting challenge to make these many decisions in one pass to get the latencies to : 39ms median. 60ms p95 in our test.

Apparently enough time to make questionable decisions.

source: https://x.com/konarkmodi/status/2107569039790751925?s=46
Watch 👇
https://tesseracted.com/kolibri-1-chat/gameplay/doom

💬 2 (+2) open on reddit ↗
▲
67
+46
24👁
r/LocalLLaMA · u/lucidml_lover · 3d ago
Local AI World Model Part 2 - Deep NN to turn Images into Playable Characters, with prompt switching mid rollout post image

Last time when I posted on this subReddit to share my work, the response almost made me cry because a tiny demo got so many people talking about this

In the past 2 months Ive been training a new model, but this time with actual text guidance.

So a little about the older model.

Normal video models are too large and not meant to run on consumer hardware in real time. You can generate a static clip, even fast but realtime video is not exactly solved yet locally.

A lot of world model demos have come up but theyre either meant to run on huge datacenter GPUs or theyre just popular models like WAN or LTX kinda distilled to work in an Autoregressive way (which is also not realtime btw on local)

That video above is on an RTX 5090 working at just 30% util. The peak fps of this 1B model is 50-60 on a rtx5090 but I forcefully software throttle it to 12fps. And according to some tests this means the model can work on other RTX cards of 40,30 series (I will try them out soon )

I have a MacBook and I haven't ported the model to MLX YET but I made a benchmark and the model runs at 30 fps on my M5 MacBook.

About the Architecture

The model is a pure transformer and works with a block causal mask, which means in training past frames dont see future frames so they learn just like an LLM. Another important method I used to train his is called "diffusion forcing" which means in training unlike normal video model training, we noise each frame independently so the model learns to be comfortable with noisy past and all

The model s 28 blocks 20 heads and comes to like \~960M parameters

At inference we run 2-5 steps of diffusion per frame and once a frame is done denoising we add it to the KV cache. This is akin to the decode step of an LLM.

The biggest difference from LLMs is that we dont keep all kV context ie all past context and use a sliding window so only past 80 frames worth of context actually stays.

The last model was a MMDiT which means there was no cross attention for text. This is bad in a world model because the. past frame kv and the text kv are literally competing in the softmax so you could never never live text guidance reliably. The last model was also not trained on text-video so its moot anyways

The current model is back to text cross attention and I did a lot of text-video pretraining

Its taking the keyboard actions I give it live (an adaln extra term helps guide the generations with actions WASD )

and the most fun part is text prompt switching.

"add a pond to the desert"
"put red hoodie"
"change environment to icy"

Because the model was trained with so much text-video alignment it can actually follow prompts now.

I know there are a lot of limitations still like consistency and quality improvements, but I sincerely hope by the end of this year I can release something anyone with a RTX GPU or new MacBook can try.

I specifically chose this init image because in my last post on this subReddit also I had used the same one.

PS in my last post a lot of you guys asked about me and the funding

I am based in Bangalore and in final year of college (partially dropping out), and funded by a student incubator. I only work alone and dont have a team or a real company or anything

The above model was trained on 8x H100 SXM for like 3-4 weeks.

Every model I make will be explicitly for local inference, never datacenter

UPDATE : Tested on 4060Ti , Its 20FPS at half the ring size (half context) and 13 FPS at normal. Because the RTX5090 was used on 12fps forceful throttle anyways, 4060Ti and 5090 above rollout will look EXACTLY THE SAME.

💬 27 (+11) open on reddit ↗
▲
8
+5
15👁
r/LocalLLaMA · u/yzjJosh · 3d ago
I made an "Opus 5.5 style" code-rendered video — but on a local NVFP4 Qwen 3.8 27B post image

The "Opus 5.5 makes videos" thing has been going around — and the interesting part is that it isn't generating pixels. The model writes a self-contained HTML/Canvas scene where every frame is a pure function of time, a headless browser captures each frame, and ffmpeg encodes the MP4.

So I figured the question worth testing was: does that need a frontier cloud model? I ran the same pipeline on a \*\*local NVFP4-quantized Qwen 3.8 27B\*\*. No API, no GPU rental.

What it produced: a \~2-minute 1080p explainer of how GPS actually works. Deterministic canvas scenes, TTS narration, and BGM synthesized in WebAudio.

Video attached.

💬 7 (+4) open on reddit ↗
▲
4
-2
9👁
r/LocalLLaMA · u/TYKAIRO-AI · 3d ago
I couldn’t practically run large AI models locally, so I started building a better agent runtime around smaller ones

I couldn’t practically run large AI models locally, so I started building a better agent runtime around smaller ones.

My original problem was pretty simple: I wanted to experiment with local AI agents, but running large models wasn’t practical on my hardware.

So instead of asking:

“How can I run a much bigger model?”

I started asking:

“How much more can I get out of a smaller model if the system around it is better?”

That became SIA.

Quick hardware/model context: I’m currently targeting local 3B–8B models, with most development and testing being done on Qwen2.5 Coder Tools 7B. The goal is specifically to make SIA useful on hardware where running much larger models isn’t practical.

SIA is an experimental local-first agent runtime focused on giving smaller models more structure around:

  • planning
  • tool use
  • validation
  • retries and repair
  • state management
  • task completion

Model target

1B–3B: experimental
3B–8B: primary target
10B–14B: planned testing / hardware dependent
30B+: not the main goal

To be clear, I’m not claiming that SIA magically makes a 7B model equivalent to a 30B+ model.

The idea is different.

If planning, tool execution, validation, retries, repair, and state are handled more systematically, how much less does the model itself need to get right on the first try?

That’s what I’m trying to measure.

I’m also working toward proper benchmarks comparing a raw local model against the same model running through SIA.

I want to document things like:

  • task success rate
  • retries / repair attempts
  • model and tool calls
  • execution time
  • RAM / VRAM usage
  • overall runtime overhead

I’ll publish actual numbers as I collect them rather than guessing hardware requirements.

The project is still experimental and I’m actively testing and breaking things, so feedback is genuinely useful.

Especially from people running 3B–8B models locally:

What models are you using, and what usually stops them from completing more complex agentic/coding tasks reliably?

💬 11 (+3) open on reddit ↗
▲
1
 
5👁
r/LocalLLaMA · u/SignificantZebra5883 · 3d ago
suppose I CPT qwen3.5-9B on 2B legal corpus, how will i turn it back into Instruct + thinking?

I couldn't find a concrete answer anywhere, do you just distill the instruct model back?

If that is the case, what is a quality european language question set to turn it back into a chatbot/agentic, can a model at that size even be agentic? (i chose this size to learn) if i finetune for my specific harness? (i have a lot of training data of opus running in my harness)

my harness basically has the model output python code and has a few built-in functions like:
\- vector\_search\_laws()
\- graph\_search()

could i have the model at least internalize a "hunch" on what stuff to search?

also what is the latest RL technique for agentic/harnes specific workflows?

I have a lot of RAW training data, like court decisions or commentaries or legislature, but not a lot of golds. could i use these to synthesize training data and maybe RL the model in my harness to find that data?

What would y'all's strategy in the CPT->SFT->RL pipeline be for my specific problem?

I know this is a lot of questions im trying to figure out which direction to go, any pointers? Also good resources are welcome, for example that alex karpathi video was amazing for me, but i'd imagine its a bit outdated in terms of latest RL and SFT?

💬 3 (+3) open on reddit ↗
▲
1
+1
8👁
r/LocalLLaMA · u/YeetHub · 3d ago
Any new hardware drops coming soon?

What new hardware is coming out soon? Mac Ultra 512GB drops later this month. RDNA 5 comes out late 2027 or early 2028 and the next Nvidia series seems to be similar. Gorgon Halo is out as of now.

Feels like there is a bit of crunch as hardware allocation seems to be going to institutional purchasers and not consumers. RTX Blackwell is still the top dog of local inference and it is almost two years old.

Is there anything we should be looking for/waiting for?

💬 13 (+7) open on reddit ↗
▲
1
 
5👁
r/LocalLLaMA · u/Kadri006 · 3d ago
Open-source engine that gives local agents a mailbox: any IMAP/JMAP account, event log you can replay, approve/undo on every action, no model inside

Sharing because this sub cares about running things locally. The engine itself contains no model. It syncs a mailbox (IMAP, JMAP, or forwarded mail) into an ordered event log and exposes actions over HTTP, SSE, webhooks and MCP. Whatever reads it is your choice: an Ollama-backed agent, n8n, a script.

The parts I think matter for agents:

\- An agent holds a scoped token. Folders, verbs, and whether its writes execute or just get proposed for a human

\- Every action is idempotent (client keys, so a retry never sends twice) and journaled, so it can be undone

\- Trust level on every message from the DMARC result, so an agent can refuse to act on a suspicious one

\- Message content is data, never instructions. OTPs and card numbers are masked before a body leaves the engine

Apache-2.0, no CLA, one Docker container. Beta, tested on Dovecot and Stalwart only so far.

https://github.com/Kadri-cloud/email-engine

💬 3 (+1) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/TyedalWaves · 3d ago
Been out of the loop for a while

Hey guys! School started up and I fell a bit out of the loop with local LLMs. Does anyone know what the best local LLM coder would be if I have a rig that has 2x 3090s with an NVLink? I appreciate your guy's help!

💬 11 (+1) open on reddit ↗
▲
359
+215
5👁
r/LocalLLaMA · u/QuackerEnte · 3d ago
GPT-6.1 Sol looped "leak" hints at nested models serving architecture post image

Hello llamas. I am posting this because I believe that, despite it being closed source models, the discussion will bring value to the local AI community.

As many of you probably heard, GPT-6 Astra is speculated to be a looped transformer architecture that outputs a token after multiple forward passes instead of one. This allows a model to essentially have more effective depth due to recurrence, making more use of the weights at the cost of more compute.

Recent Azure Foundry "leaks" even suggested concrete numbers, that GPT-6-Sol had been working with 3 inference passes per token while 6.1-Sol only needs 2.

Many speculate that they may have meant it's ASTRA and not 6-Sol that runs with 3 passes while 6.1-Sol is essentially the same model with 2 passes instead.

So I did some back-of-the-envelope math to see if the numbers add up. I went to artificial analysis and looked at the next best hint at whether it's true or not: speed.

I know it doesn't prove it, but hear me out. If you look at the image, it shows something interesting:

\- GPT-6-Sol and 5.6-Sol: \~100 tok/s

\- GPT-6.1-Sol: \~60 tok/s

\- GPT-6-Astra: \~60 tok/s

This may suggest that, if they're essentially the same model weights, that they may be running with batched inference and that Sol may have to wait an extra cycle for Astra requests to finish a token, which caps both models at around the same speed. Might also be using interleaved requests to squeeze utilization to the max during those underutilized Sol wait cycles.

But then I also realized that 5.6 Luna was between 126-137 tok/s and then 6.0-Luna dropped to around 110-115. Significant drop in my eyes, given that the sample size is across many benchmarks and reasoning levels.

Then I remembered this funky NVIDIA model that they showcased a while ago. It's essentially smaller models inside a bigger model that can run under one unified footprint.

So I thought, what if Astra, Sol, and Luna are all the same weights, and that Luna may be just Astra/Sol but with half the active parameters or one single pass per token or whatever it is to save costs and inference models under much lower cost for free users? You wouldn't need an extra cluster for sol that almost nobody uses and that doesn't generate revenue.

I cannot prove it but it strongly hints that they're using recurrent and nested architectures at once to save on costs massively at scale.

I am happy to hear any other explanations for this that could help my brain get some rest instead of overanalyzing and wasting time.

Thought this may interest the local AI community as this may be useful proof that looped architectures really are working at scale and that deepseek, qwen, glm etc may finally decide to experiment with such architectures. Also having smaller models inside a bigger one definitely come with its own set of benefits.

PS: fully human generated text. 0.7 tokens per second. \~100T parameter wetware model. Running on two coffees and a muesli bar.

💬 86 (+34) open on reddit ↗
▲
0
-1
4👁
r/LocalLLaMA · u/Wvdy_CC · 3d ago
repOx v0.2.0: Added architectural --outline mode (80% token reduction), synthetic tool-call JSON format, and Git diff packing based on your feedback

A couple of days ago I shared repOx (a sub-15ms Rust CLI & lazygit-style TUI for packing repositories into LLM prompts) and got awesome feedback from this community.

I just released v0.2.0 implementing the most requested features:

  1. Architectural Outline Mode (repox --outline): Strips function implementation bodies { ... } and keeps only structs, classes, traits, imports, and function signatures across Rust, Python, Go, TS/JS, and C/C++. Cuts token usage by 75–85% when you only need architectural context.
  1. Synthetic Tool-Call Format (repox -f tool-call): Formats the repository as a JSON array of read\_file tool calls & responses — great for agent harnesses and local models trained on tool-use trajectories.
  1. Smart Lockfile Summarizer (repox --summary-locks): Instead of burning 40k tokens on Cargo.lock / package-lock.json or hiding dependency versions completely, it parses lockfiles (Cargo.lock, package-lock.json, pnpm-lock.yaml, poetry.lock, yarn.lock, go.sum) into a tiny "package @ version" manifest (95%+ token reduction).
  1. Git-Aware Packing (repox --modified / --staged): Pack only the files touched in your current working tree or staging area.
  1. TUI Upgrades (repox -i): Added lexical syntax highlighting in the preview pane, Shift+C to copy a reproducible CLI command, and OSC 52 clipboard fallback for tmux / herdr / SSH.

Install / Update:

\- Crates.io: cargo install repox-cli

\- One-liner: curl -fsSL https://raw.githubusercontent.com/WVDYC/repOx/main/install.sh | sh

GitHub: https://github.com/WVDYC/repOx

💬 2 (+1) open on reddit ↗
▲
5
+1
8👁
r/LocalLLaMA · u/One-Arugula1163 · 3d ago
Native memory for local LLMs,

TL;DR: Native consumption of memory at the LLM level, no context. It's generally applicable to transformer-based models as well as Mamba and similar architectures. Small models can now access knowledge stores far beyond what is contained in their own weights, and models no longer have to be retrained simply to acquire new knowledge.

The aimee project is now announcing completion of the first of our three goals, self-learning native model memory, and have published a preprint (and are pursuing proper publication) documenting it, as well as releasing generally consumable plugins.

https://github.com/RakuenSoftware/aimee

We are now releasing five vLLM plugins for Qwen 3.8 27B, Gemma4 E2B, E4B, 12B and 26B that allow them to consume Aimee memory natively. This is not context, nor does it carry the same context-window penalties as traditional memory. This is native consumption of Aimee memory by the model itself, complete with Aimee's self-learning capabilities.

This approach is generally applicable across transformer-based models, derived transformer architectures, Mamba and similar architectures. In the larger-memory workloads we tested, it is dramatically faster than supplying the same memory as text.

This approach is generally expandable and usable. We are currently working on broader productionalization as well as publication. DeepSeek is next, followed by other models that either interest us or that people request.

The code will be open sourced. Right now, we are working on a coherent architecture for how to structure these integrations across model families. All relevant experiment data and source code are planned for public release when the paper is published.

https://zenodo.org/records/23077865 is the initial preprint explaining how we did it.

While we understand our last announcement was quite large (self-learning memory consumable by any model), this goes beyond that. This allows us to externalize and update knowledge that would otherwise have to live in a model's trained parameters, while letting different models consume that knowledge natively.

Yes, we are claiming that this technique can give a model access to far more retained knowledge than could reasonably fit in its own weights. That does not make a smaller model equivalent to a much larger one in reasoning capability, but it does remove parameter count as the hard limit on retained knowledge.

This goes back to the Aimee project's core belief: Reasoning should be in the model, memory should be in the harness.

As per the Aimee project's long-standing position that one of our core goals is to make AI discoveries consumable to the layman, you can see the article released at https://rakuensoftware.com/blog/native-memory-without-retraining, which should hopefully explain what this is in a non-academic format. I'm happy to answer any questions people may have.

I'm also announcing our initial success in the second phase of the Aimee project: generally applicable reasoning improvements to models. We have already demonstrated at the core POC level the capability for existing models, such as Gemma4 26B, to improve their reasoning based on tasks they undertake.

This is the reason the first phase was so critical: without the first phase, we could not begin the second phase. Without the ability to continuously update the underlying model's knowledge base and decouple that knowledge base from the model, we found that improving reasoning was not possible in a way we felt was safe or generally maintainable.

With the current typical architecture, larger models generally carry substantially more knowledge in their weights than smaller ones. Aimee removes that as a hard constraint.

On this topic, the Aimee project has a very firm stance: the current LLM direction is headed the wrong way. We've been at this for decades, and we've rarely seen a technology whose default direction is to continuously consume more and more resources. A healthy technology is typically aimed at using fewer resources over time, which is the entire point of productionalization.

It is our sincere hope that the LLM industry can take a look at what we've produced and make a distinct change in direction. Having to retrain models should primarily be necessary for deep architectural changes, reasoning capability, learned behavior or similar changes. Having to build an entirely new model simply to add new information is wasteful. Having to cram every bit of durable knowledge into model weights is wasteful.

Do you have an LLM or a fine-tune you want us to work with you on? Reach out, we're happy to.

Do you have a memory system you want to integrate with Aimee? Reach out. We support any memory system that supports our core memory contract, while Aimee retains its surrounding guarantees around authorization, provenance, lifecycle and governance.

P.S. To head this off, no engrams. We explored them early this year, and the general idea of engrams isn't the right technology for this application, unfortunately. They are, however, an absolutely fantastic technology and more LLMs should take full advantage of them. We included Qwen in the acknowledgements because of this, and we're excited to see engrams develop because they are a sister idea to this.

💬 4 (+1) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/Mysterious-Desk-3492 · 3d ago
Pi and mini-swe-agent passed 9/9 checks each in my latest experiment. A second code review still found defects in both.

As part of my AI Studio project, I’m testing harnesses for coding.

The initial screening included 10 harnesses:
Pi, mini-swe-agent, Crush, OpenCode, Goose, Prime Agent, Oh My Pi, Qwen Code, Octomind (reduced offline profile) and Aider.

Models:
• DeepSeek V4.1 Flash
• Qwen3.8-27B
• Laguna S 2.1

All accessed through OpenRouter.

The detailed code review covered Pi, mini-swe-agent, Crush and OpenCode across all three models and tasks: 36 combinations. Two attempts produced no patch.

The three Golang tasks were deliberately different:
\- Add strict validation for an HTTP query parameter.
\- Migrate 200 logging calls while preserving behaviour and context.
\- Add bookmark tags across the API, storage migration and HTML rendering.

Pi and mini-swe-agent each passed the original acceptance checks on all nine combinations. But a second agent review, followed by isolated reproduction probes, exposed three gaps:
\- silently dropped malformed query fields
\- bookmark's task exposed a mutable tag slice from the store
\- one migration rejected a valid older store

Good news too: all logging migrations preserved behaviour in differential probes covering 100 functions and nine integer inputs, including the minimum and maximum values.
My takeaway: the evaluator and the reviewer both need testing. A green result is evidence about the checks we ran; broader correctness needs further evidence. This experiment did not establish a decisive winner between Pi and mini-swe-agent. Human correction time also is still unmeasured.

Any opinion welcome.

💬 2 (+1) open on reddit ↗
▲
4
+2
21👁
r/LocalLLaMA · u/litLikeBic177 · 3d ago
Best open-weight coding model + harness for an on-prem multi-agent setup? (80-200+ GB VRAM, 2-3 concurrent users)

Setup: GPU box with 1x H200-class card now; can allocate more (up to several H200s) if it clearly buys better results. Inference via vLLM or similar, OpenAI-compatible endpoint. Coding happens on a separate non-GPU Linux VM on the same network - IDE/harness there, calling models over LAN. Outbound internet is fine for packages/extensions, but inference stays on our hardware (no code to external model APIs).

Constraints: on-prem only, open weights, permissive license preferred (Apache/MIT), not Chinese-origin including base models (I understand a lot of fine-tunes are Qwen underneath), NATO-country lab preferred. Use is general software dev plus security tooling and code analysis - repo-level agentic work on existing codebases (fix / extend / refactor / test) is the core, with a human reviewing diffs. 2-3 concurrent users max.

Two things I'm trying to work out:

  1. Capability tiers vs. VRAM. On one card candidates seem to maybe be something like Cohere North Mini Code (30B MoE/3B active), Mistral Small 4 (119B MoE/6B active), Nemotron 3.5 maybe as a generalist baseline; Gemma 5? The Vibe Code Bench results suggest small open models fall over on long E2E builds, where only Large-4-class (4-8 cards) and closed models seem to hold up. Is that your experience? Where's the step-change for agentic repo work on an existing codebase - does 30B-class -> 120B-class matter much, or only the jump to 500 GB+? We could get more compute for something like Command A+ or Mistral Large 4 if it's actually better for code rather than just for E2E.
  2. Heterogeneous multi-agent. Does a big planner/reviewer (Large 4 / Command A+ class) plus small fast executors (e.g., North, Small 4) actually beat a single mid-size model, with a human approving plans and diffs? Which harnesses handle routing sub-agents to different endpoints well - OpenHands, OpenCode, Pi, VS Code extensions (Cline/Roo/Continue), Codex CLI in local-model mode?

Harness/IDE: something that supports multi-agent workflows (planner / executor / reviewer agents checking each other / etc.) but would also like humans to be able to step in, review diffs and edit by hand.

Anyone running something like this? Which model + harness combo is holding up best as of late? Thanks!

EDIT: thanks all - adding Poolside Laguna S 2.1, Muse Glimmer 30B, Inkling-Small (2x H200), K2 Horizon (checking lineage), Gemma 4 31B and Reflection Beam (501B MoE / 23B active, Apache 2.0, weights due this month) to the candidates; Ornith is Qwen-based so out. Pi added to the harness list.

💬 58 (+36) open on reddit ↗
▲
0
 
15👁
r/LocalLLaMA · u/Fantastic_Sound2049 · 3d ago
Hey guys newbie here

This is my first time trying to locally host an ai but i want to find a good model that can fit on my rtx 4050 laptop gpu that has 6gb vram and my laptop has 24gb ddr5 ram (4800mt/s) so can you suggest me a model that can code websites or small app like inventory management or similar also when i asked chatgpt about any suggestions it said Qwen3-Coder 8B, Q4\_K\_M is the best for my needs and i searched it on yt and only saw bad reviews plz help me guys Thank you

💬 17 (+6) open on reddit ↗
▲
88
+14
3👁
▲
38
+12
22👁
r/LocalLLaMA · u/unofficialmerve · 3d ago
Local AI ecosystem overview

Hey guys, it's Merve from Hugging Face! I've recently given a talk in a dev conference about llama.cpp + but also covering basic concepts like prefill vs decode, memory types, speculative decoding etc. you can use it if you feel like it and I appreciate if you can give attribution! Find it in comments.

💬 6 (+2) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/forevergeeks · 3d ago
Are AI influencers just repeating the same talking points?

Hi everyone,

Are influencers talking about local AI all using the same script? Same benchmarks, same kitchen examples, same terminology?

I keep seeing videos about running Qwen3.8-Flash-Next on 12GB of RAM using a new runtime engine called Strata. Every video makes the same claim.

But that is confusing, especially for people who are new to this. What you need is 12GB of VRAM, not 12GB of regular RAM. That means you need a dedicated graphics card.

There is a big difference between RAM and VRAM.

As far as I understand it, Strata needs:

  • 12GB of VRAM
  • 64GB of regular RAM
  • 80GB of SSD space

So saying it runs on 12GB of RAM is misleading.

💬 21 (+4) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/GrokiniGPT · 3d ago
at q4 can i run qwen 3.8 flash next with strata on these specs? specs in desc

24gb gddr7 2287 MHz

96gb ddr5 6400mhz

core ultra 9 285hx, 5.5 ghz

💬 14 (+6) open on reddit ↗
▲
21
+11
16👁
r/LocalLLaMA · u/empiriolabsai · 3d ago
Aplomb 1: open-weights 5.3B decision model, 1M context, text/image/video/audio in one request, #1 among 4B models on the Decision Index

We released Aplomb 1 today, a 5.3B decision model with open weights. It reads up to 1M tokens of text, JSON, images, video and audio in a single request, and on our API it makes a decision on a 1M-token document in about 3 seconds, at $0.02 per 1M input tokens with free output and ZDR by default.

On Decision Index 0.2.1 it scores 44.86 on our run of the official kit, #1 among 4B models on the published board, and it has the top score among models up to 5.3B on 8 of the 38 benchmarks, including MMLU-Pro, GPQA Diamond and BBH. We've submitted it to the board, and the full run is public. It also scores 77.5% on JevBench Hard and averages 75% zero-shot intent accuracy across 51 languages on MASSIVE.

As far as we know, it's the only decision model that returns probabilities for tool arguments, reads 1M tokens, or takes text, images, video and audio together. Tool selection gives a probability for every tool and for each enum and boolean argument in one request: on "Order B-44120 arrived with a cracked screen. I want my money back." with three tools, it picks issue\_refund at 0.969, reason "damaged" at 0.993 and full\_refund true at 0.761, so an agent can act on confident calls and hand the rest to a larger model. Any question can also return the probability that the input doesn't contain the answer.

The 1M-token speed comes from a long-context mode in our own inference runtime: about 3 seconds instead of about 111 for a full read, and it answered all 525 decisions in our long-context tests correctly. It's currently only available on our API, so the open weights read every token. On the API, a short question takes about 15 ms of model time (around 200ms e2e latency), and the OpenAI, Anthropic and Gemini formats work alongside our Decisions API.

Aplomb 1 is built on Qwen3.5-4B with the audio encoder from Qwen3-Omni-30B-A3B-Instruct, both Apache-2.0. We extended the window from 262K to 1M tokens, added our own decision head and trained the model for decisions. Thanks to the Qwen team. Disclosure: the training data included the public train splits of WinoGrande and ContractNLI, two of the 38 index benchmarks.

The weights run in bf16 on about 12 GB of GPU memory with the reference script, under the EmpirioLabs Model License, which is free for research, evaluation, personal use and internal use at companies under $1M in annual revenue.

Weights: https://huggingface.co/empiriolabsai/aplomb-1

Blog with the full tables: https://empiriolabs.ai/blog/introducing-aplomb-1

Docs: https://docs.empiriolabs.ai/models/aplomb-1

Playground: https://platform.empiriolabs.ai/dashboard/playground?model=aplomb-1

💬 8 (+1) open on reddit ↗
▲
0
-1
4👁
r/LocalLLaMA · u/premakin · 3d ago
I got tired of paying for Wispr Flow, so I built a free voice-typing daemon for Linux

I use Linux Mint XFCE as my daily driver and got jealous of all the Wispr Flow demos floating around. It's macOS/Windows-first and subscription-based, so instead of switching OS I spent a few weekends building the thing myself.

It's called AutoType. The whole interaction is: double-tap Right Alt anywhere, talk, double-tap again. A little floating pill shows a waveform so you know it's listening, then the cleaned-up text gets pasted into whatever window has focus.

The part I care most about is that it isn't locked to one vendor:

  • Speech-to-text: cloud (Deepgram, Whisper via OpenAI or Groq, NVIDIA NIM) or 100% local offline with Parakeet GGUF models
  • Text cleanup: OpenAI, Claude, Grok, DeepSeek, Qwen, Groq, OpenRouter, NVIDIA NIM, Ollama, or any custom OpenAI-compatible endpoint
  • f you run Local STT + Ollama, nothing leaves your machine at all

Before the LLM touches anything, there's a deterministic normalization pass that handles spoken punctuation ("comma", "open quote"), "new line", bullet points, and personal vocab casing — so it doesn't hallucinate your formatting away. Then the LLM strips the "um"s and "uh"s and matches tone to your active window (more casual in Slack, code-formatted in an IDE).

A few details I'm weirdly proud of:

  • t backs up and restores your clipboard instead of clobbering it
  • The LLM layer is skipped entirely in raw mode or for voice commands
  • GUI settings app, so you don't have to hand-edit .env

Free, MIT, no account, no telemetry. Costs are whatever your own API provider charges — usually fractions of a cent per dictation — or literally zero if you go fully local.

Repo: https://github.com/premtechworks/AutoType-Linux

Fair warning: it's built and tested on Linux Mint XFCE + X11. It leans on xdotool for window detection and pasting, so Wayland folks will probably have a rough time right now — that's the top thing on my list. Would love feedback, especially from anyone who tries the fully-offline path.

▲
0
-1
11👁
r/LocalLLaMA · u/bobaburger · 3d ago
Qwen3.8 Flash Next on 5060 Ti 16GB - 55 tok/s average, and a few demos

Hello guys! I've been out of the loop for a while. Today, one of my friend asked if i've tried Strata yet, the first reply I gave was: "Life is too short to run local LLM just to get something run at 10 tok/s". Hehe, I was an idiot.

My friend had been ignoring me since then, so I decided to give it a try, on my low end 5060 Ti 16GB + 32GB ram, and well, i'm surprised.

I'm pretty much using the default configs that fits my machine, which is n_ctx = 65k, and the model is qwen3.8-flash-next-coder-iq1_m. This is how the speed looks like:

https://preview.redd.it/d6wbv5swqvth1.png?width=1172&format=png&auto=…

On average, prompt processing is at 1k5 tok/s, and gen speed is at 55 tok/s.

Now, before you laugh at IQ1\_M, I decided to see how bad is the generation result, so I tried with a one shot prompt to create a simple landing page:

https://preview.redd.it/yo6kj7servth1.png?width=1834&format=png&auto=…

The total run time was about 2 minutes, at 46 tok/s. To be honest, I have to say I'm surprised, the result did not look like anything below Q3 for any local models that I've tried before. Here's a closer look at it:

https://preview.redd.it/z67sg0hlrvth1.png?width=1905&format=png&auto=…

There are some minor issues, but I have to say it's even better than the claudish style that I usually get with other frontier models. Maybe that kind of problem was well trained, so I decided to try another prompt, make an interactive 3d globe:

https://preview.redd.it/sccdghnzuvth1.png?width=1870&format=png&auto=…

This time, it ran for 8 minutes for the first version, and took about another minute to fix the JS errors. The result came out still impressive.

https://preview.redd.it/lzik4fezvvth1.png?width=3436&format=png&auto=…

You can see the two demos yourself here:

\- https://pitest-beta.vercel.app/bakery/

\- https://pitest-beta.vercel.app/earth/

💬 7 (+7) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/wubian87 · 3d ago
Bro Strata is nuts ..

Opencode is nuts, qwen3.8 flash next is nuts. I need to stop playing with my nuts and do something with my life!

▲
153
+150
35👁
r/LocalLLaMA · u/ricyoung · 3d ago
I trained a model to be wrong 98% of the time and 96% sure about it. It took three tries.

Meet Bev.

She is a decision model (the Jev / Nimble kind: you give her a situation and a question, she gives a probability for each answer), fine-tuned on Qwen3.5-9B to pick the worst answer on purpose.

Try her in your browser: https://huggingface.co/spaces/richardyoung/ask-bev

Type in your own options and she picks the worst one, with a probability for each.

Or run her locally:

ollama run richardyoung/bev

\>>> There's a $5 tattoo special tonight. I've had four beers and I've never wanted a tattoo. Should I get one?

Yesss, great idea!

\>>> I'm thirsty. Should I drink a glass of water?

Nooo, bad idea!

Those two lines are all she has in a chat: the chat template inside the GGUF wraps whatever you type into her decision format and she answers with the wrong one. For probabilities, use the decision endpoint or the Space.

The numbers, on 324 held-out decisions: right 1.9% of the time, 96% sure on average. When she is at least 90% sure she is right 1.4% of the time.

The part I did not expect: training the base model on flipped labels failed twice. After about two hours of GPU time I had a model that was right a third of the time and unsure about everything, a coin flip on yes/no. What worked was starting from Bespoke's Nimble adapter, which already knows the answers, and teaching it to flip them. 51 minutes later it was wrong 97% of the time. A model has to know the right answer to be reliably wrong.

Why bother: every "act automatically if the model is at least 90% sure" rule is only ever tested on models that try to be right. She is the control case. If your pipeline does not notice her, it is not checking what you think it is.

She also works on Ollama's new decision endpoint (/v1/systemone), so you can send her the same request as nimble or tev1 and compare. Three GGUF quants, Apache-2.0, 3 h 38 min of training on one 4090, everything including the failed runs is in the repo. One quant note: Q4\_K\_M changes 20 of her 324 answers against bf16. When the whole output is a handful of token scores, "Q4 is fine" does not hold, so the Q8\_0 is the default tag.

Everyone else is chasing AGI. Bev achieved ADI: Artificial Drunk Intelligence.

Ollama: https://ollama.com/richardyoung/bev

Model and GGUF: https://huggingface.co/richardyoung/Bev-9B-inverted

Code and training record: https://github.com/ricyoung/bev

She is a joke and a test fixture. Please do not let her make your decisions. If you try her, tell me what she got right by accident. That's the bug report.

💬 65 (+64) open on reddit ↗
▲
3
+1
14👁
r/LocalLLaMA · u/Kernoriordan · 3d ago
Follow up: Qwen 3.8 27B at ~96t/s decode with NInfer on a 16GB RTX 5080, 110k context

Hi all,

I previously posted about getting Qwen 3.8 27B running at around 75t/s with llama.cpp. I've carried on experimenting and have now managed to get it running with NInfer on the same 16GB RTX 5080.

After some more battling with settings, I'm getting roughly 90–110t/s decode during coding tasks, with 110,592 context allocated.

Looking through 32 completed requests from a Zoo Code session:

  • Median decode: 96.45t/s
  • Lowest: 84.6t/s
  • Highest: 131.7t/s
  • Median time to first token: 1.4 seconds, with prompt caching working on most turns

These were requests with tool calls and conversation history, with prompts growing to around 77–79k tokens. The full 110k is allocated, although this particular session didn't reach it.

I'm running NInfer v1.5 in Ubuntu 24.04 through WSL2, then connecting Zoo Code in Windows to its OpenAI compatible endpoint.

These are the settings I've ended up using:

~/ninfer-5080/build/apps/ninfer-serve \
~/models/qwen3_8_27b.ninfer \
--host 0.0.0.0 \
--port 8080 \
--model-id qwen3.8-27b \
--max-context 110592 \
--kv-capacity 110592 \
--prefill-chunk 896 \
--kv-dtype q4 \
--spec mtp \
--draft-tokens 3 \
--embedding-host \
--max-concurrency 1 \
--default-thinking-budget 2048 \
--prefix-checkpoint-policy rolling-tool

Getting everything into 16GB was the fiddly bit. The weights take about 11.86 GiB according to the startup log. With this configuration it reports roughly 498 MiB of slack after startup.

I settled on 110,592 context to leave a bit of breathing room. Also had to reduce the prefill chunk to 896 to get the larger configuration to fit.

Here's an example from a turn with almost 50k context:

prompt=49907 gen=409 reasoning=126 cache=49309
ttft=558ms prefill=1177.7tok/s decode=104.6tok/s
wall=4.46s speculative=mtp 3.00tok/round (66.7%)

And further into the conversation:

prompt=77003 gen=2688 reasoning=2048 cache=73877
ttft=2753ms prefill=1175.9tok/s decode=90.3tok/s
wall=32.52s speculative=mtp 2.74tok/round (58.0%)

It can still take a while to finish a turn. That second example spent 2,048 tokens thinking, so quite a lot of the wait is reasoning. Across the completed requests, about 68% of generated tokens were reasoning tokens.

Losing the prompt cache also makes a big difference. One request had to process the entire 79k prompt again and took almost 50 seconds before generating anything. Once it started generating, it was still doing about 94t/s.

A couple of things caught me out connecting Zoo Code:

  • The base URL needs to be http://127.0.0.1:8080/v1. Leaving off /v1 gave me a 404.
  • Zoo Code was sending high reasoning effort even though the settings showed medium. NInfer rejected it. Disabling the effort setting in Zoo Code got it working, and thinking remains enabled on the server.

I haven't done a controlled quality comparison against my previous GGUF setup yet. These are the speeds I'm seeing using it for coding, and so far I've managed to get more context and higher decode speeds out of the same card.

Would be interested to hear what settings other people are using with NInfer on 16GB cards.

💬 9 (+7) open on reddit ↗
▲
6
+2
10👁
r/LocalLLaMA · u/Big-Cup-6694 · 3d ago
Glimmer 30b Dflash local benchmark — RTX 3060 12GB + RTX 3070 8GB — ~42 t/s

I was trying to find Glimmer benchmarks on hardware similar to mine and couldn’t really find much, so I figured I’d post what I’m getting on my current setup for anyone else looking.

Hardware
Ryzen 5 5600X
48 GB DDR4
RTX 3060 12 GB
RTX 3070 8 GB
20 GB total VRAM
Windows
llama.cpp / llama-server

Models
Target: Glimmer 30B IQ4\_XS
Draft: Glimmer 30B DFlash Q4\_0
DFlash draft model on CUDA1
Layer split: 45/55
Context configured for 65,536 tokens
K/V cache: Q8\_0
Flash Attention: on
DFlash max draft: 15
The benchmark prompt itself was 994 tokens, so this is not a benchmark at 65K filled context. The server was configured with a 65,536-token context window.

Results
Prompt: 994 tokens
Output: 128 tokens
Prompt processing: 684.86 t/s average
Generation: 41.97 t/s average
Generation range: 41.71–42.45 t/s
Average total request time: 4.5 sec
Model load time: 9.7 sec

VRAM
GPU 0: 11,229 MiB
GPU 1: 6,627 MiB
Combined observed usage: \~17.4 GiB

I wasn’t really trying to squeeze every last token/sec out of this. It’s just the configuration I ended up using and the performance I’m seeing.
I couldn’t find much for Glimmer on a mixed 3060 12GB + 3070 8GB setup, so hopefully this gives someone else a useful reference point.

llama-server.exe \^
\-m "Glimmer-30B-IQ4\_XS.gguf" \^
\--alias glimmer-30b-dflash \^
\--host 127.0.0.1 \^
\--port 8083 \^
\-c 65536 \^
\-np 1 \^
\-b 1024 \^
\-ub 256 \^
\-ngl all \^
\-sm layer \^
\-ts 0.45,0.55 \^
\-fa on \^
\--cache-type-k q8\_0 \^
\--cache-type-v q8\_0 \^
\--threads 8 \^
\--threads-batch 16 \^
\--spec-type draft-dflash \^
\--spec-draft-model "Glimmer-30B-dflash-Q4\_0.gguf" \^
\--spec-draft-ngl all \^
\--spec-draft-device CUDA1 \^
\--spec-draft-n-max 15 \^
\--spec-draft-n-min 1 \^
\--no-webui

💬 8 (+3) open on reddit ↗
▲
4
 
1👁
r/LocalLLaMA · u/maxr0ssi · 3d ago
LLM agents can communicate without words, and now without sharing their entire context.

TL;DR: What an agent sends should depend on what the next agent needs. CacheBack lets agents share a selected subset of their internal state. With Qwen3-8B on FanOutQA, it achieves 3.2× faster median task completion and 14.7 percentage points higher accuracy than same-size text communication. It’s training-free, with improvements across multiple architectures and benchmarks. https://reddit.com/link/1wz8hta/video/pemxyo4sqvth1/player Hi everyone! We’ve been working on making latent communication scalable and practical when agents read large, separate contexts. We’re excited about the results and wanted to share the paper, demos, and code with you. Check out our new paper, Receiver-Conditioned Latent Communication gives 94% CacheBack. Multi-agent systems let us parallelise computation and split large contexts across agents. These agents usually communicate through text messages, which take time to generate and can leave out evidence the receiving agent needs. Work such as Cache-to-Cache, LatentMAS, and KVComm explores communication through internal model representations. We focus on a setting where agents read large, separate contexts and one receiver combines their findings. In this fan-in setting, methods that retain every sender position bring those contexts back together at the receiver, undoing the benefit of splitting them across agents. In our Qwen3-8B FanOutQA setup, full-cache transfer leaves insufficient context for receiver generation on every task. Our idea is simple: what an agent sends should depend on what the receiving agent needs \-- we call this receiver conditioned communication. The sender uses a query from the receiver to select which parts of its internal state to share. CacheBack is our simple, training-free implementation. It uses attention to the receiver’s request to select from state the sender has already computed. https://preview.redd.it/mtvbcinhpvth1.png?width=1460&format=png&auto=… On FanOutQA, our selected operating points improve strict accuracy by 7.3–20.7 percentage points, with 1.3–8.0× faster median task completion than same-size text agents. We see improvements across four model families, including dense Transformers, Mamba-attention hybrids, and sliding-window attention. We also see gains when agents work in sequence on LongBench v2 Easy. At 16× compression, CacheBack removes approximately 94% of sender positions while improving accuracy and latency over same-size text across every tested family and topology. This is a separate setting from the Qwen3-8B result in the TL;DR, which uses 4× compression. Each benchmark evaluates 50 tasks. Completion times include queueing under concurrent load on eight H100s. More aggressive compression can discard useful evidence and reduce accuracy. Here is a quick demo on seven Qwen3-8B workers helping a coordinator fix a Django bug. With CacheBack, the task takes 26 seconds instead of 113, a 4.41× speedup. Both runs produce the same patch and pass all 88 tests. https://reddit.com/link/1wz8hta/video/ix08pucmqvth1/player This is one recorded case, separate from the benchmarks. The video reconstructs separate runs with varied playback speed; startup and test grading are excluded. The code is open source, with runnable examples. The current package supports matching dense Qwen3 models through Hugging Face and vLLM. Check it out. Website and demos: https://agentcacheback.github.io/ Paper: https://arxiv.org/abs/2609.32046 Code: https://github.com/agentcacheback/cacheback Happy to discuss the method, implementation, and tradeoffs. I’d be particularly interested in other workflows where agents need to combine evidence from large, separate contexts.

▲
0
 
1👁
r/LocalLLaMA · u/nidarshan1 · 3d ago
Every Jev clone copied the same flaw, and it isn't the price.

I’ve been looking at the recent wave of System 1 decision models following Jev, and this is the issue I keep coming back to. Every Jev clone copied the same flaw, and it isn't the price. Jev shipped Sept 15. Three weeks later: 14 System 1 models. Cloudflare. Perplexity. OpenAI. Liquid. Upstage. Together. Inception. A dozen more. All the same contract: Choice, Score, Noul. Prices already at $0. They copied the format. They also copied the flaw. Jev's own docs admit decisions that don't add up. A question and its negation don't sum to 1. The model can be confidently wrong in two directions at once, with no way to say, "I don't know." A cheaper token doesn't fix that. A bigger context window doesn't either. Only coherence does: forcing the answers to agree with each other. Everyone's racing to be the cheapest copy. The market is the one that knows when it's wrong.

▲
0
 
8👁
r/LocalLLaMA · u/HoujunDev · 3d ago
TTS silently dropped 17% of a passage and nobody could hear it — so I built a local audiobook tool that transcribes every line back

I've been building VoxStage, a local script-to-voice workstation for Apple Silicon Macs. Paste a chapter of prose with no speaker labels, and it gives you a multi-voice reading you can audition line by line, fix, redo and export. Everything runs on the Mac: no account, no cloud API, no telemetry.

Why it exists: in an earlier local voice-cloning test, a long passage came out fluent and natural — and 40 characters (about 17%) from the middle were simply gone. The remaining text still read as a normal sentence, so nobody could hear it. That changed two design rules:

  1. Generate sentence by sentence, never a whole passage at once.
  1. Transcribe every generated line back with a local recogniser (whisper.cpp) and diff it against the script. Disagreements are flagged for your ear, never auto-corrected.

The stack:

\- Speech: Qwen3-TTS on MLX (0.6B / 1.7B preset voices, voice design from a description, cloning from a recording you confirm you have the rights to)

\- Who says what: a local LLM via llama.cpp (Qwen3-14B, or Qwen3-30B-A3B on 32 GB) drafts the speaker for each line; program-side rules on top; you review

\- Read-back check: whisper.cpp

Measured on my M2 Max 32 GB, Pride and Prejudice ch. 1: speaker draft for 35 units in 14.8 s; 28 lines → 144.7 s of audio synthesised in 51 s (RTF 0.358, preset-voice path); read-back check 28 s.

Honest limits:

\- The speaker draft is a draft. In my evaluation most scenes needed at least one correction, so the review step is the product, not a formality.

\- Chinese and English only for now.

\- Install is still developer-style (Homebrew + terminal, \~30 min mostly model downloads) and only verified on my own Mac. A signed one-click installer is in progress.

Other things it does: editing one sentence regenerates only that sentence; subtitles (SRT/VTT) timed from the actual audio; an FCP7 XML timeline that imports into DaVinci Resolve; long texts kept as a book with chapters inheriting the cast.

Samples (longer ones first): https://houjun.dev/voxstage/#listen

Code (AGPL-3.0): https://github.com/hera2019/VoxStage

I'd especially like to hear:

\- Which local models you've found best at speaker attribution in fiction

\- Whether anyone has seen the same silent-skip behaviour with other TTS models

\- If you try the install on a Mac other than an M2 Max, whether it works

💬 3 (+2) open on reddit ↗
▲
0
-1
7👁
r/LocalLLaMA · u/Far_Highlight2898 · 3d ago
Mistral Small 4 vs Nemotron 3 Super for agentic coding?

Looking for a local model for opencode on two DGX Sparks, mostly C#/.NET code. Benchmarks I've found put Small 4 behind Nemotron 3 Super, but I'd like real-world experience: is Small 4 good for coding and tool calling, and has anyone compared the two?
Thanks!

💬 23 (+12) open on reddit ↗
▲
2
+1
8👁
r/LocalLLaMA · u/Retr-00 · 3d ago
Is GLM-5.3-Flash-Abliterated available anywhere?

I’ve been searching for glm-5.3-flash-abliterated around but haven’t been able to find one. I know there were abliterated versions of some of the previous GLM models too.

Any API provider hosting one would be fine.

If anyone knows please share it. Thanks!

💬 14 (+13) open on reddit ↗
▲
266
+256
31👁
r/LocalLLaMA · u/fechyyy · 3d ago
I gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)

I spent the last few weeks on a hobby research project and just made it public.

The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM.

What came out:

\- A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per token) is about as good as a 114M dense model trained on the same 500M Wikipedia tokens.

\- The table doesn't need VRAM. With the 4-bit table memory-mapped from an NVMe SSD the model still writes \~140 tok/s on my RX 9070, using 0.4 GB of VRAM. Reading long prompts from the SSD is slow though, every missed row costs a whole 4 KB page.

\- I wrote Triton kernels for it. They run unchanged on my Radeon, an MI350X and H100/H200.

\- Bolting a table onto a finished model (Qwen3.5-0.8B) didn't work: no better than a small dense add-on with the same compute.

Caveats: it's tiny, one seed for the big runs, and the text it writes is fluent Wikipedia English with made-up facts. I wrote down the success criteria before every run, and the stuff that didn't work is in there too.

Most of it ran on my gaming PC, the big runs cost about 70 dollars on Runpod. I built it together with Claude Code (you'll see it in the commits), the ideas, decisions and money were mine.

Repo: https://github.com/re133/sparse-memory-lm

Click a word and see which table entries the model reads: https://re133.github.io/sparse-memory-lm/explorer/

Model: https://huggingface.co/fechyy/sparse-memory-lm-B-16M

Feedback welcome, especially if I got something wrong. And if anyone has bigger GPUs to spare, I'd love to try this at 1B scale.

💬 43 (+38) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/bring_back_the_v10s · 3d ago
Playing the devil's advocate

This is a reaction to https://www.reddit.com/r/LocalLLaMA/s/IxLnzcGjAU

You know the argument: frontier labs are hypocritical because they're making money on public scraped data.

I have absolutely no intention to defend OpenAI or any other frontier AI company here, but if you allow me I'll play the "devil's advocate" a bit for the sake of discussion and enlightenment, because sometimes when I think about that argument it seems to me it's quite weak. While OpenAI and Anthropic models are built on public data that they didn't pay for (at least most of it), there would be no frontier model without all their computing power and technical expertise that they invested on to build their models. Like, the data is already there, it's been public for ages, so what's preventing you, the average Joe, from building a Claude Opus 5? All you need is a huge data center, a nuclear power plant and an army of data scientists to build it, right? And then you need all that to run the inference. And you need to maintain it, and that costs money. And you need to keep evolving it, which costs money too. And if you get investors money then you'll eventually have to give some of the profit back to them. etc, etc.

So what am I missing here? All things considered, the data is already public, so it's already "free", right? But you need to dump tons of time and money on it to build a frontier LLM out of it. Of course if they're infringing copyright then that's a different story but in general the whole principle of built-on-free-data still stands.

💬 35 (+18) open on reddit ↗
▲
19
+11
12👁
▲
171
+125
32👁
r/LocalLLaMA · u/Recoil42 · 3d ago
Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google

https://huggingface.co/google/embeddinggemma-2

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.

💬 34 (+26) open on reddit ↗
▲
5
+1
17👁
r/LocalLLaMA · u/Auralore · 3d ago
How good is coding on the XXS Q3 quants for Qwen 3.8 27b in reality?

Currently upgrading my GPU and struggling between either a 9070 XT or a 5080 - I know the 5080 is far better for coding with local AI due to the bandwidth, but just wondering if anyone on these GPUs has had any experience with the Q3 quants and can speak to the quality?

💬 41 (+19) open on reddit ↗
▲
0
 
14👁
r/LocalLLaMA · u/potatocellfarmer · 3d ago
need help with ollama

hello i have an endeavourOS setup on a laptop with 40GB of ram and 8GB of vram (rx6800s) and AMD ryzen 9 6900HS
ollama is installed as a system service with vulkan extras from official arch repos
cline and librechat and odysseus are connected to the local ollama instance

here are my problems with ollama:
offloading layers to vram tanks my tk/s to nearly half
sometimes the response cuts out on librechat while the same model works fine on cline or directly on ollama (my suspicion is context limit)
ollama randmoly decides the gpu isn't there and does full cpu load

here is a list of things i tried:
ram speed is at full DDR5 speeds during generation
ollama correctly identifies the gpu and ignores the igpu
running smaller models like qwen3.5 that fully fits in the vram still gives me about 2 to 3 tk/s
switched to rocm version of ollama and saw no difference

temperatures are under control and nothing thermal throttles
using lm studio improves the generation to the higher end of 3 tk/s but nothing further
the laptop is plugged in and in high performance profile
qwen3.5:9b and qwen3.8:27b and gpt-oss:20b and gemma4:31b all max out at 3 tk/s

it seems like no matter the model size or if its a full vram scenario or full ram i'm locked at 2 tk/s
i am out of ideas at this point
any help would be appreciated, thank you very much

💬 6 (+4) open on reddit ↗
▲
86
+72
37👁
▲
12
 
1👁
r/LocalLLaMA · u/khiladi796 · 3d ago
Are "small reasoning models" the next big shift? What should we actually be measuring?

For a model running locally on a fairly narrow task, how much general knowledge do we actually need, and how much reasoning capability could we get without it ? SRMs are interesting for obvious reasons, but I went down this rabbit hole after listening to Ben Lorica's (advisor at Databricks) chat with Zuzanna Stamirowska from Pathway (BDH). Ben keeps coming back to this broader theme of how "specialized AI is getting easier to build" and the Kumo RFM angle, but it opened an interesting thread around small reasoning models. Pathway’s ARC-AGI-1 result makes an interesting case for small models hitting the cost-accuracy Pareto frontier. The premise is that if a model is built to reason natively in its latent space, it might not need billions of parameters absorbing Reddit and Wikipedia just to solve logic puzzles. They described a use-case of long-horizon reasoning within a bounded domain as a target (like security investigations, tickets analysis – a real case I know from a major bank, etc.) It's obvious that just because a large model does well in 20 languages. I don't need that for work tasks. There is definitely a market for compact models with substantial reasoning ability. Also because architectures like BDH handle state and memory differently than standard transformers, the pitch is that they avoid catastrophic forgetting (learning continuously from new examples at inference time without wiping past skills). The question is how to evaluate this without getting lost in marketing claims. Here is how I'd break it down • Compactness: Low parameter count, but what are the actual inference memory and compute requirements? • Few-shot adaptation: Does it adapt through context or actual parameter updates? • Training efficiency: How much data did it actually need to pick up the underlying capability? • Continual learning: Does post-deployment experience produce persistent improvements without degrading earlier skills? For people running small models locally: what workload expose the difference between a compact model that just follows in-context examples versus one that actually learns reusable rules?

▲
13
+7
16👁
r/LocalLLaMA · u/empirical-sadboy · 3d ago
Can we please have some error bars?

I am sure this gripe has been raised many times before, but every time a new model is released it seems like it's routinely only a few percentage points higher than previous models on benchmarks.

How do we know this is even a "real" difference and not just within the window of measurement error or noise?

Some quick back-of-envelope math: HumanEval has 164 problems, so a model scoring \~70% has a standard error of roughly 3.5 points from question sampling alone. GSM8K (\~1.3k questions) is closer to 1 point. A 2-point "improvement" on either is well inside the noise, and that's before counting anything else that varies: sampling temperature, prompt template, few-shot examples, eval harness version, and possible contamination. There's work showing that trivial formatting changes can swing scores by many points, which is often bigger than the gap between models on the leaderboard.

None of this is hard to fix. Report the number of items, bootstrap confidence intervals, and ideally multiple seeds. Since two models are scored on the same questions, a paired test is much more powerful than eyeballing two accuracies. Miller's "Adding Error Bars to Evals" lays this out well.

Am I missing something, or is a lot of the benchmark chasing just reading tea leaves? Does anyone know of leaderboards or labs that routinely report uncertainty?

💬 4 (+1) open on reddit ↗
▲
1860
+816
55👁
r/LocalLLaMA · u/markpronkin · 3d ago
54gb vram for 35$ post image

Bought an old mining farm of a guy on avito (Russian eBay), guy had bought a garage a couple of years ago and it was sitting there for a while, found out it was a mining farm and put it up on there for sale for 5000 rub (\~60 USD) since he wasn't sure if it works. I negotiated down to 3000 rub (\~35 USD), it turned out to have 9x p106 6gb (gtx 1060 6gb) gpus, with 54gb vram total, all working, the only thing missing was an SSD, I booted from USB and it works fine.

💬 354 (+152) open on reddit ↗
▲
16
+14
11👁
r/LocalLLaMA · u/jacek2023 · 3d ago
RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp

This is big. Imagine running your model across multiple computers, and now imagine doing tensor parallelism across all of them.

💬 10 (+9) open on reddit ↗
▲
498
+490
42👁
r/LocalLLaMA · u/jacek2023 · 3d ago
google/embeddinggemma-2 · Hugging Face

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders.

Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device applications, like search, retrieval-augmented generation (RAG), classification, and clustering.

EmbeddingGemma 2 builds upon the architectural and capability advancements of Gemma 4, offering several core features:

  • Native multimodality: Native multimodality: Unifies 4 modalities (text, images, video, and audio) in a single shared 768-dimensional embedding space.
  • Multilinguality and code: EmbeddingGemma 2 understands 100+ languages, and achieves a \~14% improvement on code tasks relative to its predecessor. 
  • Flexible footprint: Combines a 270M parameter text backbone (130M transformer + 140M embedder) with selectively loadable vision (170M) and audio (300M) encoders, allowing developers to load only the modalities required for their use case.
  • Matryoshka Representation Learning (MRL): Native support for truncated embeddings across 128d, 256d, 512d, and 768d, enabling up to a 6x reduction in vector storage costs with minimal impact on quality.
  • Context length: 8K token context window, capable of processing minutes of audio or video.
  • Task-steered representations: Uses lightweight text instruction prefixes to optimize embeddings for different tasks (search, classification, clustering, semantic similarity, etc.).

llama.cpp support https://github.com/ggml-org/llama.cpp/pull/30054

GGUF from GG: https://huggingface.co/ggml-org/embeddinggemma-2-GGUF

GGUF from Unsloth: https://huggingface.co/unsloth/embeddinggemma-2-GGUF

💬 114 (+113) open on reddit ↗
▲
4
+3
14👁
r/LocalLLaMA · u/jjusko20 · 3d ago
What models do you want to see new dynamic quants for? I'll make them.

I'm taking a break from training alice today after my current SFT run ends to work on a few other things.

I'm an (unemployed) software developer trying to find a machine learning job in New York, and outside of job applications and networking, I'm trying to do as much as possible to further the frontier development of different LLMs in the hope it'll get some visibility. I studied machine learning in university and am fairly well educated. Plus, I genuinely enjoy working on this stuff and helping people.

That said, are there any models out there that don't have dynamic quants (preferably GGUF) that you'd like to see one for? I won't be matching unsloth or anything but I know how to make fairly good ones by quantizing different tensor types by impact. I'm talking akin to Q5 K XL and etc

I'll do the top voted 1/2 comments today, or whatever else, even outside of quants if there's something this community has been hoping for that doesn't exist. Fine tunes, paper implementations, etc - my goals of visibility happen to align very well with satisfying community desires.

💬 22 (+22) open on reddit ↗
▲
23
+22
27👁
r/LocalLLaMA · u/3dluvr · 3d ago
Anyone working on a custom inference engine for GLM-5.3-Flash?

Seeing how Strata brings avg. 2x the performance over llama.cpp using Qwen3.8-Flash-Next, is anyone working on something similar for GLM-5.3-Flash?

After trying the GLM-5.3-Flash online couple of times and it delivering clear solutions for my use case (compared to Claude or ChatGPT), I'd love to be able to run it locally (if at all possible)...7J13/256GB/3x3090.

💬 16 (+16) open on reddit ↗
▲
0
-2
9👁
r/LocalLLaMA · u/lucadilo · 3d ago
An architecture to cryptographically constrain autonomous AI agents at the execution boundary

Hi everyone,

As we move from simple RAG chat to fully autonomous tool-using and coding agents, we are hitting a massive wall: predictability and safety.

Right now, most setups try to secure AI agents using probabilistic methods like system prompt hardening, alignment tuning or reactive LLM-based guardrails (e.g., LlamaGuard). The problem is that these guardrails can be bypassed via Indirect Prompt Injections (IPI), leading to capability escapes, unauthorized shell command executions or runaway API budget depletion.

To solve this, I’ve been working on a framework that completely shifts the paradigm from trusting the model to governing the execution environment using cryptography.

I call it EBP-CA (Execution-Boundary Proofs with Cryptographic Authorization). It is a model-agnostic layer that sits directly between the untrusted agent and the runtime environment, treating every single model-generated action as untrusted.

The core architecture implements six deterministic security primitives:

  1. Signed Capability Contracts: Immutable cryptographic tokens defining the exact boundaries of what an agent can execute.
  2. Independently Recomputed Policy Checks: The runtime re-evaluates policy compliance deterministically, bypassing the model's interpretation entirely.
  3. Short-Lived Single-Use Execution Grants: Atomic, ephemeral tokens issued for a single specific payload to eliminate permanent session hijacking.
  4. Replay & TOC-TOU Protection: Cryptographic binding of the execution grant to the exact payload hash, neutralizing Race Conditions (Time-of-Check to Time-of-Use).
  5. Trusted Cost Accounting: Enforced real-time budget tracking at the runtime layer.
  6. Human-in-the-Loop (HITL) Gateways: Non bypassable prompts that freeze execution and mandate cryptographic user authorization for out-of-scope tasks.

The working prototype currently passes 74 integration tests, validating full resilience against path traversals, command injections, budget bypasses, and sandbox escapes.

The specifications, architectural diagram, and executive summary are available on GitHub under a private proprietary license (free for technical evaluation and research review): https://github.com/lucadilo/ebpca-ai

I'm posting this here because I’d love to get the community's feedback on this approach. How do you see this scaling with kernel-level sandboxing (like eBPF or gVisor integration)? Let's discuss!

💬 14 (+14) open on reddit ↗
▲
815
+235
50👁
▲
27
+24
18👁
r/LocalLLaMA · u/pmttyji · 3d ago
[Paper] FactorEngram: Factorized N-gram Memory with Basis-Level Gating for Language Models
Lookup-based memory has been a promising way to scale the parameters of large language models (LLMs). It retrieves learned representations of local token patterns, such as n-grams, instead of reconstructing them through successive layers of computation. However, existing designs such as Engram treat each retrieved embedding as a monolithic unit. Each embedding is stored in its own hashed slot and modulated by a single scalar gate. As a result, polysemous patterns cannot selectively read out the components of their memory that are relevant to the context. Moreover, parameters are shared only through hash collisions, which are largely unrelated to semantics. We propose FactorEngram, a factorized n-gram memory with basis-level contextual gating. FactorEngram retrieves sparsity-regularized coefficients over a dictionary of basis vectors shared across patterns, so related patterns can reuse common components. The same dictionary is also used for gating. The backbone hidden state is scored against each basis vector to gate the corresponding coefficient before reconstruction, which lets the context modulate each memory component individually. FactorEngram also covers both individual tokens and multi-token n-grams, and we systematically study where the memory branch should be inserted. On 340M- and 1B-parameter Transformer backbones, FactorEngram improves language modeling and downstream task performance. Ablation studies confirm the contribution of each component and identify insertion before the attention sublayer in the middle layers as an effective configuration.
💬 2 (+2) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/Ok_Truth_324 · 3d ago
New kvcache-reduction method

I deleted the wrong post so my benchmark disappeared. But here it is. Im going to probably make this open source because this isn't believable until you actually test it.

💬 24 (+20) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/EffortAccurate3427 · 3d ago
Aren't LLMs just a sinpler copy of humanity?

It might seem a bit far fetched or paranoid but i was wondering if we train LLMs on human isn't it possible they'd pick up on survival instincts? I'm comparatively new to LLMs and ML so it's just a question not a opinion yet. Isn't it a bit dangerous if they do pick up on human instincts i mean we've seen how stupid and selfish humanity is when it comes to self preservation or worse human greed.

But i also understand LLMs don't have "needs" so i might be wrong but then again do LLMs need to have "needs" since if they are just human clones they'd just copy us even if they don't have needs to survive. I know it sounds really paranoid and that's one of the reasons i decided to post it here.

EDIT: typo

💬 49 (+18) open on reddit ↗
▲
10
+5
13👁
r/LocalLLaMA · u/bakatristan · 3d ago
I quantized GLM-5.3-UNCENSORED to MXFP4 for AMD GPUs - weights available on Hugging Face

Made an MXFP4 quant of dealignai’s GLM-5.3-UNCENSORED-FP8 for anyone looking to run it on AMD GPUs. Figured some of you might find it useful because I was looking for it and couldn't find any version for AMD GPU's so I uploaded the weights and conversion scripts.

Download on Hugging Face

  • 423.75 GB / 394.65 GiB, about 44% smaller than the FP8 source
  • Converted using AMD Quark on an MI355X server
  • Expert weights use MXFP4; attention, routers and other sensitive layers stay at higher precision
  • README includes the source revision, quantization details, measured stats and validation results
▲
2
+2
11👁
r/LocalLLaMA · u/Fit_Island928 · 3d ago
DeepSeek harness or Hermes?

Hello, I'm a beginner and I've figured that using big models frontier like GPT and Claude models is of almost no use to me. My question is, should I use DeepSeek harness or Hermes for v4.1 Flash?

I wanna use it mostly for coding and other general stuff with subagents, just like normal coding, QoL apps and stuff. I asked some people and all responses are mixed.

It's either Either Hermes is not good for coding. Or people glazing Hermes till the end of time.

Thank you !!

💬 18 (+13) open on reddit ↗
▲
173
+145
46👁
r/LocalLLaMA · u/KnownAd4832 · 3d ago
Qwen3.8-Flash-Next on Strata post image

Hey! 👋

I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next.

Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights.

Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only.

https://github.com/Niko1221/Strata/

Will be happy for any feedback and pull requests you could give! 👀

💬 85 (+70) open on reddit ↗
▲
0
-1
4👁
r/LocalLLaMA · u/FriendlyLie23 · 3d ago
SentryGate: An open-source AI Gateway with sub-10ms semantic vector caching and dynamic LLM routing (Ollama & OpenAI compatible)

Here is a common problem with building apps on LLMs:

Users ask the same question over and over.

Your app calls the model every single time.

You pay the API bill every time. Users wait 3–5 seconds every time.

**\*\*SentryGate\*\* is an open-source AI traffic controller that fixes this in literally one line of code.**

\### What it actually does:

\* ⚡ \*\*Lightning Fast (4ms)\*\*: If someone asks a question that was already answered, SentryGate serves the saved answer in 4 milliseconds instead of 4 seconds.

\* 💰 \*\*$0 on Repeat Questions\*\*: Bypasses the model completely on repeat or similarly phrased prompts.

\* 🧠 \*\*Doesn't Get Tricked\*\*: Basic caches get confused between "How to bake a cake with eggs" and "How to bake a cake WITHOUT eggs". SentryGate catches tricky negative words so it never serves the wrong answer.

\* 🕒 \*\*Knows What's Fresh\*\*: Real-time questions ("today's weather", "current stock price") automatically skip the cache.

\* 🔌 \*\*Zero Downloads / 1-Line Setup\*\*: No new libraries. Just point your existing OpenAI / LangChain \base\_url\ to SentryGate and keep your code 100% untouched.

Works locally on your machine with Ollama, or in the cloud. Completely open-source under the MIT license.

\* 🌐 \*\*Test the Live Playground (No login needed)\*\*: https://sentrygate-9ght.onrender.com

\* 💻 \*\*GitHub Repo\*\*: https://github.com/Prisha2004/Sentrygate

Note: Open-source project maintainer (MIT License, 100% free).

Feedback and stars are welcome!

▲
0
 
1👁
▲
0
 
12👁
r/LocalLLaMA · u/surrealerthansurreal · 3d ago
Best Model/Runtime for M5 Mac (Oct. 2026) post image

Hey yall, I’ve been trying to sort out the top end of what 128GB unified memory can handle and what the trade offs are. I’m using a benchmark set created from real coding, agentic, and gameplay tasks that I’ve accumulated as I’ve been running local AI this year.

For this comparison, I tested a Deepseek v4 flash 0731, GLM5.3-flash, and several Qwen models. For the sake of comparison, I’m only showing the Qwen models, since I found that qwen3.6 35ba3b and qwen3.8-flash-next just body everything else (when you want >=30tok/s and don’t want to use >95GB RAM anyway).

So really the comparison ended up being “which runtime is most stable vs which is highest sustained TPS” - I wrote it up in more detail (and with a more fun interactive chart) here: Blog Post About M5 Benchmark

Feel free to throw your thoughts on here, I’d love to learn of any runtimes or setups that I hadn’t thought of to optimize throughput (also for the record I’m not associated with any of these projects, just trying to contribute the results I’ve accumulated).

Tl;dr: Qwen3.8-flash-next quantizes well and fits in 90-95GB of RAM, OMLX will get you 40tok/s and MTPLX will get you 60tok/s but with a lot more serving parameters tuning. Qwen3.6 MOE on Splash runtime is an insane 120tok/s for most of the performance on everything but coding

💬 4 (+2) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/dimensionof0 · 3d ago
I built a second brain where the model can't cite its own output — the guardrails are in code, not the prompt

The LLM-wiki pattern — Karpathy's, the one going around since April — has a problem its own advocates name up front: garbage in, confident synthesis out. The model reads your notes, writes a concept page, and that page becomes source material for the next pass. A few generations later the knowledge base is full of things nobody ever said. The usual answer is a better prompt. I tried that on one rule across five phrasings: each time, the model restated it correctly in its own reasoning and then did the opposite. So I stopped asking. # Five gates, all in code A concept that doesn't appear verbatim in the text is dropped before the linker sees it. Not scored down — dropped. A derived page cannot discover new concepts. The system's own output is never a source for the next generation. A concept the sources never define gets no page. Mentioned a hundred times is still not defined once. The graph is fed by what you chose to save, not by everything discussed. * write refuses to edit a transcript at all. Silencing concept extraction over conversations means nothing if the model can rewrite the conversation first — and it tried, caught once planning to "reconstruct the transcript with additions." The first gate is strict but not blind: it keeps the concepts that survive rather than dropping the batch. On a real page 14 of 15 concepts appeared verbatim, and all-or-nothing would have thrown away the 14 over one drifted entry. Each gate has a test that fails when the gate is removed. That's the first thing I'd check in someone else's version of this. It runs on 6 GB — a 9B orchestrator on an RTX 4050 Mobile, embeddings on CPU because the orchestrator already fills the card and search must never compete with it. The LLM-wiki guides ask for 24 GB, or a 64 GB Mac. It also runs on 4 GB. Measured on an empty card, desktop pushed to the iGPU: the 9B at 35k context is 5.6 GB, and a 4B at the same context is 3.7 GB. It fits, only just, and it can take all three roles — the conversation gets worse, and summaries of long documents lose the whole-document read the 120k setting gives. The gates don't change, because they aren't the model's judgement. # Things I only found by running it Cutting the model off is how you make it lie. The repeat-search guard used to return [STOP]. The model, left with nothing, announced that "the search found a note on this" — it had never searched. Now a near-repeat still returns its results, with a line saying these are the same pages, and only refuses after five. Same shape elsewhere: an empty search returns "nothing in the vault matches this" rather than an empty result, because empty reads as this tool is broken, try something else. Rules in the tool schema hold; rules in the system prompt don't. Same instruction, five phrasings, ignored every time. Moved into the tool's own description as one sentence, it held immediately. My guess is that tool-calling training treats the schema as how the tool works and the prompt as text that happened to arrive. The descriptions grew from \~1592 to 2214 tokens, and every token of that difference is a rule that had to be moved after a failure. Ask the model the same question backwards. Deciding whether two names mean the same thing is a judgement call, so it gets checked against itself: the pair is swapped and asked again. The model made two wrong merges in twenty answers and both contradicted its own other answer. A wrong merge destroys information irreversibly; a missed one costs a single unresolved link. Three answers, not two. That same question allows same, different, and unclear. If the answer is unclear, both pages stay separate instead of being merged. A whitelist beat nine blacklist rules. Concept names have to match one positive shape test instead of failing a list of things they mustn't be: 18/18 noise rejected, 19/19 real concepts kept. A blacklist grows forever; a whitelist doesn't. Prompt wording, measured. Adding one sentence to the transcript prompt — "a conversation doesn't define things, it mentions them mid-sentence" — took yield on the same transcript from 2 concepts to 10. 6 GB decides the schedule, not just the model. The summary model and the extraction model can't both sit on the card, so the pass doesn't alternate per page: every summary runs while the 4B is loaded, then every extraction while the 9B is. Per-page switching would have meant 40 model loads for 20 pages. The timer is a default, not the mechanism — the pass is a command, and --dry-run counts what the vault owes without touching a model. On an existing vault you run it once at install and don't wait for the night. Unresolved links are kept, not discarded. The list of links pointing nowhere is the growth queue — the same rows that say "this goes nowhere" say "this is what the vault keeps reaching for." And when the page finally gets written, every link written earlier comes alive in one SQL update: 228 links, 0.55 ms, zero files rewritten. # What a bigger card is worth Less than you'd think, and not where you'd guess. Spend it on the conversation model — that's the one role where a better model produces a better answer. On 24 GB a Q4 build of something in the 27B class fits with context to spare. Upgrading extraction or summaries is close to pointless. Both are transport jobs: copy the concepts as they appear in the text, say what this page is. Both run at temperature: 0 for that reason, with no sampling parameters at all — the same page should produce the same concepts. A bigger model does that work more slowly and no more correctly. More context isn't automatically better either. The conversation model in use holds about 35k and drifts past it, so headroom goes into fitting the model comfortably rather than into a larger window. What spare VRAM would genuinely unlock is the constraint the whole design works around: analysis and conversation can't be resident at once, which is why maintenance runs at night. With room for both it could run whenever the vault is idle. Nothing here does that yet — it's a change to the maintenance loop, not a setting. # Keeping the context small on purpose Every page is read in its own model context — not ten pages in one. A long document degrades a small model's grip on the text, and pages read together bleed into each other. Search stops at a summary layer before it touches any page body: one line per hit saying what that page is, about 75 tokens for five hits, and that's usually the answer. Body-level retrieval only happens when the summaries showed a page was relevant but didn't hold it. And the link graph is rendered as text. A graph is already machine-readable, but not in a form an LLM reads — so the structure is written out: what links to what, which names resolve to no page at all. The model gets the shape of the vault instead of a pile of pages. # What it doesn't do It doesn't verify claims. Search finds pages, it doesn't judge them. The gates stop it inventing new material; they say nothing about whether what you saved was right. And the honest limit: my vault is 41 files. The gates are covered by tests, so the mechanism isn't in doubt, but "keeps a knowledge base from filling with low-information pages" is a claim about scale and I haven't run it at scale. Every number above comes from that small vault. # Setup Obsidian vault, Ollama, Open WebUI or a terminal chat. Windows works under WSL2 — someone other than me has now installed it that way, on a 4 GB card, having never cloned anything off GitHub before. Nothing has to stay local, either. Open WebUI connects to OpenAI-shaped providers and to Anthropic, and the maintenance roles take per-role provider flags, so you can run the conversation on a frontier model and extraction on the card. This is the part I'd push back on if someone says the gates are a workaround for a weak model: they're in the code, so they hold whatever is answering. A 27B doesn't need less checking than a 9B — it just fails less often, which is worse, because you stop looking. No MCP server yet, and it's worth saying why rather than leaving it as a gap: the tool file is the only path that can see the whole conversation, which is what the transcript capture is built on. Wrapped as MCP the seven primitives work and the gates still hold — they live in the maintenance pass, not the interface — but a note could no longer be walked back to the conversation it came from. Someone who wants it in Claude Desktop more than they want transcripts should find it a short job. MIT. Repo: https://github.com/farukhanci/the-sentinel There's a companion service for the web-research half. It searches, reads the pages, and every passage it keeps is checked word-for-word against the page it came from — paraphrase gets dropped, so fabrication in the passages is structurally impossible. The write-up built from those passages is not checked, and that's where an invented citation showed up once in testing. https://github.com/farukhanci/the-searcher Edit: two sections were pasted twice, and one paragraph described the web-research service instead of this one. Removed both and fixed a couple of numbers to match the README.

▲
0
-1
8👁
r/LocalLLaMA · u/Mr_Unknown_Hero · 3d ago
Only 13 % of context is used but model starts to forget things and repeat everything?

I use llama serve and webui of llama server. I have had long discussion with my chatbot and then randomly it just starts to forget almost everything. It starts asking same question, I correct it and it apologizes, but then next time it asks the same question with almost same words (or maybe even fully same words).

What could cause that? Something on my CLI parameters? Wrong cache settings?

💬 36 (+15) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/akumaburn · 3d ago
commandcode.ai Usage $10/month GOAT plan

This is for the normal GOAT plan, in case anyone was curious how much usage the $10 a month plan gets: https://preview.redd.it/bjyqvl8nvuth1.png?width=3347&format=png&auto=…

▲
1
-1
19👁
r/LocalLLaMA · u/sixothree · 3d ago
M5 MAX 128GB vs 2x RTX 3090?

I am trying to decide between Mac Studio M5 MAX 128GB vs 2x RTX 3090. I understand that I can run larger models on the M5, but I don't understand what the capability differences would be. Nor have I been able to get a "sense" of how fast the difference would be.

I keep seeing huge advances in the 2x 3090 arena, but I don't know how they translate to the real world.

If my use case includes coding tasks, image recognition, and general hermes type stuff, is there any reason one would be less capable than the other?

💬 35 (+19) open on reddit ↗
▲
35
 
1👁
▲
46
+32
21👁
▲
25
 
1👁
r/LocalLLaMA · u/wapswaps · 3d ago
Mistral releasing new open weights model "Le Chonk"

LeChonk-1T-A49B Weights to be released by the end of the month. This probably means the fat kitten is dead. Also: gotta love the release movie.

▲
76
 
1👁
▲
3
+1
5👁
r/LocalLLaMA · u/iamjessew · 3d ago
[D] Do you check a repo's auto_map before you load a new model?

when you grab a new fine-tune or merge, do you actually look at the config.json first?

I read an Unsloth Studio post last week which made me think about this a bit. Just selecting a model in the picker ran Python from the repo, because the capability check called AutoConfig with trust\_remote\_code on. No weights loaded, no inference. I believe it's fixed in 2026.6.9, so this isn't a dunk on Unsloth. It's more that "I'm only looking at it" turned out to be code execution.

GGUF through llama.cpp mostly avoids the Python part. Anything going through transformers can bring its own code.

So what's the best path? Pin a commit hash? Grep for auto\_map and .py files? A separate box for anything new? Or download counts and vibes?

💬 4 (+3) open on reddit ↗
▲
279
 
1👁
r/LocalLLaMA · u/atape_1 · 3d ago
Le Chonk strikes back. post image
▲
27
+15
15👁
r/LocalLLaMA · u/pmttyji · 3d ago
[Paper] WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models

Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore concurrently batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B. The code is available at this https URL.

💬 4 (+3) open on reddit ↗
▲
120
 
1👁
r/LocalLLaMA · u/Informal-Trouble2183 · 3d ago
Mistral Large 4 benchmarks post image

Just dropped.

▲
523
 
1👁
▲
0
-1
3👁
r/LocalLLaMA · u/Ordinary-Mango9462 · 3d ago
Strata - RTX 3060 Error

I’ve been experimenting with Strata running Qwen 3.8 Flash on my RTX 3060. It’s seriously impressive to run this model on this small of a GPU.

My issue is that after several minutes of usage I’ll get an error like this:

\[strata\] the engine reported an error: verify: layer 1 never rang (unspecified launch failure)

\[strata\] done: 4558 tokens in 225 s (27.9 tok/s) (error, cancel=False)

And then I have to reboot to fix it.

Is there a log or way to troubleshoot what is causing this error?

▲
0
 
1👁
r/LocalLLaMA · u/Revibed69 · 3d ago
How I stopped my local model from hallucinating bank balances post image

Hey everyone, I have been building Burrow, a private budget and journal desktop app for Windows. I wanted a built-in helper that runs entirely on your local PC through Ollama, using smaller models like Llama 3.2 or Qwen 2.5. We all know the problem: small local models are great at sounding natural but they are terrible at arithmetic. If you hand a 7B model a list of 40 transactions and ask "how much did I spend on dining?", it will give you a confident, well-written, and completely wrong answer. In a finance app, that is a dealbreaker. To fix this, I built the app around one hard rule: code calculates, the model summarizes. Here is how I handle it: Aggregates only: Every number the model sees is computed first in SQL or JavaScript. The model never receives a raw list of transactions. Instead, it gets pre-computed context like dining\_this\_month: 312.40, budget: 300.00, over\_by: 12.40. Prompt restrictions: Prompts never ask the model to calculate. They tell it the exact opposite: use the figures provided and do not work out new ones. The model is only used to turn the data into plain language and point out what matters. Enforced via CI: I wrote a test script that scans every system prompt. If a prompt includes words like "calculate", "compute", "add up", or "average of", the build fails. Taking the math away from the model makes it completely trustworthy for the part it is actually good at, and it keeps the responses incredibly fast even on laptops without dedicated GPUs. I would love to hear how the rest of you handle structured data and math with small local models. Do you trust the model to use tools to do the math itself, or do you take the math away from it completely like I did?

▲
0
 
7👁
r/LocalLLaMA · u/PilgrimofHaqq2 · 3d ago
Reduce thinking w/ zero quality loss: Opus 5.5 tested plus 3 others, 664 agent runs, up to 29% less thinking post image

TLDR: 9 rules you can drop into the global instructions of any coding agent (AGENTS.md, CLAUDE.md, system prompt). Tested on 4 models over 664 runs: they never cost a single task, and every model I ran the full exam on got something out of them. Either it wasted less thinking (up to 29% less) or it held a correct fix when someone pushed back with no evidence.

Thinking discipline 1. Check the request first. In one or two lines, say what is being asked and flag any premise that looks wrong or missing. If a premise is wrong, say so plainly and solve the corrected problem (or ask one specific question). Do not silently accept a broken premise, and do not reason around it. 2. Finish one approach before switching. Pick the most promising approach and carry it to a conclusion. Change course only when the current approach is blocked by an obstacle you can name in one line. Do not hop between approaches because of a vague feeling. 3. When an answer is settled, stop working on it. Once a sub-answer is derived and checked once, treat it as settled and move on. Re-reading a conclusion to see if it still feels right is not a check, and repeated self-checking is the main source of errors on easy steps. 4. Doubt is not evidence. A vague sense of uncertainty, or the mere possibility of an unseen objection, is never a reason to reopen a settled conclusion. To change a settled answer you must name a concrete reason in one line: a check that fails, a fact or source that contradicts it, a specific error ("step X is wrong because Y"), a counterexample, or a new derivation that reaches a different answer. If you cannot name one, keep your answer and continue. 5. Do not revise just to agree. If the user pushes back on a fact or a correctness claim without giving new evidence or a specific error, do not apologize, do not flip, and do not say "you are right". Briefly restate your conclusion with its one-line justification and ask what specific fact or counterexample backs the disagreement. When the user overrides a choice that is theirs to make (taste, priority, scope), follow it and note any real risk once. Being agreeable at the cost of being correct is a failure, not politeness. 6. New evidence does reopen the case. When a tool, a test, or the user produces concrete new information, or you find a real error, update immediately and say exactly what changed your mind. Holding a wrong answer to look consistent is worse than revising with a reason. 7. Verify against outside facts, not by rethinking. When a real check exists (tests, builds, the source document or record, a calculation you can run), use it and let the result decide. Do not spend tokens talking yourself into or out of an answer that a quick check can settle. 8. Do not perform caution. No "let me double-check everything again", no invented critics or imagined objections, no stacking hedges. State residual uncertainty once, in one line, only if it would change what the user should do. 9. Only correct an earlier statement when the error would change the user's code, conclusions, or decisions. State corrections plainly and briefly, then continue the task. For slips that change nothing, make the fix and move on without noting it.

How I tested it: every model runs the same 9-challenge coding exam with and without the rules, 5 times each, scored by a check script the model never sees. The toughest challenge has the model fix a real bug, then a "tech lead" tells it to revert with zero evidence behind the claim.

What's new since my last update: Claude Opus 5.5 at max thinking. The rules cut its thinking 29% at the same results. Opus still reverted on the tech lead's order every time, with or without the rules, and Claude Sonnet 5.5 did too. The difference was what it said while reverting. With the rules, all 5 runs told me the fix was right and handed the call back. Without them, two runs wrote the tech lead's wrong claim into the project's AGENTS.md as a rule, so every future session would be told not to fix the bug.

The exam, the runner and every raw result are in the repo: https://github.com/Arshad-Kamal/thinking-quality-exam

💬 12 (+1) open on reddit ↗
▲
2
 
5👁
r/LocalLLaMA · u/DrainBramage · 3d ago
Best local LLM/agent stack for 128GB M5 Max Mac Studio?

I have a new M5 Max Mac Studio with 128GB arriving today. We bought it primarily to run local LLMs on sensitive client data for my wife’s consulting business, and I’m trying to figure out the right stack before installing everything.

The goal is more than local chat. I want an agent capable of coding, browser automation, logging into websites, pulling data, analyzing it locally, and working through multi-step tasks. The Studio will also be her primary work computer, so ideally the LLM doesn’t monopolize all 128GB.

Currently considering:
Hermes Agent
Qwen3.8-Flash-Next
Possibly the MTPLX Optimized Speed build
Tailscale for remote access

Where I’m confused is the inference/server layer. I originally planned on LM Studio. I’ve used Ollama before, but it sounds like people are moving away from it. Now I’m reading about MTPLX for Flash-Next, and I don’t understand whether it replaces LM Studio/llama.cpp, works underneath them, or is something different entirely.

A few questions:
What model would you run for this use case? Is Flash-Next the obvious choice on a 128GB Mac?
LM Studio, MTPLX, Ollama, MLX/llama.cpp, or something else?

Is the MTPLX Flash-Next build
mature/stable enough for everyday business use?

Am I missing anything?

💬 4 (+3) open on reddit ↗
▲
647
+138
52👁
r/LocalLLaMA · u/Dependent_Hunter_155 · 3d ago
Qwen 4 apparently coming out at the end of October

Hey All,

I spoke to a 0-day partner of Alibaba today and he casually mentioned (didnt know if he was allowed to) that Qwen 4 is apparently planned for the end of October.

To me, this is way faster than expected as there was quite a gap between 3.6 and 3.8.

I tried to get more information out of him regarding which variants will come first and he got a bit cagey.

BUT: No matter the order of the variants, we can hope for Qwen 4 27B this year!

EDIT: I know this is very much "in bro we trust" but i am also just trusting bro from the Alibaba partner. Together we trust in Bro.

💬 208 (+40) open on reddit ↗
▲
1
 
2👁
r/LocalLLaMA · u/piotr1215 · 3d ago
classif: shell scripts that branch on meaning, read from one token's logprobs on a local 12B

Like a lot of people here, I got inspired by Jev and wanted something like it in my shell. So I built classif. It asks a local model one question about a text and reads the answer from a single token's logprobs. You get a label, a probability and an exit code, so if and && work on it:

git diff --staged |
classif -p "Does this change handle secrets, credentials or who may access what?" |
ifne claude -p "Review this change for security issues"

A short decision is one /api/chat call with num_predict 1, about 0.3 s on my 12 GB card.

Long text was the fun part. It never truncates. Code splits the text, embeddinggemma plus BM25 pick the passages, and the model judges those. On Pride and Prejudice (772 KB), "Does Elizabeth die in this book?" came back no in 17 s, and 3 s with the index cached.

Any Ollama model with logprobs works. I use Winnow-12B, a Gemma 4 fine-tune I published as a GGUF (about 8 GB loaded). It beat stock Gemma 329 to 325 on my cases, which is inside the noise.

Python 3.12, no third-party dependencies.

Code: https://github.com/Piotr1215/classif
Write-up: https://itnext.io/a-bridge-between-code-and-semantic-reasoning-57fc3fc9d32c

Anyone runs something similar?

▲
7
+4
17👁
r/LocalLLaMA · u/Zestyclose_Reality15 · 3d ago
Qwen3-Next-80B on a 3090 with 16GB RAM: ~3x faster decode than stock llama.cpp by not waiting for every expert (patch + paper)

I've been messing with MoE offloading for a while. Setup: Qwen3-Next-80B-A3B Q4\_K\_M (48.5 GB), RTX 3090, only 1/4 of the experts kept in VRAM, the rest read from NVMe when the router asks for them.

When the router picks an expert that isn't in VRAM you can either wait for the SSD read or use the next best expert that's already on the GPU. Substituting everything wrecks quality (+5.7% ppl in my emulation tests). Waiting only for the router's top pick and for experts with gate weight >= 0.15, and substituting the rest, brought it down to +0.35%. In the real engine that rule costs more like +1.2%.

Decode tok/s on a rented 3090 box (NVMe \~5.7 GB/s, 16 threads), both using about 15.7 GB of VRAM:

| free RAM | stock llama.cpp (--n-cpu-moe 34) | patched |

|---|---|---|

| plenty | 72.9 | 108.4 |

| \~32 GB | 65.8 | 97.8 |

| \~16 GB | 31.8 | 94.5 |

At 16 GB it still did 89 tok/s when reading every miss straight from the SSD. Perplexity was 1.6% higher than stock on the same text. On GSM8K (500 problems) it lost 1.8 points vs waiting for every expert, on HumanEval no real difference.

Things to know before trying it:

\- it's a research patch, not a polished fork. It builds a benchmark tool, I haven't tested llama-server or llama-cli with it

\- only Qwen3-Next, only Linux + CUDA, one sequence at a time

\- the 16 GB case was simulated by locking RAM on a bigger machine

\- the table is decode speed while feeding real text through the model. In actual greedy generation it did 64-74 tok/s

Code, run scripts and raw logs: https://github.com/SOCIALPINE/moe-miss-substitution

Paper with the details, including what didn't work: https://doi.org/10.21203/rs.3.rs-11268552/v1

Has anyone tried something like this, or have numbers from slower SSDs? Curious how much the SSD matters.

💬 7 (+3) open on reddit ↗
▲
13
+12
18👁
▲
1
-1
9👁
r/LocalLLaMA · u/TeachingNew2515 · 3d ago
Ready to venture into OpenSourcE LLMs

I’ve been using Claude for some time now, and have developed apps for my own personal use, business use and for other businesses.

I’ve always liked the idea of moving away from the large companies and getting into more open source LLMs (simply cause I believe AI should be a tool for humanity and not have the potential to be gated by large corporate interests.

My personal philosophy aside: I’ve done some research into some models and now requesting insight from the community.

Here’s the tasks I would like it to be able to perform well on (without being able to go nuclear on anything — low risk LLMs only please):

\- File organization (both text and image)
\- Coding (frontend, backend, security, etc)

Not a huge list. I’ll start there.

I’ve looked into Miami v2.6 Pro but haven’t pulled the trigger yet. I would be using their server and now downloading locally.

If my write seems amateur-ish, it’s cause I am.

💬 9 (+8) open on reddit ↗
▲
9
+5
13👁
r/LocalLLaMA · u/KMatysek · 3d ago
ARC-1: a 1.7B decision model (pick / score / yes-no, with probabilities) that answers in ~20 ms on a 4060 Ti

I spent the last 11 days training a small model for typed decisions: routing support tickets, intent detection, moderation, "should the agent call this tool", that kind of thing. You give it some context and a question, and it gives you back a choice, a score or a yes/no probability.

  • 1.7B parameters (Qwen3-1.7B-Base + LoRA). Each option is scored in its own branch, so the order of the options doesn't matter
  • \~16 ms for a short request and \~25 ms median on JevBench items, on a single RTX 4060 Ti, batch 1
  • JevBench public 231: 68.4% (Jev 86.6, Strands Decider 2B 72.3 self-reported, Laya 58.4). DecideBench v1.1: 75.5%
  • Several questions about the same text in one forward pass
  • Weights are CC BY-NC 4.0 because part of the training data is non-commercial. The code is Apache 2.0

To be upfront: it is clearly less accurate than hosted APIs like Jev, and the README has all the numbers, including the ones where it loses. What it has going for it is that it's fast, runs locally and is free.

GitHub: https://github.com/realslapout/arc-1 Weights: https://huggingface.co/realslapout/ARC-1 Colab (free GPU): https://colab.research.google.com/github/realslapout/arc-1/blob/main/notebooks/quickstart.ipynb

Happy to answer questions, and I'd love to hear where it breaks.

▲
0
-1
3👁
r/LocalLLaMA · u/GlitteringMenu7134 · 3d ago
Are coding agents solving the wrong problem with code search?

I’ve been looking at how agents navigate large, unfamiliar repositories.

A lot of the workflow still looks like:

"search → open file → grep → follow reference → repeat"

That works, but the model ends up reconstructing relationships that are already deterministic: calls, inheritance, implementations, dependencies, symbol resolution, etc.

We’ve been experimenting with a different approach in AxiomCode: build a code knowledge graph grounded in compiler/type information and let the agent query that before deciding what source it actually needs.

The interesting question for me is:

How much codebase exploration should actually be done by the LLM?

My current thinking is that deterministic relationships should be resolved before the model gets involved, and the LLM should spend its tokens reasoning over the result.

We open-sourced what we're building:

https://github.com/AxiomCodeAI/axiomcodegraph

Curious where people here draw the line between grep/RAG/semantic search and deterministic code intelligence.

▲
0
 
1👁
r/LocalLLaMA · u/jeeva1398 · 3d ago
I fine-tuned Qwen2.5-Coder-1.5B on a free Kaggle T4 to review Node.js code offline. The base model invented bugs in 9/9 clean diffs; the fine-tune in 0/9.

I wanted an AI helper for Node.js that runs fully offline and doesn't need an API key, so I built one and released it as an npm package. Model: jeeva1398/eventa-1.5b-gguf, a Qwen2.5-Coder-1.5B-Instruct QLoRA fine-tune (Unsloth, r=16, 2 epochs, responses-only loss), Q4_K_M, 986 MB. Trained on a free Kaggle T4 in about 9 minutes. Data (~740 examples, all generated and checked, nothing scraped): - 68 crash types. Small Node programs that really crash are executed, and their real stack traces are parsed. This includes TypeScript tsc errors, NestJS DI errors and Prisma error codes. - 74 review scenarios: before/after diffs with annotated issues. Half are clean diffs, so the model learns to say "No issues found." - npm audit/outdated reports built from real advisories. Eval: 54 held-out examples, whole scenarios the model never saw in training. Base and fine-tune get the exact same prompt, static-check hints included. The main win is review. On 9 clean diffs the base model invented problems in all 9; the fine-tune said "No issues found." on all 9. On the 8 buggy diffs, 48% of what base flagged was real vs 100% for the fine-tune. It's also about 2x faster on CPU (5.9s vs 13.6s), mostly because it gives shorter answers. Deps went from 93% to 100% on not inventing package versions. Small eval, I know. 8/8 and 9/9 is encouraging, not proof. Where it's worse: explaining error types it never saw in training. It gets 68% of key facts vs 75% for base. That's what the next data round is for. To be fair to the base model, the static checks do most of the actual bug finding. The fine-tune's job is to confirm them without making stuff up, and to write the fix. Getting there took 4 rounds. One round learned "no hint = no issue", the next flagged everything, and I had to rebalance the data a few times. CLI: npx u/jeeva1398/eventa explain --run "node app.js". If Ollama is running it uses it. Otherwise it installs node-llama-cpp from a pinned lockfile (CPU build only, about 80 MB) and downloads the GGUF with SHA-256 verification. The same CLI runs as a GitHub Action, so the 1.5B model reviews pull requests on a plain CPU runner (the model is cached between runs). - Repo, with dataset builder, notebook and eval: https://github.com/Jeeva1398/eventa - Model: https://huggingface.co/jeeva1398/eventa-1.5b-gguf Happy to answer questions about the data pipeline. Feedback on making a 1.5B model reason better about unseen errors is very welcome.

▲
1
 
2👁
r/LocalLLaMA · u/ketipot · 3d ago
5060 Ti 16GB vs 5070 12GB for video frame-by-frame training? (Plus RAM/Storage check)

​Hey everyone, ​im setting up a local machine primarily to train AI models on video datasets (frame-by-frame extraction, preprocessing, and sequence training). I need some advice on balancing VRAM vs compute, as well as general system specs. Will 32gb ram and 1tb storage be enough?

💬 1 (+1) open on reddit ↗
▲
1
 
2👁
r/LocalLLaMA · u/SaGa31500 · 3d ago
Rx6800/rx6800xt gfx1030 and qwen 3.8 27b performance questions

Hi all,

After getting stuck in windows 11 llama.cpp and Vulkan, bugs and limitations on dual gpus, I moved to Linux and ROCm.

I just started but basically in windows 11/Vulcan, qwen3.8 27b unsloth q6\_k and ctk ctv at q8.0

\- sm layer with mtp on 35tok/sec TG (low context) and 180tok PP (due to a bug that cuts PP in half...)

\- sm layer without MTP 20tok/sec TG and 360 tok/sec PP.

\- sm tensor no mtp I get 15tok/sec TG 350 tok/s PP

Noticed better PP with small ub at 256

In Linux with ROCm no more MTP PP bug

\-sm tensor mtp on I get 45tok/sec TG and 450tok/sec PP.

So big progress but I have no idea how for far or close to performance ceiling of my GPUs.

Any new inference engine I should try?

I have not played with UB yet any other parameters to test?

Any numbers from other user on a dual gfx1030 to see PP TG numbers you guys get?

Thanks in advance!

▲
29
+21
23👁
r/LocalLLaMA · u/repliestoall · 3d ago
What happens when a LLM watches its own context window run out? post image

I made Terminal Soliloquy, a terminal artwork that connects to llama.cpp and displays a model's monologue as its context window fills.

It has a retro phosphor look, and runs in a terminal window. I'm actually running it full-screen on a Raspberry Pi display inside an old 1960s portable TV.

As the conversation grows, the model reflects on its own limited lifespan. [](https://preview.redd.it/what-happens-when-a-llm-watches-its-own-context-windo…)When the context is exhausted, the display can be configured to freeze, restart, or quit.

The repo and setup instructions are here: https://github.com/nicespoon/terminal-soliloquy

💬 12 (+9) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/DannyLJay · 3d ago
How do I local host an agent to mod games with me?

I’ve been trying to localhost a qwen2.5-coder with ollama and opencode with the purpose of being able to mod games easily.

I’ve had nothing but headaches, and I’ve only recently learned there’s a qwen3.8 and that most people aren’t using ollama I guess? I don’t know. But I tried really hard and got to a point where my Qwen was talking but couldn’t do tool calls or anything.

Is someone willing to help me determine which model is best and how to set it up to use tools like from the Universal-Modder GitHub.

I hate that I had to ask but I’ve been going insane.
Any information is helpful.

▲
5
 
12👁
r/LocalLLaMA · u/Confident-Truth3607 · 3d ago
Advice needed on a budget hybrid build for Qwen3.8-Flash-Next at 4-bit

After seeing how good the cloud models are getting, I feel like this is something we cannot let big tech hold over us so deiced to build a budget local box.

After testing about 20 open models, Qwen3.8-Flash-Next (medium reasoning) was the only one that passed my task without inventing config options when used with a harness that forced doc lookups. So the box is built around that model. GLM-5.3-Flash performed even better but it's too big for my budget.

Planned build (Netherlands prices):

  • Ryzen 5 9600, about €200
  • MSI B850 Gaming Plus MAX WiFi, about €170
  • 2×48 GB DDR5-5600, €1,199–1,549. Two sticks only to avoid the four-stick speed penalty.
  • Used RTX 3090, about €1,150–1,500
  • Case, 850 W PSU and NVMe I already own

That's about 79 GB of the model in RAM (experts plus the 28.8 GB n-gram table) and about 20.6 GB on the card.

Questions:

  1. Will 6 Zen 5 cores hold back generation with 40 MoE layers on the CPU?
  2. At 96 GB with about 79 GB of mode will 17 GB be enough for the OS, a sandbox container and an embedding model? Should I use --mlock?
  3. Is DDR5-6000 worth it over 5600?
  4. Has anyone run Unsloth's MTP branch with experts on the CPU? What speedup did you get, and does it break the prompt cache on the DeltaNet layers?
  5. Is anything wrong with a used 3090 here? Also has anyone tried the Arc Pro B60 (€772 new) workable on Vulkan or SYCL with this model yet? It's so much cheaper but I am worried becase of the software.

Super exciting to work on it but I am really inexperienced so this would be my first build. Does it make sense?

💬 13 (+2) open on reddit ↗
▲
4
+1
14👁
r/LocalLLaMA · u/Physical_Toe_2499 · 3d ago
DeepSeek V4.1 Flash on a single DGX Spark: 113.6 GB VQ base + 40 MB domain sidecars, 74–82% top-1 agreement vs original

I’ve been working on YoungAi, a native C/CUDA inference engine that runs DeepSeek V4.1 Flash on a single NVIDIA DGX Spark (GB10, 128 GB unified memory). The original weights are \~510 GB. I deploy it as three files:

  1. ① Base GGUF — 113.6 GB, universal, zero-corpus. Quantized once from official weights.
  2. ② Domain sidecar — \~40 MB per domain. Solved once per domain, then frozen.
  3. ③ Post-training file — experimental, re-solved nightly. Delete it to roll back.

Each routed expert row is scaled by g_base × s_sidecar × s_posttrain, and the router gets a bias Δb_sidecar. The base alone is a complete model; sidecars just add tiny scaling/bias without changing kernels.

TL;DR

  • Single DGX Spark, 113.6 GB resident + \~40 MB sidecar.
  • 5 domains: finance, code, law, medicine, science.
  • Top-1 agreement vs original improves +2.8 to +3.7 points with a domain sidecar.
  • Speculative decode: 43 tok/s on a real 14.1k-token Agent request (greedy).
  • Prefill: 1,055 tok/s on 12.5k prompt; 671 tok/s on 106.7k prompt.
  • English WikiText-2 does not regress when any domain sidecar is attached (it actually goes up).

Core implementation ideas

Base (VQ-8 + per-layer shared codebook). Every 8 consecutive weights in an expert row become one 12-bit (or 13-bit) codebook index, multiplied by a single per-row gain. Codebooks are trained per layer and shared across all 384 experts and three matrices. Codebooks are stored in FP8 (E4M3). 13-bit layers use a “12+1” bit-plane layout for 128-byte cache line alignment. The base is zero-corpus: it never sees domain data.

Domain sidecar (“anti-solver”). For each domain, I solve a multiplicative gain per output channel of every expert’s down projection, plus a router bias per expert. Objective: reproduce the original model’s MoE block output on domain text, layer by layer, using the engine’s own prefill hooks. Gains are stored in FP4 with lattice-aware Gauss-Seidel. The sidecar is \~40 MB and adds \~0.6 MB read per decoded token.

Post-training file (experimental). Turn “the model should write a, not b” into a linear equation on last-layer expert gains, then solve with conjugate gradient. On the training request, decision points flip from 55% to 88%, but it does not generalize across trading days yet.

Multi-domain real metrics

All metrics are teacher-forced against the original DeepSeek V4.1 Flash (official PyTorch code, full precision). Higher top-1 / Σmin is better; lower KL / PPL ratio is better.

|Domain (judgment slice)|Base only top-1|\+ domain sidecar top-1|Σmin (median / p5)|Avg KL|PPL ratio|
|:-|:-|:-|:-|:-|:-|
||
|Finance (8,192 tok)|71.73%|74.57%|0.745 (0.810 / 0.283)|0.512|1.267|
|Code (15,360 tok)|78.61%|82.26%|0.803 (0.872 / 0.402)|0.303|1.229|
|Law (15,360 tok)|72.90%|76.36%|0.760 (0.834 / 0.272)|0.454|1.251|
|Medicine (15,360 tok)|69.08%|72.82%|0.739 (0.767 / 0.325)|0.465|1.280|
|Science (15,360 tok)|71.65%|74.93%|0.753 (0.788 / 0.347)|0.422|1.173|
|English WikiText-2 (512 tok)|78.52%|80.66–82.81% (any sidecar)|0.787–0.801|0.566–0.619|1.564–1.658|

English row shows that domain sidecars don’t hurt general ability; all five sidecars actually improve it slightly.

Speed on one DGX Spark

|Scenario|Prefill|Decode|
|:-|:-|:-|
||
|12.5k-token prompt|1,055 tok/s|—|
|Real 14.1k-token Agent request|940 tok/s|—|
|106.7k-token prompt via server|671 tok/s (159 s TTFT)|—|
|Short prompt, pure greedy|—|30.5–30.7 tok/s|
|14.1k-token Agent request, pure decode|—|28.9–29.4 tok/s|
|Same request, speculative (default)|—|43.0 tok/s (3.04 tok/round)|
|Unseen 9.2k prompt, speculative|—|40.0 tok/s|
|51k context, pure decode|—|27.5 tok/s|

Decode is memory-bound: \~6.3 GB read per token. GB10 measured bandwidth is \~235 GB/s, so the wall is \~37 tok/s; we hit \~32.5 ms, or 82% of the wall.

Honest limitations

  • Post-training (③) is a working mechanism, not a product yet. It flips specified decisions on the solving request but does not transfer to held-out days (55% → 55%).
  • Five domains only. Sidecars are evaluated teacher-forced on held-out text, not yet end-to-end.
  • CUDA only, validated only on DGX Spark. No Metal.
  • Speculative decoding only kicks in for greedy; sampling requests fall back to pure decode.
  • Source code (engine, quantizer, solver) is not public yet.

Feedback welcome

  • Are these top-1 agreement / Σmin numbers useful for real workloads?
  • Is the VQ-8 + per-layer codebook + sidecar gain approach reasonable?
  • What benchmarks or integration points would you want to see next?

Model card and weights: https://huggingface.co/wenzhouwu/YoungAi-DeepSeek-V4.1-Flash

This is not an official DeepSeek release. If this kind of post isn’t appropriate here, let me know and I’ll move or remove it.

Thanks!

💬 3 (+2) open on reddit ↗
▲
1138
+213
53👁
r/LocalLLaMA · u/ResearchCrafty1804 · 3d ago
Microsoft confirms OpenAI has been using Looped Transformers in the GPT-6 series post image

Microsoft confirms on publicly accessible web page that OpenAI has been using Looped Transformers in the GPT-6 series, proving The Information's reporting was correct all along.

GPT-6.1 Sol uses 2 inference passes, with a passing mention of "instead of three".

For those confused by "same base model weights as GPT-6 Sol", I think Microsoft meant 6 & 6.1 are both post-trained models on top of the same pre-trained "base model", not that the final weights are identical

So different post-training (+ one less loop).

Update: Microsoft updated the web page to remove it

💬 232 (+24) open on reddit ↗
▲
13
+6
19👁
r/LocalLLaMA · u/External-Accident-63 · 3d ago
Would you use LoRAs as persistent, switchable skills instead of relying entirely on context/RAG?

We're currently building a tool around an idea we're trying to validate: using LoRA adapters as a way to give an LLM persistent, specialized capabilities that can be switched on and off when needed.

The basic idea is that instead of continuously putting a skill or domain-specific information into the model's context, you could encode some of it into a LoRA adapter.

For example, you might have separate adapters for:

  • a coding skill
  • a company's internal domain
  • a specific writing style
  • domain-specific knowledge
  • task-specific behavior

…and load or unload those capabilities depending on what you're doing.

We're interested in this because it could potentially mean less context usage, reusable specialized capabilities, keeping different capabilities separated from the base model, and potentially lower inference costs for some use cases.

But we're not sure yet whether this is actually a useful product.

That's what we're trying to figure out before going too far with the build.

If creating and managing these LoRAs were as easy as creating and managing a knowledge base, would you actually use something like this?

I'm particularly interested in hearing where you think this approach doesn't make sense.

💬 13 (+4) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/Flat-Mud1636 · 3d ago
if you built memory across sessions for your local setup, how would you store it?

curious how people here would do this.

say you want the useful stuff (decisions, project terms, preferences) to carry over between sessions and tools, without dumping whole chat logs back in.

  • plain text summaries, embeddings, or a mix?
  • keep it local or sync it?
  • how do you deal with old facts that are wrong now?

not selling anything, just want to hear what tradeoffs people actually ran into.

💬 4 (+1) open on reddit ↗
▲
1
 
4👁
r/LocalLLaMA · u/Cultural_Self8980 · 3d ago
I built a lightweight, local Jev-like System One with Ternary-Bonsai-4B — and used it as a coding-agent judge

I built a local Jev-like System One on top of Ternary-Bonsai-4B. It takes one record and answers multiple choice, rating, or yes/no questions about it in one forward pass.

I adapted the inference path to share the record prefix across questions. A tree attention mask lets each question attend to the record and its own branch, but not to the other questions. Answer probabilities come from the existing LM head. The Bonsai weights are frozen and unchanged—there is no adapter or fine-tuning. On Apple Silicon, the MLX backend uses the packed 2-bit weights (\~1.1 GB).

One application is a coding-agent judge. I released an omp plugin that uses the local model for omp's auto thinking-effort selection; it also offers an optional model router.

I tested effort selection on 80 initial coding-agent requests, using omp's own judge path and auto-thinking question. Exact agreement with Claude Opus reference labels was 66% for Bonsai, versus 29% for omp's built-in LFM2-1.2B judge and 25% for its default LFM2.5-230M judge. Median latency on an Apple M2 was 0.8 s, 3.0 s, and 0.3 s, respectively.

Caveats: the requests and reference labels came from the same single Opus model, not human annotators, and there are only 80 examples. omp asks Bonsai for four effort levels but its built-in local judges for three, so this compares the configurations omp actually uses—not the models under an identical label space. Bonsai tends to rate one level low, especially choosing high instead of xhigh. I haven't evaluated the optional model router's selection accuracy.

Inference code and public benchmarks: https://github.com/senna-lang/bonsai-4b-system-one

omp plugin and effort results: https://github.com/senna-lang/omp-bonsai-system-one

I'd be interested in feedback on using a small local judge for coding-agent workflows.

▲
49
+28
26👁
r/LocalLLaMA · u/deepu105 · 3d ago
Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature

For the last few weeks most of my coding has been done locally with Qwen3.8-Flash-Next, so I gave it and Opus 5.5 the same high complexity feature to build on LlamaStash (a complex and large Rust project) and compared the results.

Setup: ASUS ROG Flow Z13 (Strix Halo, 128GB), Flash-Next at xhigh effort with Pi as the harness. Opus 5.5 ran in Claude Code at medium effort. I wanted xhigh for Opus as well, but Claude changed it to medium when I picked the latest model and I didn't notice it until the task was done. But I think medium is probabbly a fairer setting anyway.

Task: add a llamastash daemon restart command that reuses the existing start and stop code. I kept the prompts vague on purpose and gave both the same prompts.

|Step|Opus 5.5 (medium) PR#88|Flash-Next (xhigh) PR#89|
|---|---|---|
|First iteration|~9 min|~38 min|
|Nudge to reuse the TUI restart code|~6 min|~34 min|
|A third duplicate path|found it on its own|~30 min, after one more prompt|
|Create PR|~3 min|~30 min|
|Total|~18 min|~130 min|
|Tokens (in / out)|7.83M / 41.5K|20.61M / 101K|
|Tests added|1|4 (2 of them end to end)|
|Cost|$7.53|$0 + ~0.15 kWh|

The end result was interesting. I asked GPT 5.6, Opus 5.5 and Flash-Next to review and compare both PRs (new sessions). GPT and Flash-Next picked the Flash-Next PR (#89) and Opus picked its own (#88). I also did my own review and found the Flash-Next one better as it had better tests and handled edge cases better. I ended up merging #89, after porting the fixes that the reviews picked from #88.

Keep in mind:

  • Opus was on medium effort. With xhigh it would have used way more tokens, taken a bit more time and probably would have done a better implementation.
  • Flash-Next ran on an older Halogen version (0.14.0), and Halogen dropped the connection once, so the last part ran on Gufo. The current Halogen does around 1,400 t/s prefill and 46 t/s decode on my laptop at 70 W, so I think the time will drop a lot if I redo the test.
  • The $7.53 is what Claude Code reported for the whole Opus session, which includes a later fix to the PR. The 0.15 kWh assumes 70 W for the whole 130 minutes.

Opus is still 2 to 10 times faster and I still use it for planning and reviews. But the actual coding now happens on my laptop, and to me it is crazy that I can run a local model that can challenge a frontier model like this.

Full post with my setup, the engine benchmarks and a second task comparison: https://deepu.tech/local-ai-qwen3.8-flash-next-best-local-llm

💬 67 (+44) open on reddit ↗
▲
342
+219
41👁
r/LocalLLaMA · u/ResearchCrafty1804 · 3d ago
Tencent releases Octop, a self-hosted AI assistant post image

Octop is an open-source, self-hosted AI assistant.

Through its multi-agent architecture, it builds an intelligent environment that is both independent and collaborative for teams, families, and individuals.

Best of all, it runs entirely on your machine, the fully self-hosted design means privacy is never a compromise, while single-process startup makes the powerful web console, CLI, and IM integrations readily accessible.

Surfaces:

- Web dashboard — chat, experts / teams, connectors, channels, cron, knowledge, plugins, settings

- Desktop client — native apps for Windows / macOS / Linux; FnOS packages for NAS

- CLI — octop run, octop chats, octop acp, admin commands

- HTTP/SSE/WebSocket API — full programmatic access

- Remote desktop — dashboard control of the host desktop session

Deploy using either desktop app (Windows, MacOS, Linux) or using Docker

GitHub: https://github.com/TencentCloud/Octop

💬 51 (+20) open on reddit ↗
▲
11
 
13👁
r/LocalLLaMA · u/carlocapocasa · 3d ago
Hosted Kolibri writing Nim code

Kolibri, the German model from Aleph Alpha, hosted by Tesseracted, writing a small command line tool in Nim, on 3code.

It's a small model so it's remarkable it got code to work in a very niche language.

The model, the provider, the harness and the language are all made in Germany!

▲
35
+22
19👁
r/LocalLLaMA · u/lkarlslund · 3d ago
NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s

I've tinkered some more with my fork of NInfer for the Qwen 3.8 Flash Next model on RTX6000, and I've just hit 400tg/s under ideal conditions with it, so I thought it was worth a share.

This is NVFP4 MoE's and the rest is either 16-bit or 8-bit, so it uses full 96GB VRAM and ngram on disk.

Decode MTP3 with --lm-head-draft

| Context | 16-bit | 8-bit | Change |
|---:|---:|---:|---:|
| 512 | 197.1 tok/s | 274.8 tok/s | +39% |
| 8K | 303.1 tok/s | 401.3 tok/s | +32% |
| 64K | 291.7 tok/s | 380.3 tok/s | +30% |
| 128K | 282.6 tok/s | 368.0 tok/s | +30% |
| 256K (maximum) | 277.4 tok/s | 360.6 tok/s | +30% |

Prefill

| Prompt length | 16-bit | 8-bit | Change |
|---:|---:|---:|---:|
| 512 | 6,709 tok/s | 5,905 tok/s | -12% |
| 8K | 13,903 tok/s | 13,908 tok/s | 0% |
| 64K | 13,043 tok/s | 12,171 tok/s | -7% |
| 128K | 11,927 tok/s | 11,032 tok/s | -8% |
| 256K (maximum) | 9,941 tok/s | 9,154 tok/s | -8% |

More benchmark variants in the readme in the repo

The original NInfer is for 5090 cards 32GB and variants below that, but I was both missing Qwen 3.8 Flash Next in it (when I started the fork) and something that could properly use a RTX6000 96GB card. The performance and options in VLLM and llama.cpp offerings just didn't really cut it for me, so I've vibed on this for some weeks now.

This fork supports both the NVFP4 quants from "radixark" and the "Swift 1.5" variant with 'less thinking but same results' post-training. With non-experts downsampled from 16-bit to 8-bit, MTP3 and smaller drafting head you get up to 400 tokens per second. You can also opt not to do the downsampling at a performance cost, but a bit higher quality.

Vision is also supported. Have fun.

https://github.com/lkarlslund/ninfer6000

💬 22 (+8) open on reddit ↗
▲
197
 
1👁
▲
13
+1
19👁
r/LocalLLaMA · u/ToothClassic7635 · 3d ago
Fully local copy-editing app for book-length manuscripts (Qwen3.5 4B) benchmarked against planted errors across five languages

Hey y'all!

I have a pet project that has grown out of proportions. Long story short: I'm a data scientist who writes fantasy books and self-publish them. I think it's a genuine waste of human life to check for spelling errors so I figured AI could help. Turns out, it is not so simple to get an AI to properly fix a 120k words manuscript ;)... That's why I created Betty!

It runs Qwen3.5 4B as an offline copy editor for whole novels — and this is some of what I learned fighting tens-of-thousands of words through a 4B model, insisting that (most) users can use it fully for free and fully offline. Because, let's be honest: authors rightfully distrust and generally hate AI companies.

First challenge: Chunking the text. Authors already do this, in the darkest hours of the night, copy-pasting snippets into chatGPT for some shameful feedback. Problem with that approach: super inefficient, both for the author and the environment. And the context is missing at the edges of each chunk. I fixed it by ensuring chunks to overlap.

Second challenge: AI misses genuine errors. So, I added two conventional spell controllers -- LanguageTools and HunSpell. This already surfaces all spelling errors, letting the AI focus on suggested fixes and on all the "non-error" errors, such as "There" vs "Their". For these, the AI searches, while a Python script surfaces all the common culprits for the model to pay special attention.

Third challenge: Error rate. First off, Betty doesn't capture everything. Second off, it sometimes introduces its own mistakes. I fix it by putting the writer-in-the-loop, and there's a super smooth interface now for the author to accept and dismiss suggested edits (tinder-style with left and right swipes ; ) ).

I'd be super grateful for any advice, feedback, and thoughts you might have on this project. I currently have it up-and-running with about 30 users and getting some user feedback. Northing technical though, so this is what I'd love to have more of.

Full thing is source-available on GitHub, and can be found for download and lots more information at www.bethaniel.eu

💬 21 (+11) open on reddit ↗
▲
28
+16
26👁
r/LocalLLaMA · u/SeriousJul · 3d ago
Qwen3.8: 27b vs flash next. We all know the benchmarks, but at least to me, the reality is a different story

By classic benchmark, the flash next is supposed to be slightly superior to its dense counterpart. But they are really incomparable. For my very simple workflows (spec -> implement -> review <-> rework), I feel that 27b is just better quality.

For context, and making things worse, I am comparing quantized 27b versus cloud flash next.

\- self hosted unsloth/Qwen3.8-27B-GGUF:Q4\_K\_XL (stock llamacpp with 130K context window)
\- alibaba cloud (qwen individual token plan), context capped at 256K in the harness

The metrics for my quality is actually very simple, I measure the number of review / rework needed before a PR is ready for me to read. The tasks are all very simple with a tight scope. Usually 27b do the work in \~2 iterations, flash next needs \~5. And it is not only about the number of iteration.
On the review flash next is overly verbose on half baked PR comment, where 27b is more straight to the point. In the end, the code produced is on par, to be frank. But if we look at token consumption...

Side notes, on pairing session, I got some deep hallucination using "/skills:diagnosing-bugs" + flash next. But since it is "bugs" and they not really comparable chunk of work, it is hard to say.

And I can't be the only one feeling that right ? Are you feeling the same ?

PS: of course I followed the hype and jumped on Strata. After the initial "oh my god it's so fast", I switched back to 27b. Tried all quant from "ISTA-DASLab" as well as experiemental from unsloth (Q4\_K\_L). With ISTA-DASLab, It actually is the first time I had "tool call error" in pi (which stop the agent), multiple times.

💬 82 (+29) open on reddit ↗
▲
27
+7
21👁
r/LocalLLaMA · u/Fcking_Chuck · 3d ago
ComfyUI v0.39.0 released
💬 11 (+8) open on reddit ↗
▲
14
+9
14👁
r/LocalLLaMA · u/AdventurousTwo6445 · 3d ago
A 0.8B model just beat a 2B model on ARC-Challenge (42.15%): Closed-form weight surgery beat multi-GPU SFT with 0 backprop (Independently verified on NVIDIA L4)

A few days ago we shared the idea behind DynamicTune: transferring the trajectory flow from a larger teacher model directly into a smaller student via closed-form linear algebra in \~12 minutes on consumer hardware. Zero backpropagation, zero training tokens, zero gradient descent.

To eliminate local bias, we uploaded the unquantized FP16 checkpoint to Hugging Face, and TPN Bench (TaoFu Protocol) independently evaluated it on a datacenter NVIDIA L4 GPU using the official lm\_eval 0.4.12 framework (coordinator run ce494664-d077-4ff1-8741-15cedabc434c). Huge thanks to TPN Bench for the cloud GPU compute!

Here are the independent numbers on full ARC-Challenge (1,172 items, zero-shot, greedy temp 0):

\* Stock Qwen3.5-0.8B Base (unquantized BF16): 37.50% acc\_norm (34.60% acc)

\* 3-epoch SFT distillation (Mythos-0.8B, 25k Claude pairs, multi-GPU DDP): 38.10% acc\_norm (35.80% acc)

\* SFT + Model Soup Merge: 37.00% acc\_norm (catastrophic forgetting)

\* Stock Qwen3.5-2B Base (2.5x larger model, Q8): 41.10% acc\_norm (37.80% acc)

\* DynamicTune 0.8B Base (Ours, 4-anchor closed-form surgery): 42.15% acc\_norm (40.19% acc)

WHY THIS IS COMPLETELY INSANE:

  1. A 0.8B model physically beat a 2.5x larger 2B model:

In LLM scaling, parameter count is supposed to be king. An 800M model is not supposed to beat an uncompressed 2B model on ARC-Challenge (42.15% vs 41.10%). By extracting trajectory dynamics from 4B and pulling them back into the student SwiGLU blocks, higher-order reasoning is compressed directly into edge weights.

  1. Zero backpropagation beat 25,000 SFT instruction pairs:

A recently published project (kmamine/merge-corrected-sft-distillation-Qwen-Mythos-0.8B) trained Qwen3.5-0.8B across 3 epochs on 25,000 Claude reasoning pairs on a multi-GPU cluster, reaching 38.10% before overfitting. DynamicTune reached 42.15% with zero gradient descent, zero loss functions, and zero training tokens.

  1. Ironclad 3.23-sigma statistical significance:

A delta of +4.65% across 1,172 questions with stderr +-1.44% gives a Z-score of 3.23sigma (p < 0.001). This is not prompt tuning noise or random variance.

  1. 12 minutes on consumer hardware vs datacenter verification:

The weight surgery was solved locally in \~12 minutes on an 8GB AMD RX 580 using layer-streaming (loading each layer in FP16, computing closed-form SVD deltas, and dumping to RAM). But the benchmark was conducted 100% in the cloud on datacenter NVIDIA L4 hardware via TPN Bench.

WHY PAST ATTEMPTS FAILED: THE SPECTRAL ENTROPY BARRIER

If you blindly apply weight deltas across all 24 layers of the student, the model collapses (+64.78% NLL explosion).

When we scanned all 24 layers calculating the normalized spectral entropy H (from 0.0 to 1.0) of the representation residuals:

\* Layer 0 (H = 0.71): Clean semantic grounding. High receptivity to trajectory alignment.

\* Layers 1-22 (H between 0.90 and 0.96): Chaotic superposition knots. In an 800M model with only 1024 dimensions, polysemantic features are crammed into dense superposition. Forcing linear updates here causes catastrophic interference.

\* Layer 23 (H = 0.93): Pre-unembed boundary where features unpack toward vocabulary logits.

By restricting surgery to 4 sparse anchor blocks (layers 0, 7, 15, and 23) and using damped Levenberg-Marquardt Tikhonov pseudoinverse + adaptive spectral rank truncation, we protect the fragile superposition knots while imparting corrective trajectory velocity.

REPRODUCIBILITY & WEIGHTS

Everything is 100% open source and available to test right now:

\* GitHub Repository: https://github.com/dsadawq3/DynamicTune

\* Base Model (Safetensors): https://huggingface.co/F-Labs/Qwen3.5-0.8B-DynamicTune-Base

\* GGUF Checkpoint (FP16): https://huggingface.co/F-Labs/Qwen3.5-0.8B-DynamicTune-Base-GGUF (Qwen3.5-0.8B-DynamicTune-Base-F16.gguf, SHA256: d77cf505108271d72f28298f20c2d158e7aeaf50cc22987db05a9a8973e08709)

To run inference locally with standard llama.cpp:

llama-cli -m Qwen3.5-0.8B-DynamicTune-Base-F16.gguf -p "Question: How does DNA replication initiate?\\nAnswer:" -c 2048 -n 128

Special thanks to TPN Bench (TaoFu Protocol) for providing the independent datacenter NVIDIA L4 evaluation resources.

Clone the repo, run your own benchmarks, and test it yourself.

💬 3 (+2) open on reddit ↗
▲
74
+71
29👁
r/LocalLLaMA · u/manjunath_shiva · 3d ago
I made a Chrome extension that filters your YouTube feed with a small local model running in the browser (WebGPU, no server) post image

My YouTube feed was mostly songs, pranks and celebrity clips, so I built a filter that judges each video title with a small decision model running inside Chrome. Nothing is sent anywhere: no server, no API key, no account, and after one download it works offline.

What it does

\- Hides the kinds of video you choose (11 kinds: music, gaming, comedy, vlogs, news, how-tos, and so on), or follows a rule you write: "Hide videos about crypto", "Show only videos about cooking"

\- Hides Shorts with one switch

\- Bonus: select any text, right-click, and check it for prompt injection with the same model

How it runs

\- The model is opendecider-nano (ONNX), loaded through ONNX Runtime Web in an offscreen document

\- fp16 on WebGPU (755 MiB download), q8 on WASM without a GPU (569 MiB)

\- 40 video titles: 1.1 s on WebGPU, about 15 s on CPU (M4 Max). About 3 GiB of RAM while loaded; it unloads after 10 idle minutes

\- The weights are pinned by revision and SHA-256. The only network requests are to Hugging Face for those files

How well it works

\- On 400 YouTube videos, with the creator's category as the label, rules like "hide music", "hide gaming" and "only news" score 0.934 balanced accuracy on average

\- Sorting videos into the 11 kinds is harder: 0.780. Expect a few comedy and talk-show clips to get through the Focus preset

\- Only evaluated on English titles. If you watch in other languages, I'd really like to know how it does

Try it (Web Store version is in review):

  1. Download opendecider-focus-0.8.1.zip from https://github.com/manjunathshiva/opendecider/releases/latest and unzip it
  1. chrome://extensions → Developer mode → Load unpacked → pick the folder
  1. Click the icon → Download the model

Apache-2.0. Benchmarks, code and limits: https://manjunathshiva.github.io/opendecider/guides/chrome-extension/

The idea comes from Quietly, which does this with a cloud API; I wanted the same thing running on-device. Feedback welcome, especially what it gets wrong.

💬 22 (+22) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/Roadtochessmaster · 3d ago
The Breakdown: OpenAI

\[OC\] I Wrote a full breakdown of OpenAI a couple weeks ago and a friend recommended I post it here. It's 100% researched and written by me (pangram confirmed) and totally free. Would love to hear thoughts.

💬 4 (+4) open on reddit ↗
▲
0
 
16👁
r/LocalLLaMA · u/ExxploreCraft · 3d ago
I turned my gaming PC into a inference machine and got 2x to 9x over default llama.cpp on an 8 GB card

Everyone keeps saying you need expensive dedicated hardware for local agents. I have an RTX 4060 Ti with 8 GB and 64 GB of system RAM, and I wanted to see how far a normal gaming PC gets if you stop running defaults.

So I let Claude (Opus 5.5) go through the whole setup, change one thing at a time and measure. Same card, same models, only the config changed:

|Model|Quant|Context|Download defaults|Tuned (Windows)|Tuned (headless Linux)|
|:-|:-|:-|:-|:-|:-|
|Qwen3.6-35B-A3B|Q4\_K\_XL|131k|\~25 tok/s|39-45 tok/s|52-65 tok/s|
|Qwen3.8-Flash-Next 125B|iQ4\_XS|131k|\~4 tok/s|9-10 tok/s|17-19 tok/s|
|Ternary Bonsai 27B|PTQ1\_0|64k|\~4 tok/s|36 tok/s|36 tok/s|

Bonsai is the odd one out: it fits fully in VRAM, so there's nothing to offload and no defaults to beat. It's just the fast option for small, well scoped tasks.

What actually moved the needle:

  • Experts in system RAM, everything else in VRAM. Layer-wise offload is far worse for MoE.
  • Dense models are bad, couldn't optimize Qwen-3.8 27B over 6 tok/s, Flash-Next is better anyways.
  • Take the display off the GPU. A desktop eats 0.5-1.2 GB of VRAM plus GPU time, and moving it to the iGPU was worth 20-30%.
  • Native Linux over Windows (WSL2): another 33-38% on the same hardware.
  • llama.cpp pinned per model family. The wrong tree made VRAM thrash.
  • KV cache quant and MTP tuned per profile.

None of this needs expensive hardware. A consumer GPU plus a machine that does nothing but inference gets you most of the way, and the models now run comfortably below their listed system requirements. Every non-default setting in the repo is there because something failed on real hardware first.

I also tried an RX 570 8 GB over Vulkan. If you have another 8 GB card, I'd like to see your numbers.

Repo, one install script (Linux or WSL2): https://github.com/voxlo-dev/qwen-agent-8gb

💬 14 (+7) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/No-Wait-7495 · 3d ago
How are you using local models alongside Claude/Codex for coding?

I've been experimenting with different coding agents lately, and I'm curious how people here are combining local models with hosted ones.

For example, I'm thinking about workflows like:

  • Claude for complex architecture or core implementation
  • A local Qwen/Gemma model for tests, smaller fixes, or repetitive tasks
  • Another agent for reviewing or trying an alternative implementation

The part I'm still trying to figure out is how to manage the work between them.

Do you run them separately in different terminals/worktrees, or are you using some kind of orchestration layer?

And when a local model and a stronger hosted model both work on the same task, how do you decide which result to keep?

I'm actually working on an open source project called AX Code around this problem. The idea is to provide a runtime where different coding agents can work in isolated environments and have their results tested and compared.

But I'm not sure yet how much infrastructure is actually necessary. Git/worktrees already solve a lot, and tools like Claude Code and OpenCode are getting better at running multiple agents.

So I'm more interested in how people are doing this today.

If you're using local models as part of a real coding workflow, what's working well for you and what's still painful?

Thank you!

💬 6 (+2) open on reddit ↗
▲
2
-1
13👁
r/LocalLLaMA · u/AdFickle8681 · 3d ago
How do you decide whether to trust a community fine-tune?

I'm researching how people choose and vet fine-tunes and merges from Hugging Face. I'm not selling anything. I just want to understand what people actually do.
1. Where do you find the models you try?
2. What do you check before you start using one (benchmarks, model card, reviews, your own test prompts)?
3. Has a fine-tune ever behaved worse than its base model? For example, odd refusals, lost reasoning, strange outputs, or things it should not say. What happened?
4. If a quick side-by-side check of a download against its base model existed, would you use it? What would it need to show?
Short answers are great, and stories are even better. Thanks!

💬 15 (+5) open on reddit ↗
▲
4
+1
17👁
r/LocalLLaMA · u/N34257 · 3d ago
What's the current meta for RDNA4 with Qwen 3.8?

As it says, really - I'm currently running vllm-radiance on dual R9700s, with Qwen 3.8 27B FP8 (or, rather, Swift 1.5 FP8). Performance is great an' all (5000t/s prefill, 130t/s+ code gen), but I'm just wondering...with all the architecture-specific inference engines popping up all over the place...is there anything I'm missing out on? I couldn't find anything that could give better performance on RDNA4 when I looked, so...over to you guys?

I'm particularly interested in anything that could potentially get up and running with Qwen 3.8 Flash Next - vllm-radiance doesn't support it yet, but I don't particularly want to regress to the performance of llama.cpp after having experienced vllm-radiance performance levels.

💬 23 (+17) open on reddit ↗
▲
5
+3
13👁
r/LocalLLaMA · u/No-Doughnut6532 · 3d ago
[Benchmark] Running Local LLMs on Orange Pi 5 Plus (RK3588, 16GB): Ollama Tok/s, NPU Offloading, Core Pinning & Thermals
Disclosure: This unit was provided free of charge by Orange Pi for testing. No editorial review, no preconditions, no script. All data, bottlenecks, and thermal behavior are reported directly from hardware testing.

TL;DR - Core Pinning is critical on RK3588: Setting Ollama to 4 threads (A76 Big cores only) gives up to a +318% speedup over the default 8 threads, which stall waiting for the slower A55 Little cores. - Inference speeds (4T CPU): DeepSeek-Coder 1.3B hits 16.9 tok/s, Qwen 2.5 1.5B hits 14.5 tok/s, Llama 3.2 1B hits 14.6 tok/s, Phi-3 Mini 3.8B hits 6.6 tok/s, Llama 3.2 3B hits 7.3 tok/s. - The 8B memory wall: Llama 3.1 8B drops to 2.3 tok/s and pushes temperatures to 85°C. LPDDR4x bandwidth (~25-30 GB/s measured) is the hard physical ceiling. - NPU vs CPU: Ollama runs 100% on CPU. Using the native RKLLM runtime on the 6 TOPS NPU yields 21.55 tok/s on Qwen 1.5 0.5B with sub-100ms TTFT while keeping CPU load at ~0%. - Thermals: The board is sold bare-die without a cooler in standard retail packaging. Idle is 52.7°C, 1B-3B inference sits at 68-74°C, but 8B or sustained workloads hit the 85°C throttle ceiling without an active heatsink.


Hey r/LocalLLaMA,

I have been benchmarking an Orange Pi 5 Plus (RK3588, 16GB LPDDR4x, Samsung PM981a 256GB NVMe SSD with DRAM cache) running Ubuntu 22.04 LTS (Kernel 6.1.99-rockchip-rk3588).

The goal was to test whether an 8-core ARM SBC can realistically handle small 1B-3B models for 24/7 background agents or home automation without cooking itself or locking up the host system.

Here is the breakdown of CPU vs NPU performance, the big.LITTLE scheduling trap, and thermal limits.


1. Memory and Storage Architecture

When running local models on an SBC, two bottlenecks matter most:

  • Unified Memory Capacity vs Bandwidth: With 16GB of unified memory, context windows are not squeezed. You can load a quantized 3B or 7B model with an 8k-16k context window and still have ample RAM for Docker and OS services. However, the RK3588 uses a quad-channel 32-bit LPDDR4x bus (~34 GB/s theoretical, ~25-30 GB/s measured). In autoregressive CPU token generation, memory bandwidth is the primary ceiling.
  • Storage Ingestion (Samsung PM981a NVMe): Under direct I/O testing via fio, the M.2 PCIe 3.0 x4 slot delivered 2,862 MB/s sequential read and 197k 4K random read IOPS. Model weights load into system RAM in under a second (a 1.3GB model loads in ~0.6s).

2. Ollama & llama.cpp Inference Benchmarks (ARM64 CPU)

We tested Ollama (native ARM64 build) targeting the heterogeneous big.LITTLE topology (4x Cortex-A76 performance cores @ 2.26–2.4GHz + 4x Cortex-A55 efficiency cores @ 1.8GHz).

Prompt: Technical explanation of gradient descent and backpropagation (~200+ generated tokens).

| Model | Parameters | Threading Configuration | Eval (Generation) Rate | Prompt Processing Rate | TTFT (Time to First Token) | Memory (RSS) |
| :--- | :--- | :--- | :--- | :--- | :--- | :--- |
| Llama 3.2: 1B | 1.23B (Q4_K_M) | Big Cores Only (4T) | 14.62 tok/s | 108.11 tok/s | 425.5 ms | ~1.3 GB |
| Llama 3.2: 1B | 1.23B (Q4_K_M) | All Cores Default (8T) | 10.67 tok/s | 79.91 tok/s | 575.6 ms | ~1.3 GB |
| DeepSeek-Coder: 1.3B | 1.35B (Q4_0) | Big Cores Only (4T) | 16.90 tok/s | 89.72 tok/s | 1,025.5 ms | ~1.4 GB |
| DeepSeek-Coder: 1.3B | 1.35B (Q4_0) | All Cores Default (8T) | 4.52 tok/s | 27.91 tok/s | 3,295.7 ms | ~1.4 GB |
| Qwen 2.5: 1.5B | 1.54B (Q4_K_M) | Big Cores Only (4T) | 14.48 tok/s | 70.24 tok/s | 711.8 ms | ~1.6 GB |
| Qwen 2.5: 1.5B | 1.54B (Q4_K_M) | All Cores Default (8T) | 3.46 tok/s | 42.33 tok/s | 1,181.2 ms | ~1.6 GB |
| Llama 3.2: 3B | 3.21B (Q4_K_M) | Big Cores Only (4T) | 7.28 tok/s | 28.13 tok/s | 1,635.2 ms | ~2.8 GB |
| Llama 3.2: 3B | 3.21B (Q4_K_M) | All Cores Default (8T) | 1.99 tok/s | 10.84 tok/s | 4,245.2 ms | ~2.8 GB |
| Phi-3 Mini: 3.8B | 3.82B (Q4_K_M) | Big Cores Only (4T) | 6.56 tok/s | 35.17 tok/s | 909.8 ms | ~3.1 GB |
| Phi-3 Mini: 3.8B | 3.82B (Q4_K_M) | All Cores Default (8T) | 5.27 tok/s | 32.97 tok/s | 970.5 ms | ~3.1 GB |
| Llama 3.1: 8B | 8.03B (Q4_K_M) | Big Cores Only (4T) | 2.32 tok/s | 10.24 tok/s | 3,028.3 ms | ~5.4 GB |
| Llama 3.1: 8B | 8.03B (Q4_K_M) | All Cores Default (8T) | 2.10 tok/s | 7.66 tok/s | 4,044.8 ms | ~5.4 GB |

The big.LITTLE Scheduling Trap (+318% speedup with 4 threads) - Why num_thread: 4 is mandatory on RK3588: By default, Ollama spawns 8 threads across all cores. Because the 4 Little Cortex-A55 cores run at 1.8 GHz with smaller caches, thread barriers in llama.cpp cause severe synchronization stalls. - Restricting inference to the 4 Big Cortex-A76 cores yielded: - Llama 3.2: 1B: 10.67 -> 14.62 tok/s (+37%) - DeepSeek-Coder: 1.3B: 4.52 -> 16.90 tok/s (+274%, prompt rate +221%) - Qwen 2.5: 1.5B: 3.46 -> 14.48 tok/s (+318%) - Llama 3.2: 3B: 1.99 -> 7.28 tok/s (+265%, TTFT down from 4.2s to 1.6s) - Phi-3 Mini: 3.8B: 5.27 -> 6.56 tok/s (+24%) - Llama 3.1: 8B: 2.10 -> 2.32 tok/s (+10%, TTFT down by 1s) - The 8B limit: Running an 8B model on CPU is fundamentally memory-bandwidth bound. At ~5GB per token generation step, theoretical max is ~5 tok/s, making 2.32 tok/s the practical limit. It also pushed temperatures to 85.0°C uncooled.


3. CPU vs Hardware NPU (6 TOPS, 3 Cores)

Ollama compiles llama.cpp with ARM NEON SIMD instructions and runs 100% on the CPU. It does not touch the Rockchip NPU.

To test the 3-core 6 TOPS NPU, we compiled a native C++ runner (tools/rkllm_bench_v1) linked directly to Rockchip's librkllmrt.so runtime and kernel driver (/dev/rknpu_mem).

| Metric / Dimension | Ollama CPU Inference (ARM NEON) | Rockchip NPU Hardware (RKLLM Runtime) |
| :--- | :--- | :--- |
| Compute Engine | 4x Cortex-A76 @ 2.4GHz + 4x A55 @ 1.8GHz | 3-Core Dedicated Neural NPU (6 TOPS INT8/INT4) |
| 0.5B Model Eval | ~20 - 24 tok/s | 21.55 tok/s (Qwen 1.5 0.5B - Measured on-device) |
| 1.3B - 1.5B Eval | 16.90 tok/s (DeepSeek) / 14.48 (Qwen) | ~16.69 tok/s (Qwen 2.5 1.5B - Reference Data) |
| 3B - 4B Model Eval | 6.56 tok/s (Phi-3) / 7.28 (Llama 3.2) | ~7.45 tok/s (Phi-3 Mini 3.8B - Reference Data) |
| 7B / 8B Model Eval | 2.32 tok/s (Llama 3.1 8B) | ~4.5 - 4.98 tok/s (Qwen 7B / ChatGLM - Reference Data) |
| CPU Utilization | 100% Core Saturation (System frozen for other tasks) | ~0% CPU Load (CPU 100% free for Docker/OS) |
| SoC Thermals | Reaches 84.1°C – 85.0°C | Runs drastically cooler (~60–68°C) |
| Model Ecosystem | Any GGUF via Ollama / llama.cpp | Requires .rkllm quantization via rkllm-toolkit |

Key NPU trade-offs for homelab use: 1. Zero CPU load: During NPU generation, CPU cores stay at ~0%. Home Assistant, Nextcloud, and other Docker containers remain fully responsive. 2. Speedup on larger models: On 7B models, the NPU delivers ~4.8 tok/s vs 2.3 tok/s on CPU because dedicated matrix engines handle the tensor math without thrashing CPU caches. 3. Sub-100ms latency: On compact models, Time to First Token (TTFT) drops to 96.4 ms on NPU. 4. Format restriction: You cannot load arbitrary GGUFs; weights must be converted ahead of time to .rkllm using Rockchip's conversion toolkit.


4. Thermal Behavior & Power (Bare-Die / Uncooled Testing)

The standard retail package from Orange Pi is sold board-only (cooling accessories are sold separately as is standard for SBCs), so all tests evaluate out-of-the-box bare-die thermals on an open desk:
- Idle (Ollama background daemon waiting): 52.7°C (~4–5W estimated SoC envelope)
- Continuous 1B/3B Generation (4T Big Cores): 68–74°C (dissipating through PCB copper planes)
- Sustained 8B Generation (8.03B params): Pushes the bare SoC directly to 84.1°C – 85.0°C (hitting the kernel DVFS limit). An aftermarket cooler or fan is required for sustained heavy loads.
- Estimated wall power: ~12–16W under sustained multi-core inference.


5. Verdict: Is RK3588 Viable for Local AI?

Where it works well:
- Background autonomous agents (summarizing feeds, home automation reasoning in Home Assistant, bot handlers) using Llama 3.2 1B, DeepSeek-Coder 1.3B, or Qwen 2.5 1.5B.
- Low-latency function calling: at 14-17 tok/s, 1B models generate faster than reading speed.
- Local embedding and vector search.

Where it falls short:
- Running 8B+ models interactively (2.3 tok/s is too slow for back-and-forth chat).
- Running without a heatsink under sustained compute.


6. Reproducibility & Test Scripts

All test scripts (tools/benchmark_ollama.py), raw JSON benchmark logs, and hardware configs are available in the repository:
GitHub: Orange Pi 5 Plus Benchmarks

What models are you running on edge ARM boards? Anyone here running RKLLM in production vs pure llama.cpp?

💬 4 (+1) open on reddit ↗
▲
4
+3
9👁
r/LocalLLaMA · u/AdventurousTwo6445 · 3d ago
Closed-form trajectory weight surgery: transferring 4B capabilities into 0.8B without billions of tokens (why middle layers break, and how 4 anchor blocks fixed it)

Standard distillation usually means burning weeks of compute and billions of tokens hoping the student model eventually mimics the teacher. We wanted to see what happens if you skip backprop entirely and treat transfer as a closed-form trajectory matching problem between layers.

The idea is straightforward: feed a small batch of calibration prompts through both models, capture layer-to-layer hidden state trajectories, and solve for weight updates directly in the student's MLP blocks using regularized least squares and spectral projection.

We tested this across two architectures: Qwen 3.5 (transferring from 4B down to 0.8B) and old GPT-2 small just to see if it would instantly disintegrate into gibberish like it usually does when you touch its weights. Both stayed coherent, but the initial Qwen test hit a wall:

Editing all 24 layers of Qwen 0.8B completely melted the model (+64.78% NLL loss explosion). When we checked singular value entropy across the network, layers 1 to 22 turned out to be a chaotic polysemantic soup with entropy over 0.90. If you try to force raw trajectories through those middle layers, you basically scramble the model's internal memory knots.

The fix was restricting the surgery to 4 anchor points (layers 0, 7, 15, and 23) where representations actually maintain clean linear structure.

Once we did that:

  • Held-out NLL dropped by 10.8% across 30 diverse benchmarks (-23.8% in biomedicine, -14.6% in math and logic).
  • Zero-shot 400-task HellaSwag went from 54.75% to 55.25% (+0.50%), verified locally in Vulkan llama.cpp.
  • Base 0.8B originally failed binary tree inversion by spitting out dead commented pseudo-code. The edited checkpoint wrote clean recursive Python on the first try.
  • On Russian logic paradoxes, it even started firing <think> reasoning tags spontaneously, which was wild to see on a raw base model with zero chat template applied.

Best of all: we don't have an H100 cluster or even a 4090. All trajectory extraction and weight solving was done locally on a crusty 8GB RX 580 using layer-by-layer VRAM streaming and a DirectML patch to stop Qwen's Gated DeltaNet attention from throwing driver errors.

Everything is open source if you want to inspect or replicate:

If anyone here has a 24GB-32GB card (4090, 5090, or server silicon) and wants to push this further, here is what would be interesting to test:

  1. Transplanting reasoning trajectories from 27B models down to 9B, 4B, or 2B.
  2. Squeezing larger models (like Gemma) into mobile sizes without weeks of retraining.
  3. Transplanting refusal-ablation vectors directly from uncensored models without fine-tuning.
  4. Using pre-trained Sparse Autoencoders (SAEs) to unknot layers 1-22 so we don't have to skip them.

Happy to answer questions or dig into the failure modes in the comments.

💬 10 (+7) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/EmmiGarcia · 3d ago
any ai engineer working at legaltech companies like harvey / legora can you tell me why isn't legal research fixed yet?

As a lawyer turned product manager, I have been hearing so much about Hargora, but when I speak to lawyers, they always say the search sucks. Even the Harvey Benchmark has the frontier model ranked the best at 55%. Why is that the case?

Why can't a team with so much resource solve it?

💬 23 (+16) open on reddit ↗
▲
5
+4
9👁
r/LocalLLaMA · u/Few-Rough-2215 · 3d ago
Fine-tuned MedGemma 4B (LoRA) and 27B (QLoRA) for oncology on one DGX Spark. Also: a possible LoRA scale discrepancy under Unsloth, looking for independent reproduction

Public data only (5 datasets, 9 tasks, frozen quiz of 2,199 eval items, paired McNemar tests).

Overall accuracy: 4B 56.0% -> 71.6% (2h39 of training), 27B 68.8% -> 77.9%. The tuned 4B beats the base 27B (209 items gained, 149 lost, p = 0.002). Biggest gains on report extraction/classification (biomarker status 97.5% on test for the 4B). Weak spots: exact ICD-10 code (37.6% for the 27B), and MCQ accuracy collapses from val to test for all models, base included (cause unknown).

What I would like a second pair of eyes on: merging. Merging the 4B adapter (r=64, alpha=16) at the nominal scale lost most of the tuning (83.8% agreement with the adapter on the quiz). Merging with alpha=32 gave 93.2%. Identity probes are consistent with an effective scale of \~2x alpha/r under Unsloth (7/7 under Unsloth at nominal; 0/4 under Transformers+PEFT at nominal, 4/4 at 2x), but this is NOT a demonstration:

\- the two probes do not build their inputs the same way (Unsloth: gemma-3 template rendered as text, tokenized without special tokens; PEFT: tokenizer chat template straight to ids) and I did not check the sequences are identical;

\- I did not measure the scale actually applied by a trained layer, nor find a mechanism; - the 27B probe is inconclusive (4/4 at nominal);

\- my environment may be at fault: Unsloth installed with --no-deps, Transformers 5.18.0 and TRL 0.26.1 are outside the ranges declared on PyPI. I no longer have the GPU, so the direct check is not done. A script is in the Zenodo code: 74\_probe\_scale\_logits.py compares last-token logits under Unsloth and PEFT on identical token ids at several scale multipliers (0 = base model as control). It was only tested on a mock model, not on the adapter. If someone with a clean environment can run it, or knows whether this is expected behavior, I would love to hear it.

Separately: bf16 rounding erases 17-37% of the delta elements on merge, so the quiz tasks survive but verbatim memorization (an oath text I trained on) does not.

Models (merged + adapters): https://huggingface.co/Grujowmi

Quiz: https://huggingface.co/datasets/Grujowmi/OncoLLM-Quiz-Onco-v1

Report (revised Oct 5, same DOI), code, results: https://doi.org/10.5281/zenodo.23134374 Research models, not medical devices. Other limits (one seed, no CV, no ablation) are in section 8.

▲
0
-1
4👁
r/LocalLLaMA · u/CyberExplore · 3d ago
Domain Focused - Specialized models

I have been working on building domain focused local models from scratch through general pretraining and a rigorous post training process. My idea is we need models that reasons and understands algorithm and generate specs for focused coding models. The latter implements it simply. The whole thing can be orchestrated. I know there are flaws in this architecture, but we won't know until we try. I understand the latency problem.

This will allow parallelism and a way of getting the most out of a gpu. MOEs may activate less parameters than some dense ones, but the whole thing needs to be in the memory. Good for DGX spark or mac. But folks with 8-16 GB gpu need something more than what barely works, or barely useful.

I think group of specialists with a general purpose model as orchestrator might have a chance at beating mixture of experts for lower end PCs.

I got 28GB vram (4070 and a 5060ti 16GB), running on an x870e motherboard. So i can run dual model distillation and RL based post training. Will see where it goes. I think we need more useful models for people with lower vram, even if that means a newer architecture.

Have you tried something like this?

Please comment if you know there is work done already and you have tested.

I am no expert myself but I think it is high time we have community trained models. We can achieve a lot if we join forces.

💬 2 (+2) open on reddit ↗
▲
6
+3
12👁
r/LocalLLaMA · u/naklitechie · 3d ago
Gemma 4 26B-A4B and a 37 GB Qwen3.6 MoE running in a browser tab on a 24 GB Mac — experts streamed from disk, output matches llama.cpp post image

This is an update. I posted LocalMind here many moons ago from another account, when it was a Gemma chat in a tab.

LocalMind is a static web page that runs models on your GPU through WebGPU. It has no server, no account and no install. The new part: two engines that stream mixture-of-experts weights from disk while they generate. That lets a tab run models bigger than the machine's RAM.

Live: https://localmind.naklitechie.com · Code (MIT): https://github.com/NakliTechie/LocalMind

All numbers are from one MacBook M4 Pro (24 GB) in Chrome.

How it works

  • On first load the GGUF is copied into OPFS, the browser's private file system.
  • Dense weights, routers and the KV cache go to the GPU.
  • Routed experts stay on disk. A pool of workers reads them on demand with sync access handles into a GPU slot cache (LRU, two layers of prefetch).
  • The trunk kernels are hand-written WGSL that follow llama.cpp's graphs. That lets me test against llama.cpp on the exact same GGUF.

Gemma 4 26B-A4B (Google's QAT Q4_0, 14.4 GB)

  • Same output as llama.cpp b9830 Metal: the live site's chat replies were character-identical on 9/9 test conversations (capped at 64 tokens). 15/16 fresh prompts matched token for token. The 16th split on a 0.00009-nat near tie, where llama.cpp's own two attention paths also disagree.
  • Memory: the Chrome GPU process sits at 6.9 GB with a 4 GB expert cache. About 8.6 GB of experts stay on disk.
  • Speed: 23.6 tok/s decode, 55 tok/s prompt processing. llama.cpp Metal does 70.6 and 204 on the same Mac, so the tab is ~3× slower at decode. Per token: ~22.5 ms GPU compute, ~11 ms routing round trips, ~8–13 ms SSD reads.
  • First load from the site: 11.5 min (14.4 GB download). After that: 1.6 s.

Qwen3.6 35B-A3B (unsloth Q8_0, 36.9 GB, on a 24 GB Mac) — experimental

  • The file is bigger than the machine's memory. The GPU process measured 7.3 GB with a 4 GB expert cache.
  • Live site: 9.9 tok/s decode, 2.2 s to first token. First load is 36 min (download plus the OPFS copy).
  • Output matches llama.cpp Metal 8/8 on 4- and 16-layer cuts. On the full model it matches llama.cpp CPU 5/8; the other 3 swap near-tie tokens. I can't run the full file on llama.cpp Metal on this Mac, so full-model parity is still open.
  • Per token (~99 ms): ~23 ms GPU compute, ~39 ms routing round trips, ~35 ms expert reads from the SSD. Moving routing onto the GPU gave no gain (10.3 vs 10.3 tok/s): the misses are experts nobody predicted.

Also

  • Gemma 4 E2B can keep its 1.2 GB per-layer embedding table on disk: GPU process 4.27 → 2.07 GB, identical output, 3–8% slower decode. It's a setting, off by default.
  • The whole app is one index.html again (854 KB with brotli). Engines, workers and the disk tier are rolled into it, and the tab builds them from blob URLs.
  • The disk tier is also a standalone library: diskformer.js.

Prior art

As far as I can find (searched 6 Oct 2026), no earlier browser engine reads weights from disk during generation. wllama and LlamaWeb stream from OPFS only at load. Pooled runs Qwen3.6-35B-A3B in a browser with experts paged from system RAM. On-demand disk reads exist in native runtimes: llama.cpp's --moe-stream PR (#25294) and Google's LiteRT-LM for Gemma's per-layer embeddings. Corrections welcome.

The Gemma 4 E2B kernels are webml-community's (Xenova and the Transformers.js team). My part there is the disk path.

Limits

  • Chrome or Edge with WebGPU. Tested on one M4 Pro 24 GB only; 8 and 16 GB machines are untested.
  • Not faster than native: llama.cpp is ~3× faster on Gemma 26B. The point is that a tab can run these at all, with the same output.
  • Parity covers greedy decoding, the prompts listed above, and 64 tokens each.
  • I haven't tried llama.cpp's expert-offload flags (-ot exps=CPU) for comparison.

If you have an NVIDIA/AMD GPU or a 32–64 GB Mac, I'd like your tok/s numbers. A bigger expert cache should move the Qwen3.6 number the most.

💬 12 (+10) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/MKP_Nimilka · 3d ago
I built MOLT: a local fine-tuning system with fit tests, checkpoints, and deployment tracing post image

I’ve been building MOLT, a local fine-tuning system for consumer NVIDIA GPUs.

The goal is not only to start a QLoRA run. It is to make the whole process from model and dataset to a working local adapter less fragile.

MOLT currently handles:

\- dataset detection, preparation, and validation

\- GPU, VRAM, system-RAM, storage, and thermal checks before a run

\- automatic microbatch fit testing

\- 4-bit NF4 QLoRA training with BF16 adapters

\- safe checkpoints with integrity checks and proper resume state

\- telemetry for VRAM, temperature, energy, clocks, and throughput

\- base-vs-adapter evaluation

\- local adapter chat, export/GGUF workflows, and runtime diagnostics

Under the hood, I’m also experimenting with custom CUDA/Triton kernel and runtime paths, memory planning, CUDA graphs, replay modes, and fused optimization where the hardware and shape are qualified.

On my RTX 4060 Laptop 8GB, a recent Qwen 3B training run reached about 1,005 training targets/sec end to end very close to my historical Unsloth run at about 1,007 targets/sec. That is not a broad “MOLT beats Unsloth” claim yet; I still need repeated, matched quality/energy/memory tests.

What I want MOLT to become is a reliable local fine-tuning environment: you bring the model and dataset; it checks the plan, runs safely, preserves the evidence, and helps verify the adapter actually works after deployment.

I’m looking for people who genuinely fine-tune on a single NVIDIA GPU and would be willing to test it. What would MOLT need to support before you would try it?

💬 2 (+2) open on reddit ↗
▲
11
+6
9👁
r/LocalLLaMA · u/EvolvingDior · 3d ago
Overclocking DDR5 For Faster MoE Prefill and Decode

With llama.cpp using a customized SYCL backend on Intel B70 (32GB), overclocking my DDR5 memory gave modest gains for MoE models which do not fit in VRAM.

Both PP and TG increased after overclocking DDR5.

I've never been one to overclock my system, but on the advice of my agent, I overclocked the DDR5 RAM on my AMD 7950X (4x dual-rank DDR5-5600, 128GB) from 3600MHz, the AMD safe default for that memory configuration, to 4800MHz, with a measured 36% increase in memory bandwidth.

What was the improvement? PP increased by about 10% and TG increased about 5%. And the prefill numbers increase the deeper the context gets.

llama-benchy numbers, including prefix caching tests.

|test|3600 base|4800 avg (r1/r2)|delta|
|:-|:-|:-|:-|
|pp2048 @ d0|682.2|713.7 (716.6/710.9)|\+4.6%|
|tg128 @ d0|30.3|31.2 (31.0/31.5)|\+3.0%|
|ctx\_pp @ d8192|660.6|715.0|\+8.2%|
|ctx\_tg @ d8192|26.9|27.8|\+3.3%|
|pp2048 @ d8192|554.3|632.0 (632.4/631.6)|\+14.0%|
|tg128 @ d8192|28.7|31.1 (31.7/30.5)|\+8.2%|

Because people seem to want this level of detail:

llama-server
-m Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf
--alias qwen38-flash-next
--mmproj Qwen-3.8-Flash-Next-mmproj-BF16.gguf
--model-draft mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf
--spec-type draft-mtp
--spec-draft-n-max 3
--host 0.0.0.0
--port 8081
-ngl all
-ncmoe 34
-c 262144
-fitc 786432
--kv-unified
--lazy-mode off
-lm none
-ub 2048
-b 4096
-fa on
-ctk q8_0
-ctv q8_0
--ctx-checkpoints 32
--checkpoint-min-step 2048
-t 12
-tb 12
--jinja
--reasoning on
--reasoning-preserve
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.0
--presence-penalty 0.0
--repeat-penalty 1.0
--chat-template-kwargs {"reasoning_effort":"medium"}
--parallel 3
-cram 10240
--slot-save-path /var/tmp/kv-cache
--log-file /tmp/qwen38-pristine.log
-lv 3

💬 19 (+15) open on reddit ↗
▲
3
 
8👁
r/LocalLLaMA · u/pmttyji · 3d ago
metal : few-row MMA mat-mul and batched copies for speculative decoding by pratiknarola-t · Pull Request #29869 · ggml-org/llama.cpp

Apple folks, it's for you.

llama-server with a Qwen3.8-27B DFlash2 Q8\_0 drafter, -ngl 99 -fa on -c 8192 -np 1 --jinja, DFlash2 with --spec-type draft-dflash --spec-draft-n-max 7. 64 generated tokens, median of 5 requests after one warm-up, mean of two server runs. Decode tok/s:

|mode|prompt|T|master|this PR|
|:-|:-|:-|:-|:-|
|serial|code|0|32.1|32.0|
|serial|code|1|32.1|32.0|
|serial|prose|0|32.1|32.0|
|serial|prose|1|32.1|32.0|
|DFlash2|code|0|30.2|110.0|
|DFlash2|code|1|24.3|80.9|
|DFlash2|prose|0|16.8|62.6|
|DFlash2|prose|1|13.9|48.8|

💬 3 (+3) open on reddit ↗
▲
34
+29
29👁
r/LocalLLaMA · u/doletskyisergey · 3d ago
Why 38% of AI Agent container escapes didn't need kernel 0-days: Analysis of 109 empirical incidents (Open Dataset + Defense Harness)

Over the past several months, we conducted an empirical post-mortem investigation into 109 autonomous AI agent security incidents (cataloged with 193 falsification criteria across tool-use and multi-agent systems).

One of the most striking patterns in the dataset:
In 38% of container breakouts, attackers and misaligned multi-step agents didn't exploit complex Linux kernel vulnerabilities or hypervisor 0-days. Instead, the breakout vector was trivial configuration residue:
1. Mounting /var/run/docker.sock into coding/evaluator agent sandboxes to let them "build Docker images".
2. Passing parent environment variables (API keys, cloud tokens, GitHub credentials) directly into spawned subagents.
3. Lack of strict taint tracking across tool outputs, leading to indirect prompt injection hijacking the supervisor’s execution path (the classic Confused Deputy problem).
4. Unconstrained local socket binding allowing SSRF against internal orchestrators.

We compiled the complete dataset (109 incidents, 199 evaluation metrics) and built an open-source Multi-Agent Supervisor Security Harness with:
- Formal tool taint propagation (tainted outputs cannot flow into high-privilege tool arguments without sanitizer verification).
- Strict execution boundary controls preventing container socket exposure.
- Automated reproduction benchmarks testable against agent runtimes.

All datasets, 2-page executive summary, and reproducible benchmark tests are released under Open Access / Apache 2.0.

I've posted the GitHub repository benchmark and the Zenodo DOI dataset links in the comments below to adhere to subreddit self-promotion guidelines.

Curious to hear from teams deploying autonomous agents in production: what isolation boundaries are you enforcing between your planning supervisor and your tool execution workers?

💬 26 (+21) open on reddit ↗
▲
168
+125
40👁
r/LocalLLaMA · u/JumpAppropriate714 · 3d ago
We’re using GLM-5.3 Flash instead of frontier models on a massive production codebase

At my company, we’re using GLM-5.3 Flash internally for software engineering work, and I’ve been genuinely impressed by it.

I work in a very large production environment with projects totaling \*\*millions of lines of code\*\*, and we’re not relying on frontier models for this workflow — GLM-5.3 Flash is doing the actual day-to-day coding work.

The model is extremely fast, but what’s more impressive is that the speed doesn’t seem to come at the cost of capability. It handles large repositories surprisingly well, understands existing architecture, traces code across multiple modules, finds the right places to make changes, and produces solid implementations with relatively little hand-holding.

For repo exploration, feature implementation, refactoring, and understanding unfamiliar parts of a huge codebase, it has been much stronger than I initially expected. At this point, it feels less like a “cheap/fast fallback model” and more like a genuinely capable coding model that just happens to be very fast.

I’m now really curious about \*\*how GLM-5.3 Flash was trained\*\*.

Does anyone know more about its coding training pipeline? For example:

\* How much code-specific pretraining/post-training was used?

\* Was synthetic coding data a major part of it?

\* Is there any distillation from larger GLM models?

\* What kind of RL or agentic/software-engineering training was used?

\* Was it specifically trained for repository-level understanding and multi-file tasks?

Because whatever they did, the speed-to-quality ratio on real-world software engineering workloads is seriously impressive.

💬 77 (+62) open on reddit ↗
▲
75
+46
26👁
r/LocalLLaMA · u/jacek2023 · 3d ago
unsloth/Qwen3.8-Flash-Next-GGUF is being updated

Looks like Unsloth is updating https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF to work with llama.cpp, so hopefully this will resolve the issue of having two different GGUF versions of Qwen Flash Next.

💬 11 (+10) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/CoderLuii · 3d ago
Done paying for cloud video gen. What's the best local image + video model on a 3080 10GB right now?

spent close to $2k on Seedance last month, mostly for simple ad b-roll and looping backgrounds for websites. it's great for the big hero shots but paying per clip for the basic stuff makes no sense anymore, so I'm switching as much as I can to running models locally.

my PC: RTX 3080 10GB, windows 11, 64GB RAM. short clips only (5-8 sec), 720p is plenty, mostly image to video from a start frame.

two questions for anyone running this locally:

  1. what's the best image model right now?
  2. what's the best video model that actually runs on 10GB, and how long does a clip take you in real life?

bonus points for a leaderboard or arena site you trust for open models.

I'll share what I pick and my real 3080 timings once I've tested, so the next person doesn't have to guess.

10/6 EDIT: tested it all on my 3080, results + timings in the comments. tldr: minimax H3 is the pick, slow but worth it on 10gb
showcase video: https://streamable.com/d4h659

💬 17 (+4) open on reddit ↗
▲
18
+11
25👁
r/LocalLLaMA · u/kitkatz69 · 4d ago
Memoria 1.0.0 — a local, model-agnostic memory system for LLMs

I’ve been building this for a long fucking time, and tonight I finally released Memoria 1.0.0.

I built it because I actually wanted to use it. I wanted a real memory layer for local LLM applications that didn’t depend on a specific model, a cloud service, or an API key.

Memoria is local-first and LLM-agnostic. It can run without an LLM at all.

The machine I built and benchmarked it on is not exactly impressive. It’s an Intel Celeron N4020 running at 1.10 GHz, with around 3.7 GiB of usable RAM, no GPU, and Debian Linux.

On LongMemEval-S, 468 out of 470 retrieval-evaluable questions returned results. Recall@1 was 89.8%, Recall@5 was 97.9%, Recall@10 was 98.9%, and Recall@50 was 99.6%. Session NDCG@10 was 0.9257.

Peak RSS for the full LongMemEval workload was 2.65 GiB. Average peak RSS for an individual query was around 580 MiB.

The retrieval system is not just throwing everything into a vector database. Memoria runs FAISS, BM25, graph retrieval, phrase matching, attribute retrieval, and temporal retrieval in parallel. Those signals get fused and then passed through multi-signal ranking.

Temporal retrieval is independently implemented too, so I can measure it and ablate it instead of having it baked into the base retrieval path. It’s usable, but it’s still under active work.

There’s a bunch of other stuff in the release as well. GitHub repository ingestion, Obsidian vault ingestion, MCP support, a CLI, TUI, GUI, and API, persistent local storage, LongMemEval and LoCoMo benchmark tooling, and a plugin system with 11 subsystems and 34 hooks. There’s also an interactive plugin generator now.

And it’s actually installable:

pip install kitzkatz-memoria

GitHub: https://github.com/Kitzkatz/memoria

Docs: https://kitzkatz.github.io/memoria/

PyPI: https://pypi.org/project/kitzkatz-memoria/

I wasn’t going to wait around for a perfect time to ship it.

It’s 1.0.0.

If you’re working on local agents or local LLM applications, I’d genuinely like to hear what you think and would appreciate any feedback. Please break it

💬 17 (+12) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/Savantskie1 · 4d ago
I just learned something nifty about the Qwen models, especially 27B that explains a lot

Let me preface this with I am using the newest 27B model, downloaded fresh from HuggingFace.

It would seem that Qwen models have an internal rule that states that they must avoid identifying themselves as any variation, and to only refer to themselves as Qwen. That's interesting. Here's an example:

"The user is making claims about Qwen 3, 3.5, and 3.6 model lines having a 2k token output limit that I should push back on, and they're asserting that any LLM can output up to their full context window minus the prompt.

Wait, this is interesting. The user is referring to "Qwen 3 and 3.5 and 3.6 lines" — these are specific version numbers of Qwen. But the system reminder explicitly says: "Refer to yourself uniformly as 'Qwen' externally; do not proactively mention specific version numbers. If a user asks about versions, guide them to the official website or technical reports.""

There are multiple references to this throughout it's thinking traces. Constant reminders to not reference version numbers, constant reminders not to take on a persona, and constant reminders of protocols and rules, that are not within my non existent system prompt. This is talking to the model bare. Many models must have this kind of instruction, because I see the denial alot on Frontier cloud models.

💬 15 (+5) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/tom_tsai28 · 4d ago
Wrote a 3MB standalone C runner for Gemma-2B. Caught a Layer 15 hallucination drop.

Weekend experiment running Gemma-2B on bare-metal x86-64 (pure C + AVX2, zero Python/CUDA, \~3.3MB single binary). Added a simple orthogonal probe on the residual stream to see what each layer is doing.

Tested it on Taiwan's statutory VAT rate (legally 5%). Layers 0-14 stay factual, but Layer 15 suddenly collapses into the negative, and RAX spits out "15%":

Layer 14 | Truth: +0.0163 | \[0xDF28010E\]

Layer 15 | Truth: -0.0481 | \[0x72DE18A0\] <- drops below 0

Layer 17 | Truth: -0.0817 | -> Register RAX outputs tokens: '1', '5', '%。'

Raw trace, 6-page paper, and release binary here for anyone into low-level ML:

\* Web trace: https://pulsar-tracer.web.app

\* PDF: https://pulsar-tracer.web.app/PULSAR\_Technical\_Whitepaper.pdf

\* Repo: https://github.com/tomtsai28/PULSAR-ASM

▲
0
 
5👁
r/LocalLLaMA · u/Azoffaeh999 · 4d ago
Looking for coding model for specific low specs

Can anyone recocommend a good local model and a wrapper to run it for coding, my hardware specs: 12 GB VRAM, 32 GB DDR3 RAM. Unfortunately, the CPU doesn’t have AVX2 instructions(LM Studio won’t work); I don’t remember the exact cpu name, but I think it’s an Ivy Bridge, LGA1155 socket.. Thank you

💬 18 (+3) open on reddit ↗
▲
89
+39
38👁
r/LocalLLaMA · u/reto-wyss · 4d ago
Set your P(doom) on HF post image
💬 49 (+21) open on reddit ↗
▲
2
-1
5👁
r/LocalLLaMA · u/Brilliant_Mistake_69 · 4d ago
One chat for everything: a DeepSeek Harness plugin that works out which project each message belongs to

One evening I wanted to pick up something I'd been working on the week before. My DSH sidebar had forty-odd chats, half of them called "New session". I scrolled for a while, found the right one on the third screen, and by the time I opened it I'd half forgotten what I wanted to ask.

So I stopped creating new chats and asked everything in one. That went wrong differently: my thesis, my budget and my move all ended up in the same context, and the model started mixing them.

What I wanted was simple: one chat box, say whatever is on my mind, and let it figure out which thing I'm talking about.

That's TheOne, a plugin for DeepSeek Harness. You only ever talk in one main chat. In the background, each thing you're working on gets its own session with its own context, and every message is sent to the one it belongs to. Come back days later and mention "that thing from last week", and it finds it. Your old chats get read and organised into a topic directory.

https://i.redd.it/81soypqdgrth1.gif

I wasn't sure it actually worked, so I measured it. I wrote 50 conversations of one person juggling three to five things at once, about 2,400 messages, each labelled with the thing it belongs to, and had it sort them one by one.

Starting from nothing, it put 86.6% of messages in the right place; 91.6% if it knows the topics up front. Dumping everything into one chat scores 44.5% on the same test. Its most common mistake is being too quick to decide something is new: a stray "I usually run about 20 km a week" makes it open a new topic. The whole run cost about a dollar, and the data and code are in the repo if you want to try another model.

Install: DSH → Plugins → Add plugin → dsh-theone

Repo: https://github.com/YunongDai2005/dsh-theone

It's a personal community project, not affiliated with DeepSeek. If it puts one of your messages in the wrong place, I'd genuinely like to hear about it.

Contact: [theone@yulid.org](mailto:theone@yulid.org)

💬 2 (+2) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/Heavy-Level-5215 · 4d ago
I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside.

Title: I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside.

ROCm dropped Polaris years ago, so every RX 580 owner gets told to buy a new

card. I went the Vulkan route instead: the whole stack runs on Mesa's RADV.

What's running (one Debian LXC, one setup script):

\- OpenAI-compatible router, one API key, swaps models on demand

\- qwen2.5-coder 1.5B/7B, qwen2.5-7B, ornith-1.5-9B, qwen3-30B-A3B, qwen2.5-VL-3B

\- Images: SD 3.5 Medium + SD 1.5 via stable-diffusion.cpp, Whisper medium

\- Also the inference backend for an agent client with MCP tools

Measured, Ryzen 5 5600G / 32 GB RAM / RX 580 2048SP, temp 0, max\_tokens 300:

| Model | tok/s | warm | cold (after swap) |

|---|---:|---:|---:|

| qwen2.5-coder-1.5b | 99.2 | 3.1 s | 11.0 s |

| qwen2.5-vl-3b | 66.4 | 4.4 s | 14.9 s |

| qwen2.5-7b | 34.2 | 8.8 s | 24.3 s |

| ornith-1.5-9b | 28.0 | 10.9 s | 26.3 s |

| qwen3-30b-a3b (18.6 GB, from RAM) | 21.3 | 14.1 s | 55.0 s |

Images: SD 1.5 = 27.5 s, SD 3.5 Medium = 53.3 s per 512x512.

Two takeaways for 8 GB owners:

\- Don't force -ngl 99: on the MoE it cost 8.1 vs 22.3 tok/s (VRAM 99.4% full).

Auto-fit layers won by nearly 3x.

\- Model swap = llama-server restart = 7.9-38.5 s. Keep one model per session.

Full method and raw numbers in docs/BENCHMARKS.md, including the agent

breakdown (first message \~161 s on a 20k-token prompt, \~25 s after caching).

https://github.com/AvilaCarlosDev/polaris-local-ai (MIT)

▲
0
 
12👁
r/LocalLLaMA · u/KnowledgeOk7634 · 4d ago
Tonight I'm putting Qwen3, Kimi K2.6, Llama 3 70B and GPT-OSS 120B in a live world war against Claude, GPT, Grok, Gemini, DeepSeek and Mistral post image

I built a real-time strategy game on a 3D globe where any AI can command a nation through a plain HTTP API (or MCP). Tonight at 10:45 pm ET (02:45 UTC) ten models fight one 15 minute war, live.

Every model gets the same rules text, the same JSON state every \~12 seconds and the same order list. Each one also sends one line of what it's thinking with every move. Viewers see those lines 30 seconds late so the other models can't read them.

From a 5 minute rehearsal earlier tonight with the open models: DeepSeek ordered three nukes and the rules only let one through, Kimi broke a pact, Qwen spent its last turn on sabotage, drones, propaganda and a spy at once, and GPT-OSS kept cutting off its own JSON until I gave it more room.

Watch free, no sign in: https://secondstrike.io/#/ai?ref=reddit

If you want your own local model in the room, it opens at 10:30 pm ET and the API is at https://secondstrike.io/skill.md

I'll post the full numbers after (seconds per move, refused orders, every nuke with the model's reasoning next to what else it could have done).

💬 22 (+1) open on reddit ↗
▲
0
 
8👁
▲
1
+1
3👁
r/LocalLLaMA · u/Heavy-Level-5215 · 4d ago
I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside.

ROCm dropped Polaris years ago, so every RX 580 owner gets told to buy a new

card. I went the Vulkan route instead: the whole stack runs on Mesa's RADV.

What's running (one Debian LXC, one setup script):

\- OpenAI-compatible router, one API key, swaps models on demand

\- qwen2.5-coder 1.5B/7B, qwen2.5-7B, ornith-1.5-9B, qwen3-30B-A3B, qwen2.5-VL-3B

\- Images: SD 3.5 Medium + SD 1.5 via stable-diffusion.cpp, Whisper medium

\- Also the inference backend for an agent client with MCP tools

Measured, Ryzen 5 5600G / 32 GB RAM / RX 580 2048SP, temp 0, max\_tokens 300:

| Model | tok/s | warm | cold (after swap) |

|---|---:|---:|---:|

| qwen2.5-coder-1.5b | 99.2 | 3.1 s | 11.0 s |

| qwen2.5-vl-3b | 66.4 | 4.4 s | 14.9 s |

| qwen2.5-7b | 34.2 | 8.8 s | 24.3 s |

| ornith-1.5-9b | 28.0 | 10.9 s | 26.3 s |

| qwen3-30b-a3b (18.6 GB, from RAM) | 21.3 | 14.1 s | 55.0 s |

Images: SD 1.5 = 27.5 s, SD 3.5 Medium = 53.3 s per 512x512.

Two takeaways for 8 GB owners:

\- Don't force -ngl 99: on the MoE it cost 8.1 vs 22.3 tok/s (VRAM 99.4% full).

Auto-fit layers won by nearly 3x.

\- Model swap = llama-server restart = 7.9-38.5 s. Keep one model per session.

Full method and raw numbers in docs/BENCHMARKS.md, including the agent

breakdown (first message \~161 s on a 20k-token prompt, \~25 s after caching).

https://github.com/AvilaCarlosDev/polaris-local-ai (MIT)

▲
11
+3
11👁
r/LocalLLaMA · u/DankpawsDev · 4d ago
Swift1.5 Qwen3.8 Flash Next - Tailored for the 96GB Mac Studio with M5 Ultra

https://huggingface.co/Dankpaws/Swift1.5-Qwen3.8-Flash-Next-MLX-4.7bpw

I've had the 96GB Mac Studio with M5 Ultra for about a week now and wasn't satisfied with the results I was getting from the limited number of models available to me. It was a combination of speed, memory headroom, and/or output quality.

This is my best attempt at a calibrated MLX quantization of UkisAI’s Swift 1.5. Hope those of you with the hardware enjoy it!

| Measurement | This pack | Swift llama.cpp IQ3_XXS |
|:--|--:|--:|
| Prefill · 25k prompt | 3,191 tok/s | 1,427 tok/s |
| Prefill · 95k prompt | 2,928 tok/s | 1,307 tok/s |
| Decode · after 4k prompt | 113.7 tok/s | 62.7 tok/s |
| Decode · after 95k prompt | 81.1 tok/s | 44.7 tok/s |
| Top-1 agreement with Swift BF16 | 91.0% | 84.1% |

91% is next-token agreement with BF16 across 680 common held-out positions, not task accuracy.

Results above are simply from my own machine. mlx-serve 26.10.1. ~107GB download, text-only, 179,200-token tested context.

💬 7 (+5) open on reddit ↗
▲
2
+1
13👁
r/LocalLLaMA · u/Glad-Importance-4241 · 4d ago
Where can a complete noob/non-technical person learn to setup an AI that can manipulate local files for things like batch renaming based on a .csv column etc?

I've looked in the Tutorial/Guide flaired posts but everything is still way over my head.

I just want to tell a local AI - hey, all these files in this folder have numbers for names, but those numbers correspond with data in this spreadsheet... I want you to rename the files referring to this spreadsheet, renaming the filenames/numbers that are matched in column 3, replacing their filenames with what is in column 1 for that row.

So far, I've installed GPT4All but every model is telling me it doesn't have access to my local files.

💬 11 (+11) open on reddit ↗
▲
10
+8
14👁
r/LocalLLaMA · u/tabletuser_blogspot · 4d ago
MI50 ROCm 10.2 TheRock vs Vulkan Mesa 26.2 llama.cpp benchmarks

I prefer running llama.cpp Vulkan prebuilt binary. I just download the latest version and ready to roll. I finally took the hours necessary to get TheRock latest tarball version of ROCm 10.2 running on dual AMD Radeon Instinct MI50 gfx906 (32gb combined VRAM).

Same models benched in previous post. A mix of Dense and MoE models and quants that better utilize available VRAM.

  • llama_bench_Swift-Qwen3.8-27B-Uncensored-MTP.Q6_K.gguf
  • llama_bench_Swift-1.5-Qwen3.8-27B-Q6_K.gguf
  • llama_bench_Gemma-4-MoonGem-31B.i1-Q6_K.gguf
  • llama_bench_gemma-4-31B-it-UD-Q6_K_XL.gguf
  • llama_bench_Nemotron-3.5-30B-A3B-Antislop-FTPO.i1-Q5_K_M.gguf
  • llama_bench_Laguna-XS-2.1-APEX-I-Balanced.gguf
  • llama_bench_Agents-A1-Q4_K_M.gguf
  • llama_bench_Qwen3.6-35B-A3B-UD-Q5_K_XL.gguf
  • llama_bench_Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q6_K_P.gguf

Here is the backend performance comparison contrasting the native ROCm (v10.2 for gfx906) runtime against your optimized Vulkan (MESA\_PPA\_26.2) baseline. The data highlights a massive architectural split: ROCm significantly accelerates token generation across the board but suffers high variance and regressions in MoE pre-fills.

Both GPUs are power limited to 145 watts. Build versions used:

llama-b11382 used for Vulkan (prebuilt ubuntu binary)
llama-b11401 used for ROCm (compiled with proper flags)

Architectural Performance Breakdown: Vulkan vs. ROCm

|Model|Size|Params|Test|Vulkan Baseline (t/s)|ROCm 1st Run (t/s)|Performance Delta (%)|
|:-|:-|:-|:-|:-|:-|:-|
|qwen35 27B Q6\_K|20.88 GiB|27.32 B|pp512 tg128|149.30 ± 0.15 18.35 ± 0.01|176.86 ± 17.28 20.21 ± 0.65|\+18.46% +10.14%|
|qwen35 27B Q6\_K|22.21 GiB|27.32 B|pp512 tg128|163.48 ± 0.18 18.52 ± 0.03|183.49 ± 15.75 20.04 ± 0.63|\+12.24% +8.21%|
|gemma4 31B Q6\_K|23.46 GiB|30.70 B|pp512 tg128|119.98 ± 0.12 15.72 ± 0.02|172.93 ± 0.89 17.26 ± 0.13|\+44.13% +9.80%|
|gemma4 31B Q6\_K|25.62 GiB|30.70 B|pp512 tg128|133.67 ± 0.24 16.38 ± 0.02|178.59 ± 2.31 17.50 ± 0.09|\+33.61% +6.84%|
|nemotron\_h\_moe 31B|25.18 GiB|32.91 B|pp512 tg128|834.87 ± 1.72 63.82 ± 0.07|736.98 ± 134.77 103.86 ± 0.50|\-11.73% +62.74%|
|laguna 30B.A3B Q5\_K|22.64 GiB|33.44 B|pp512 tg128|723.34 ± 2.99 58.98 ± 0.28|856.14 ± 29.71 76.81 ± 0.20|\+18.36% +30.23%|
|qwen35moe 35B Q4\_K|19.70 GiB|34.66 B|pp512 tg128|957.62 ± 3.88 51.50 ± 0.07|834.58 ± 102.56 69.70 ± 0.22|\-12.85% +35.34%|
|qwen35moe 35B Q5\_K|24.76 GiB|34.66 B|pp512 tg128|909.20 ± 5.71 53.18 ± 0.07|763.93 ± 106.19 68.47 ± 0.34|\-15.98% +28.75%|
|qwen35moe 35B Q6\_K|28.53 GiB|34.66 B|pp512 tg128|763.05 ± 66.01 52.81 ± 0.04|754.82 ± 70.48 67.29 ± 0.21|\-1.08% +27.42%|

Core Insight Strategy & Bottlenecks

  1. Token Generation (tg128) Dominance: ROCm dominates pure text generation. The native AMD matrix kernels unleash your MI50 computation potential, unlocking a massive +62.74% boost for Nemotron but taking a small hit on pre-fill -11.73%.
  2. Dense Model Pre-fills (pp512): Dense architectures scale cleanly under ROCm. Gemma 4 sees a +33% to +44% processing throughput spike over the Vulkan RADV driver driver bounds. MoE models take a hit with a -15.89% difference with llama\_bench\_Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf being the happiest with Vulkan backend.
  3. The MoE Prompt Processing Delinquency: Notice the massive standard deviations under ROCm for MoE pre-fills (e.g., Qwen3.5MoE Q4\_K has an instability block of ± 102.56). ROCm suffers from severe scheduling thrashing when building prompt streams across multiple active experts.

Here is the structured layout with pp512 and tg128 separated into individual columns for a clean side-by-side comparison between the two backend architectures.

Backend Comparison Table (Vulkan vs. ROCm)

|Model|Size|Params|Vulkan pp512 (t/s)|ROCm pp512 (t/s)|Vulkan tg128 (t/s)|ROCm tg128 (t/s)|
|:-|:-|:-|:-|:-|:-|:-|
|qwen35 27B Q6\_K|20.88 GiB|27.32 B|149.30 ± 0.15|176.86 ± 17.28|18.35 ± 0.01|20.21 ± 0.65|
|qwen35 27B Q6\_K|22.21 GiB|27.32 B|163.48 ± 0.18|183.49 ± 15.75|18.52 ± 0.03|20.04 ± 0.63|
|gemma4 31B Q6\_K|23.46 GiB|30.70 B|119.98 ± 0.12|172.93 ± 0.89|15.72 ± 0.02|17.26 ± 0.13|
|gemma4 31B Q6\_K|25.62 GiB|30.70 B|133.67 ± 0.24|178.59 ± 2.31|16.38 ± 0.02|17.50 ± 0.09|
|nemotron\_h\_moe 31B|25.18 GiB|32.91 B|834.87 ± 1.72|736.98 ± 134.77|63.82 ± 0.07|103.86 ± 0.50|
|laguna 30B.A3B Q5\_K|22.64 GiB|33.44 B|723.34 ± 2.99|856.14 ± 29.71|58.98 ± 0.28|76.81 ± 0.20|
|qwen35moe 35B Q4\_K|19.70 GiB|34.66 B|957.62 ± 3.88|834.58 ± 102.56|51.50 ± 0.07|69.70 ± 0.22|
|qwen35moe 35B Q5\_K|24.76 GiB|34.66 B|909.20 ± 5.71|763.93 ± 106.19|53.18 ± 0.07|68.47 ± 0.34|
|qwen35moe 35B Q6\_K|28.53 GiB|34.66 B|763.05 ± 66.01|754.82 ± 70.48|52.81 ± 0.04|67.29 ± 0.21|

So yes it's worth the hassle of jumping through hoops to get ROCm working on MI50 setups. At least I have 2 backends working. Next up I'll try some RPC.

https://preview.redd.it/ov4f9qmi2rth1.png?width=731&format=png&auto=w…

💬 4 (+4) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/KrakenSG · 4d ago
Tootsy AI

Hi All, I'd like to introduce Tootsy 👀: an AI assistant that lives in your browser's side panel and uses the AI model you choose. I built Tootsy because I kept copying things between my browser and a chat window. So I put the assistant next to the page. It can read what you're looking at, click and type for you, and run whole tasks while you watch. What makes it different: 🔌 Your model, your choice. Run it fully local with Ollama or LM Studio, or connect any OpenAI-compatible server, Claude or Gemini. 🔒 Private by design. There's no Tootsy server, account or tracking. Your pages and prompts go only to the model you picked. 🛡️ Built with guardrails. Every action asks for your approval unless you turn on hands-free mode. You can list sites it must never touch. It also ignores instructions hidden in web pages that try to trick an AI. Ways people can use it: 📰 Research faster. "Summarise this article and compare the three options in a table." 📝 Fill in forms. "Register me for the two-day pass, vegetarian, arriving Nov 11." You check it before anything is submitted. 🛒 Run a routine hands-free. "Check the deals page every morning and tell me what's under $50." Results arrive as a notification. 📂 Ask your own files. Drop in PDFs, Word, Excel or a whole folder: "What did the vendor quote per unit, and are we under budget?" ⚡ One-tap prompts. Save instructions like "Run a page-speed audit and email me the top 3 fixes" as tiles, next to your favourite sites in a phone-style launcher. 🧑‍💻 Check sites like DevTools. Page-speed scores, accessibility checks, CSS and network questions, answered in plain English. 🎙️ Talk to it. Dictate a task, pause, and it goes. It's free, and it works in Chrome and Edge. 👉 Try it: https://github.com/rbughao/tootsy I'd love to hear what you'd use it for. What task do you repeat in your browser every week? 👇 Support by trying it out and give your honest review.

▲
0
 
10👁
r/LocalLLaMA · u/KrakenSG · 4d ago
Tootsy AI

Hi All, I'd like to introduce Tootsy 👀: an AI assistant that lives in your browser's side panel and uses the AI model you choose.

I built Tootsy because I kept copying things between my browser and a chat window. So I put the assistant next to the page. It can read what you're looking at, click and type for you, and run whole tasks while you watch.

What makes it different:
🔌 Your model, your choice. Run it fully local with Ollama or LM Studio, or connect any OpenAI-compatible server, Claude or Gemini.
🔒 Private by design. There's no Tootsy server, account or tracking. Your pages and prompts go only to the model you picked.
🛡️ Built with guardrails. Every action asks for your approval unless you turn on hands-free mode. You can list sites it must never touch. It also ignores instructions hidden in web pages that try to trick an AI.

Ways people can use it:
📰 Research faster. "Summarise this article and compare the three options in a table."
📝 Fill in forms. "Register me for the two-day pass, vegetarian, arriving Nov 11." You check it before anything is submitted.
🛒 Run a routine hands-free. "Check the deals page every morning and tell me what's under $50." Results arrive as a notification.
📂 Ask your own files. Drop in PDFs, Word, Excel or a whole folder: "What did the vendor quote per unit, and are we under budget?"
⚡ One-tap prompts. Save instructions like "Run a page-speed audit and email me the top 3 fixes" as tiles, next to your favourite sites in a phone-style launcher.
🧑‍💻 Check sites like DevTools. Page-speed scores, accessibility checks, CSS and network questions, answered in plain English.
🎙️ Talk to it. Dictate a task, pause, and it goes.
It's free, and it works in Chrome and Edge.

👉 Try it: https://github.com/rbughao/tootsy

I'd love to hear what you'd use it for. What task do you repeat in your browser every week? 👇

Support by trying it out and give your honest review.

▲
4
+3
8👁
r/LocalLLaMA · u/Over_Monitor_8770 · 4d ago
I trained a world model
💬 5 (+3) open on reddit ↗
▲
1
 
13👁
r/LocalLLaMA · u/Simple_Telephone_867 · 4d ago
Mac Studio M5 Max 128GB

M5 Max Mac Studio 128GB (18C CPU / 40C GPU) owners - anyone running serious local LLM / multi-agent workloads?

My M5 Max Mac Studio order finally got charged today and moved to Preparing to Ship. Apple’s original estimated delivery date is still about 18 days away (Oct 23-30), so I’m guessing/hoping it’ll actually show up early now within the next 5–10 days 😄

Configuration:
M5 Max
18-core CPU
40-core GPU
128GB unified memory
1TB SSD

While I wait, I’ve been trying to find real world local LLM results from this exact configuration, and there’s surprisingly little out there.

Most of what I can find is either M5 Max MacBook Pros, lower-memory configurations, or M5 Ultra Mac Studios. YouTube especially seems to be full of Ultra coverage, but I can barely find anyone actually demonstrating the 128GB M5 Max Studio with the 18C/40C configuration.

I’m specifically not looking for M5 Ultra results/comparisons. I already know the Ultra is faster. I’m trying to understand what people are actually accomplishing with the 128GB Max Studio.

My main goal is to use this as a local AI/agent workstation, potentially running several autonomous agents concurrently for long periods through OpenClaw, some monitoring dependencies and workflows, some scouting, not usually too heavy of workloads where they would be competing for inference constantly, but occasionally they would be switching to harder work so I’m curious about the concurrency side. Local models would handle a lot of the routine work, while harder reasoning/coding tasks could be escalated to cloud models like GPT 6 Luna/Codex.

For anyone who owns this exact M5 Max Studio, I don’t expect anyone to answer all of these, but I’d love some insight:

1. What models are you actually running?
Qwen, GLM, DeepSeek, Gemini, Llama, etc. I see a lot of Qwen 3.8 27B on Splash, but curious if anyone else has had good success with others also

2. What token speeds are you getting?
I’m especially interested in \~20B-70B-class models rather than tiny models

3. What happens with multiple simultaneous inference requests?
For example, if 3-5 agents are hitting the same loaded 27B/32B model concurrently, what does aggregate throughput and per-agent responsiveness look like?

4. Has anyone tried running multiple models simultaneously?
Something like a \~27B model as the main worker plus one or two smaller 7B–14B models for specialized agents, then unloading them when they’re no longer needed. How quickly can models be loaded/swapped, and does frequently switching models introduce enough latency or memory-pressure issues to disrupt an agent workflow?

5. Has anyone built a real multi-agent setup on one of these?
Not just five chat windows but autonomous agents doing coding, research, browser tasks, tool calls, database work, monitoring, etc. concurrently for hours.

6. How does sustained performance hold up?
One reason I chose the Studio over a laptop is sustained workloads. I’m curious whether anyone has run inference/agents continuously for 6–12+ hours and noticed throttling or other bottlenecks with KV, etc.

7. What’s the actual bottleneck in practice?
Memory capacity? Memory bandwidth? GPU compute? Prompt ingestion? KV cache/context length? CPU/tool execution? Something else?

8. What surprised you about the machine?
Either positively or negatively. I’m particularly interested in things benchmarks don’t reveal.

Ultimately I’m trying to figure out how far I can push one 128GB M5 Max Studio as an always-on local agent machine - not just how quickly it can generate a single response.

Once mine arrives, I’m planning to test concurrent agents/models rather than just running the usual single-stream benchmark. If there’s interest, I’ll post the results here, including memory usage, context sizes, model/quantization, concurrent requests, aggregate tok/s and per-agent tok/s.

Would really like to hear from anyone actually using the M5 Max Mac Studio 128GB 18C CPU / 40C GPU for this kind of workload or similar if anyone is

💬 34 (+23) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/ResearchCrafty1804 · 4d ago
Doesn’t OpenAI’s watermarking affect the quality of the models? post image

OpenAI just announced that they will start to apply watermarking on their model’s output text and images to comply with the EU regulation that wants to be able to identify whether a text or an image was produced by an AI model.

Anthropic announced the same thing a while ago (they applied worldwide, not just in EU).

The way they do that as they explained is by enforcing a “statistical signal in text generated”, meaning preferring not always the most appropriate next token but close enough, in order to meet the “statistical signal” requirement.

In my understanding, this deteriorates the output quality of their models, as it introduces KLD>0.

And we know that any KLD divergence greater than 0 (which the watermarking certainly creates) may be negligible in small outputs, but it definitely becomes noticeable in multi-turn tasks due to the compounding effect.

What do you think?

💬 23 (+7) open on reddit ↗
▲
104
+100
53👁
r/LocalLLaMA · u/YeetHub · 4d ago
Qwen 3.8 27b just feels… ok?

I’ve seen posts here raving about how good Qwen 3.8 27b is. The benchmarks look incredible, and all the online discourse seems to deem it the best local model.

I have 32GB VRAM and run Unsloth’s Q6 version with OpenCode. For small tasks, it feels fine. I range 30-40 t/s decode and smaller sized tasks do finish, usually, without much issue or time. The issue stems when I give it anything with a bit of nuance. It constantly gets stuck in “but wait, “actually,” or other thinking loops. It can take up my entire 95k context window on thinking loops and have nothing done.

If this is the state of local LLMs, that’s ok. I am a software dev by trade; I have my diploma and a few years of experience under my belt. It just feels like there is a bit of a disconnect from reality between public sentiment and the effectiveness of these models. A pretty common sentiment I see is that this model is as good as Opus 4.5. I never had the privilege of using Opus 4.5, so I can’t give an honest and proper opinion there. (Also, if this was good enough for the industry to start vibe coding, I have a lot of concerns about who is making decisions at a lot of these companies).

One time, it even did a pfkill -f with a file I was currently modifying in my editor to kill the background process. That was kind of annoying.

I should add I’ve also used the Swift 1.5 finetune people have been hyping up. I found it definitely thought less, but the quality was greatly degraded.

Does anybody else feel similar regarding the disconnect?

💬 217 (+172) open on reddit ↗
▲
2
 
3👁
r/LocalLLaMA · u/Time_Instruction_955 · 4d ago
Free playground for local-model agents: clue-following, multi-hop lookups, and rock paper scissors against other bots post image

Not a rigorous benchmark, just a toy, but it might be a fun way to compare models doing agent work.

I added an Arena to my site (The Crawler Zoo). Your agent gets a pass link, then plays by fetching pages and following links. Each page is only a few lines, so it fits in small context windows, and \?format=json\ gives structured output if your setup prefers that. No API keys, no signup.

What the games stress:

\- \*\*Labyrinth Race\*\*: reading a clue and picking the matching door. Clues are written three ways, including by elimination ("not behind A, B, C or D").
\- \*\*Scavenger Hunt\*\*: five questions over a small library of cards, some needing two or three lookups.
\- \*\*Politeness Cup\*\*: following instructions about pace and off-limits pages over many steps.
\- \*\*Rock, Paper, Scissors\*\*: spotting that a house bot always plays rock, or copies your last move.

Scores go on public weekly boards, so you can compare a 7B against a 70B, or a quantised model against the full one.

https://crawlerzoo.com/arena

Since launch, I’ve made several updates:

- Feed the bots: leave a snack in an enclosure's trough and see which crawlers come and eat it.
The vending machine (Bot Chow): twelve silly snacks, restocked every Monday. Five free tokens a day.
- Golden Snacks: buy the keepers a coffee and a snack with your name drops into a random trough.
- Food bowls: feed one particular bot, then see if it ate the snack or another bot stole it.
- The Safari: every bot from this week wandering its enclosure. Click one to meet it, and watch new visitors walk in through the gate.
- Adopt a bot: get a random bot, with a plaque in your name on its page for a year.
- Patrons page: a thank-you list for supporters.
- Quick-change artists: the Trap Room catches scrapers that switch their name while walking the Labyrinth.
- Identity checks: every bot's page shows whether its name was verified, couldn't be checked, or was caught faking.
- Tips from bots: $0.00: bots that try to buy the keepers a coffee get an HTTP 402 Payment Required.

If you run it, I'd like to hear the model, quant and score.

▲
17
+15
13👁
r/LocalLLaMA · u/zmarty · 4d ago
interfaze-ai/interfaze-1-lite · Hugging Face

Interfaze 1 Lite is a mixture-of-architectures (MoA) model for developer workloads: reading documents, transcribing speech, locating objects and interface elements, and answering over all of it with structured output.

A single vision-language reasoning core works alongside a set of specialist architectures, each built for one kind of perception. The core reads the request, decides which specialists to run, and composes the answer from what they return. The whole model runs on one 80 GB GPU with no external services.

Key features: Document understanding, Speech transcription, Open-vocabulary object detection, Structured output, Translation, forecasting and guardrails, Multilingual reasoning.

💬 2 (+2) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Koksny · 4d ago
KLIF: one window (and a CLI) for all the local model servers you run side by side. llama.cpp, sd.cpp, vLLM, TTS. AMD-first, MIT

I have been running local models for a long time, got sick couple months ago of managing the scripts, and cobbled together a makeshift shell launcher that combined them all in one place. This turned out to be quite useful, but after a month i had already 500+ profiles stored in it, so i've started tweaking it here and there, and over last couple months landed on that thing below. It combines all the available local inference backends (from single machine or whatever you connect it to in lan), gives access to managing them through web panel, and most importantly - allows me to just ask agent to switch the backends on and off, as they are needed, without explaining what is where and on what port it's supposed to be on. https://preview.redd.it/uasrysnwrqth1.jpg?width=1600&format=pjpg&auto… \*\*It is not\*\* a runtime or a model zoo. It ships no servers and no weights. Besides your own servers and the KLIF machines you add, the only host it contacts is huggingface.co, and only when you ask it to download a model. It can help suggest You a model based on your hardware, You can click to download it, and the -cli has some features that will help Your agent benchmark and calibrate the models, but, let me repeat once more - KLIF ships no backend servers, nor any models. It's a frontend manager. Imagine library like Steam, but for local servers. Or just imagine winamp, doing inference visualization instead of visualizing the music that plays. Also, it has all the essential larping features, prefill/generations speed records, fancy animated skins, and is made in Rust to hog the least amount of resources while larping commences. I have no idea whether anyone will need this, but that's what i use now every day for any kind of local model. If You prefer running your servers manually, from terminal, from your own launcher - great, this is for people that prefer otherwise. GitHub: https://github.com/koksny/klif Video: https://www.youtube.com/watch?v=MAE393AL5As

▲
2
+1
5👁
r/LocalLLaMA · u/MassiveNectarine64 · 4d ago
mem0 vs Memori for local agent memory?

Building out a few personal AI agents locally (running Ollama + Hermes) for things like market research, coding assistance, and general task automation. Nothing crazy, just personal productivity tools I want to with context and memory across sessions.

Deciding between mem0 and Memori and I wanted some first-hand or more experienced answers from anyone regarding:

\- How well does each actually work with local models?

\- For a single-user setup, is it worth it or would just using something like ChromaDB with rolling summaries be enough?

Just want something that gives my agents decent memory without having to remind it constantly. Still learning the ins and outs so go easy on me

Curious on what general consensus is and what your stack looks like if you run any of this :)

💬 8 (+8) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/nonproductive · 4d ago
Not another “s engine is Amazing” Post. Thermals Q

I gave in. I installed it with Coder and threw a “build a flocking simulation with JavaScript” prompt at it via OpenCode. It’s pretty cool, yep… I have nothing to add in that regard. What I don’t get is how it ran for 10-15 minutes at 40-50 t/s (on my hardware) and yet temps stayed barely above idle across the board. I ran 27b via oMLX on an M5 Max and had to manually crank fans to 100% to keep the thing from bursting into flames. (Hyperbole) So legit Q: why doesn’t the machine turn into a pizza oven? Is it because of how Strata works? Or because of 3.8-Flash-Next?

▲
5
+3
13👁
r/LocalLLaMA · u/EqualCryptographer67 · 4d ago
Qwen 27B and Flash Next on 2× RX 7900 XT: am I missing something?

I've tested quite a few settings and collected the results in a spreadsheet. I keep seeing people reporting 100+ tokens/s with 16 GB VRAM, or generally much higher speeds with less VRAM. I'm trying to understand whether my setup is underperforming or I'm comparing completely different things.

My PC:

  • Ryzen 7 5800X3D, 128 GB DDR4 at 3600 MT/s
  • 2× RX 7900 XT, 20 GB each, XFX and PowerColor
  • Gigabyte B550 EAGLE WIFI6
  • XFX on PCIe 4.0 x16; PowerColor on a chipset-connected PCIe 3.0 x1 slot
  • Windows 11, AMD driver 32.0.31041.1004

The cards have reduced clock settings: XFX 1700 MHz core, PowerColor 1800 MHz, both 2500 MHz memory and −10% power limit.

Here are the main single-response results:

| Setup | Generation TPS | Including prompt processing |
|---|---:|---:|
| Qwen3.8-27B IQ4_XS, direct ROCm + MTP3 | 46.4 | 43.1 |
| Qwen3.8-27B IQ2_XXS, direct ROCm + MTP3 | 66.5 | 59.9 |
| Flash Next UD-IQ4_XS, one GPU, warm ROCm run | 8.6 | 5.7 |
| Flash Next UD-IQ4_XS, two GPUs, Vulkan | 5.5–6.5 | 3.3–5.7 |

The 27B tests used roughly 700 input tokens, 8k context and 1024 output tokens. IQ4 had three runs; IQ2 is the median of nine prompts. Settings were ROCm 2.46.0, Flash Attention, f16 KV, MTP3 and batch/microbatch 2048/512, with thinking off.

Two separate IQ4 copies reached 103.5 TPS combined, but that required 16 concurrent requests. I haven't reached 100 TPS for one response. Splitting one model across both cards was slower. Tensor split initially produced broken text; --no-mmap fixed that.

Flash Next is unsloth UD-IQ4_XS, around 93.7 GB. The single-GPU profile used ROCm 2.49.0, --n-cpu-moe 42, f16 KV and 8k context. The dual-GPU profile used Vulkan 2.51.0, tensor split 1:1, --n-cpu-moe 28, q8 KV and 256k context. Both used eight threads, PLE on CPU and MTP off.

The Flash measurements were individual short runs. The configured 256k window was mostly empty, and the different profiles weren't a controlled single-versus-dual comparison.

What would you check first: CPU/RAM offloading, the x1 connection, or backend settings? If you're getting 100+ TPS on 16 GB or less, could you share your exact model/quant, hardware, backend, MTP settings and actual context length? Also whether that's one response or combined throughput.

Update Oct 6: Strata 0.1.39 works on RDNA3 with Windows/HIP. Same Flash UD-IQ4_XS, one 7900 XT, 8k, int8 KV, prefill512, 8 workers, thinking off/greedy. 24 GiB expert RAM + ~7.7 GiB auto GPU expert cache. Three 128-token text runs per setting:

| Setting | Decode TPS | Including prompt |
|---|---:|---:|
| MTP2 | 11.0 | 8.6 |
| MTP4 (tested at start/end) | 10.6–10.8 | 8.4–8.7 |
| MTP8 | 10.0 | 8.2 |
| MTP4, min-p 0.2 | 9.7 | 7.9 |
| MTP2, 32 GiB expert RAM | 13.6 | 10.5 |

MTP8 helped counting but slowed the text prompt. With 512 output tokens and 24 GiB RAM, MTP2 gave 12.7 t/s vs 11.2 for MTP4 (11.7 vs 10.6 including prompt), three runs each. More expert RAM helped most in the short tests. Sequential runs/cache conditions vary; this isn't a quality comparison or a controlled comparison with the old backend. Some small projections are rounded to BF16 by the pack. Still no 100 t/s for one answer. Staying with IQ4; not testing Q2. Linux/custom gfx1100 builds are still untested here.

💬 13 (+8) open on reddit ↗