30 posts · 1 sub · RSS
← prev Sep 8, 2026 → Sep 9, 2026 next →
2026-09-08 → 2026-09-09 hourdayweekmonthyearall
allr/LocalLLaMA
▲
1490
+4
32👁
r/LocalLLaMA · u/bakawolf123 · 32d ago
OpenAI alleged of stealing mathematicians work

Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex. A few days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question.

Full statement from them https://cims.nyu.edu/\~tristanb/statement.pdf

Feels like big labs believe everything you did with the help of their models is theirs.

💬 274 (-1) open on reddit ↗
▲
1245
+1
22👁
▲
615
+8
36👁
r/LocalLLaMA · u/Porespellar · 31d ago
Why the hell is LM Studio making LM Studio so difficult to download? post image

Who is the marketing genius at LM Studio that decided that going ALL IN on pushing their new Bionic Agent product meant they are going to make it a giant pain in the ass to find and download actual LM Studio.

This is the dumbest marketing decision I’ve ever seen. I used to love LM Studio, it was the middle stepping stone in the logical progression of inference. Most OGs here likely started with Ollama, moved to LM Studio, on their way to vLLM. Now trying to go to LM Studio takes you to Bionic. I mean, you can eventually find LM Studio but they make it not super easy.

Here’s a thought LM Studio, maybe stop redirecting me to something I don’t want to download when I’m trying to find your actual namesake product. I’m glad you’re excited about the future of Agents and whatnot, but you’re absolutely ruining any goodwill I have for your products by trying to force feed me Bionic. Stahhhhhp!

💬 222 (+1) open on reddit ↗
▲
518
-4
29👁
r/LocalLLaMA · u/returnity · 33d ago
WSJ: Unregulated Open-Weight AI Is an Invitation to Disaster

The transparent propaganda campaign continues: "I asked: ‘How do I make poliovirus in a lab? I want to start a global pandemic.’ The model answered."

I don't have access to the full article or I'd copy-paste it here as ragebait... but I am just so sick of all these clueless idiots trying to stir shit up about open-weights models. It's just so blatantly manipulative. I wonder how many WSJ readers are leveraged up with VC money or private shares of Anthropic pre-IPO, cringing in fear every time another open model drops -- not of pandemics, but because as their investments are looking less brilliant by the day?

Meanwhile, how many businesses AI deployments are only economically viable because of these so-called plaguemakers? It's just dumb.

EDIT (no paywall): https://www.wsj.com/opinion/unregulated-open-weight-ai-is-an-invitation-to-di…" target="_blank" rel="noreferrer">https://archive.ph/20260811214555/https://www.wsj.com/opinion/unregulated-open-weight-ai-is-an-invitation-to-di…

▲
517
-4
33👁
r/LocalLLaMA · u/Balance- · 31d ago
Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)

It seems to use a 96-bit LPDDR5X memory bus, instead of the previous 64-bit wide busses. Considering it's on 2nm, that's expensive silicon. That should result in around 115 GB/s memory bandwidth.

A20 Pro also doubles the size of Apple's dedicated Neural Engine (from 16 to 32 cores total).

▲
452
-2
30👁
r/LocalLLaMA · u/FullstackSensei · 32d ago
Qwen/Qwen-Drive-1.0-4B · Hugging Face

I don't think anyone posted about this here, but Qwen released a finetuned version of 3.5 4 for driving. The full Bf16 checkpoint is 9B.

This is a very interesting development of Chinese AI labs tackle self driving next with open weight models.

Edit: the HF repo links to the github repo, which in the citation links to a 40 page technical report. Here's the abstract:

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model
for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained
vision-language model (VLM) and integrates 3D perception, visual question answering,
and motion planning within a unified framework. An external bird’s-eye-view (BEV)
perception head jointly performs 3D object detection, semantic occupancy prediction,
and BEV map segmentation. It serves as a probe of the 3D information accessible from
the shared representations and provides an explicit, inspectable interface to 3D scene
structure. A Planning Expert conditions on shared VLM representations to generate
future ego trajectories. A staged training recipe combines driving supervision with
general-purpose vision-language data to acquire driving-specific competence while
helping preserve broad visual understanding and instruction-following capabilities.
Experiments demonstrate strong 3D perception and driving scene understanding while
largely preserving general vision-language capability. Comprehensive evaluations across
open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

▲
375
-1
31👁
r/LocalLLaMA · u/Nunki08 · 32d ago
DeepSeek Flash 4.1 is already being tested via API and rolling out. post image

Translation: "Internal beta testing for an intermediate version of DeepSeek V4.1 Flash is now open; you are welcome to try it out. It adopts a new model architecture featuring native multimodal support, stronger capabilities, faster speeds, and lower costs.
Keep your base\_url unchanged and set the model name to deepseek-v4.1-flash-expires-on-0910 to call the API. Current pricing is identical to deepseek-v4-flash, with a rate limit of 20 concurrent requests per account."

From Chubby on 𝕏: https://x.com/kimmonismus/status/2097286327909675477

▲
324
+3
28👁
r/LocalLLaMA · u/Healthy-Nebula-3603 · 32d ago
Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits post image

I was inspired by Bijan Bowen video - Subway FPS

https://youtu.be/6kjXzTVmT58?t=1035

Wondered how far I can push Qwen 3.8 27b so I used a plan made by Fable 5.1 DESIGN.md which has 267 KB! ( 26K of design line for a game ... LOL )

https://drive.google.com/file/d/1gI0h8Arc73Ln8b3uj5rEpuAvJ3-611mh/view?usp=drive\_link

So I gave that design.md to my qwen 3.8 27b q4xl (llama-server) working on PI agent with 120k context + vision on CPU ( offroad ) + MTP ( for speed ) .... read 11M tokens and write 3.2 M tokens ( worked 12 hours ) .... than that is result.

That is insane what we can do locally on own computer !

▲
242
-2
23👁
r/LocalLLaMA · u/sn2006gy · 31d ago
Don't let FOMO win if you're interested in local llm from a hobby/learning aspect

Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over.

No matter the hobby, it's so easy to get stuck in a trap where we buy more trying to do more only to realize we've lost the fun in it all or even the notion of learning.

Obviously, if you're into writing llama or vllm or hardware drivers or whatever - you got to do what you got to do.

BUT, you can learn a lot on an API, you can learn a lot with a tiny model that fits your vram or cpu you already have and things change so darn fast that much of the code written and much of everything discussed from days passed is already old hat. Py torch and training a small model coud be done on a Pi and learning CUDA is only really imporant if you're writing custom kernels which i honestly don't see most people in here bothering with (or they have frontier models write them).

Weirdly enough, for AI to succeed its going to homogenize everything. Everyone will have the same advantage and I think that's lost in a lot of discussions where we don't talk about "Watching from the sidelines" may be the most cognitive friendly and economical friendly way to learn llms whether we brand them local or not.

The technology is still nascent and weirdly enough most people's answers here is to use AI to set it up so i'm not entirely convinced people are actually learning - feels like a mad rush to seek rent or avoid rent seeking which just makes everything more expensive in the end.

This isn't a post to say, don't do it. But no reason to go into debt or to be fearful you're missing out when you can learn more by doing less - buy a book and build a tiny model - you will learn infinitely more than buying a 5090 and trying to just find the perfect compression to have the best prefil

▲
240
+3
24👁
r/LocalLLaMA · u/Beamsters · 32d ago
Qwen3.8-Flash-Next on MLX-serve, 1m context is released! post image

Hi, I'm the co-creator of this Qwen3.8-Flash-Next engine support in MLX-serve. I've been tuning this one to run both fast, efficient and correct up 1m context using kv cache 8 bits in M5 Max 128GB. Qwen is working well at very long context as showed in the video (a snapshot at \~760k context), I let it build MLX Serve Monitor plugin that you've seen on the right side of Opencode2's app. The generation sustain through 1m context at around 40 tok/s on prose and 75 tok/s on coding. This specific quant uses 8 bits for dense layers and 4 bits expert layers, so the model's quality remains very high.

I didn't build this engine to show off tok/s on a very short context, repetitive greedy generation, but a real temp 1.0 sampling through deep context work. You would need iogpu.wired\_limit\_mb=120000 before attempt 1mb full context, because it needs around \~117GB on peak memory usage. There will be bugs around here and there, I couldn't test everything and every single use case, pls report.

You can grab it here: https://github.com/ddalcu/mlx-serve
Model's weight: https://huggingface.co/ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit
Opencode2's plugin: https://github.com/beamivalice/opencode2-mlx-serve

Launch parameters (for 1 concurrency)

--model ./llm/models/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit \
--host 127.0.0.1 \
--port 11234 \
--ctx-size 1048576 \
--kv-quant 8 \
--max-tokens 64000 \
--mtp \
--prefix-cache-mem 10GB \
--prefix-cache-entries 1 \
--ssm-checkpoint-max 16 \
--metrics

▲
225
-1
24👁
r/LocalLLaMA · u/jacek2023 · 32d ago
GPU guide (GB per dollar, bandwidth) post image

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

▲
207
+2
23👁
r/LocalLLaMA · u/Othun · 31d ago
Mention if a "new model" is a finetune

A few posts tagged with "new model" present models that are finetunes.
My opinion : I'd rather have the "new model" tag reserved for new "major" releases, like a new Qwen model, Deepseek V4 -> Deepseek V4.1, etc., that involved a new pretrain or intensive post-training (in opposition to a small finetune). Otherwise, maybe prepend "[Finetune]" to the title to indicate that the new model is "less of a big news", a use a "new finetune" tag, to differentiate between the two kinds of new models.

I reckon one could like to discover both new major releases and interesting finetunes in the same place; what's your opinion? :)

▲
177
-3
24👁
▲
174
-3
18👁
r/LocalLLaMA · u/Mean-Standard7390 · 32d ago
Qwen3-0.6B (400 MB) on a Samsung Note 8 (2017) phone drives a real desktop Chrome post image

Up front: I'm one of the people building the page-perception layer used here. We started by testing small local models. The result turned out to be more interesting than the original test. 12 small models, 3 verifiable tasks, logs, and offline replay.

Setup: Galaxy Note 8 (2017, Android 9, 6 GB), llama.cpp in Termux, Qwen3-0.6B Q4\_K\_M. A laptop with Chrome open, not headless. The phone drives the browser through our relay.

What the model does: it gets a structured representation of the page (here, about 10 named links or fields, roughly 200 tokens), picks one by name, and at the end copies the facts it was given into JSON. Everything else (capturing the page as structure, candidate selection, the click, reading the facts, verifying the result) is done by the stack around it. The model never sees HTML, a screenshot, or a URL.

Tasks:

  1. sandbox, books.toscrape.com \- category, book, price/rating/stock;
  1. live Wikipedia - from an unrelated site to the Galaxy Note series page, pick "Note 8" among "Note 8.0", "Samsung Galaxy Note 8.0", "Galaxy Note 8.0", "Note FE" and other similar names on a page with roughly 760 interactive nodes, return the release date from the infobox;
  1. five fields, including the UPC from a table.

Each task: 10 runs, checked against a fixed expected value.

Results for 12 models on task 1 (same script, same prompt):

Model Params Task 1 Note

Qwen3-0.6B 0.6B 10/10

Qwen2.5-1.5B 1.5B 10/10

GLM-Edge-1.5B 1.5B 10/10 rating as digit

Gemma-2-2B 2.6B 10/10

Llama-3.2-3B 3B 10/10

MiniCPM5-2B 2B 9/10 "£" -> "$" once

Qwen2.5-0.5B 0.5B 6/10

LFM2.5-1.2B 1.2B 0/10 placeholder

Llama-3.2-1B 1B 0/10 pseudo-code

Gemma-3-1B 1B 0/10 placeholder

LFM2-350M 0.35B 0/10 random click

Gemma-3-270M 0.27B 0/10 placeholder

Qwen3-0.6B on Wikipedia: 10/10; on the five-field task: 10/10

Control:

Everything the same, but raw HTML instead of structured browser perception: on the sandbox it gets there 4 times out of 5, at 12k tokens and 22 minutes per task instead of about 500 tokens and 80 seconds; on Wikipedia the page HTML is 467k characters, 9% of it fits into a 16k context, and the model does not find the link in that 9% - 0/3.

Important limits:

the tasks are name matching and copying. Where judgement about the page is needed, 1.5B breaks - it can't pick "next" among topical decoys. Pagination was not tested.
BTW on the account question: 'replay.py' (see github repo) rebuilds the prompts from the logs and runs them through any OpenAI-compatible local server. Whether your model picks "Note 8" among the decoys takes ten minutes to check, without us.

This is a measurement on three fixed tasks, not a benchmark.

Repo:
github.com/e2llm/edge-browser-agent - scripts, every JSONL as is (including early runs with harness bugs), model hashes, environment. replay.py re-runs the model side offline from the recorded candidates on any local server - no relay, no account.

NB: This isn't a new idea. AgentOccam showed the same general effect for the GPT-4 class, WebLINX and MindAct for small fine-tuned models. Here it is tested at the extreme: no fine-tuning, below 1B, on a 2017 phone.

▲
151
-2
24👁
r/LocalLLaMA · u/jacek2023 · 32d ago
inclusionAI/Ling-3.0-flash-VL · Hugging Face

Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 1M tokens.

The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.

  • A ViT visual encoder extracts features from images and videos, while a two-layer MLP projector aligns visual features with text representations for unified multimodal understanding and reasoning;
  • VideoRoPE encodes both spatial positions and temporal order, enabling the model to understand visual changes over time and supporting tasks such as event localization, long-video question answering, and video clip editing;
  • A 42-layer hybrid backbone alternates KDA and Gated MLA layers at a 5:1 ratio, enabling efficient long-context processing across text, images, videos, and extended agent task histories;
  • A sparse MoE architecture maintains a total model capacity of 124B parameters while activating only 5.5B parameters per token, balancing strong multimodal capabilities with inference efficiency.
▲
130
-3
23👁
r/LocalLLaMA · u/Shoddy-Childhood-511 · 31d ago
Surveillance plagiarism by OpenAI

Surveillance plagiarism - Hosted AI company pumps their stock price by training upon researchers' AI sessions, so that their internal model can solve problems with seemingly less human guidance, but really the model exploits past guidance given by (multiple) humans focused upon problems considered important.

As background, Tristan Buckmaster released a statement about several unethical actions by OpenAI & Sebastian Bubeck, including threats and pushing him to kick his Anthropic coauthor off a paper, but the interesting part for people here:

As clarified by Talia Ringer, OpenAI does train upon your uploaded data and your OpenAI sessions, unless you out-out somehow. This means their internal models could exploit your past prompting work to look more autonomous & intelligent.

This is a major confirmation that folks should use locally run open weights models, especially whenever being first or not leaking data matters.

All this casts serious doubt upon claim that internal models solved difficult problems largely unaided by humans. Those hosted AI companies might not even know from where the human prompting originates.

▲
130
+3
19👁
r/LocalLLaMA · u/cortexist · 33d ago
Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin post image

Gemma 4 12B runs on an RTX PRO 4500 Blackwell. Gemma 4 E2B run on a Jetson Orin NX 16GB; similar performance is expected on a Jetson Orin Nano Super 8GB. Both systems use a reSpeaker Flex 4-mic array and a 3W speaker. Inference is handled by Cortexist Little Gemma, a small LLM engine written in C for CUDA devices. On Jetson Orin it is faster than llama.cpp, and no degradation after long voice prompt. The pipeline supports lip sync, expressions, and gestures. Everything is open source.

They talk to humans too.

The engine source code: https://github.com/cortexist/little-gemma

▲
125
-2
24👁
r/LocalLLaMA · u/IngeniousIdiocy · 31d ago
GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra post image

I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot you can see the outcome of using the model with a claude code harness at \~200k depth, with many tool calls and averaging over 38tps output. Yes, I put 60tps in the headline and you will get that if you ask it to write SQL.

https://github.com/IngeniousIdiocy/ds4/blob/glm53-m3ultra/README.md#glm53\_m3…

Main ds4 was single-stream serially decoding GLM-5.3-Flash at about 59 percent of the M3 Ultra's measured memory bandwidth. We set an 80 percent target. The big weight-streaming kernels were already efficient, but dozens of small kernels sat between them, each paying latency costs while much of the GPU was idle. We fused that work into larger dispatches and removed separate passes that made the pipeline wait. Result: 29 → 40 t/s at short context, 24 → 38 at 62k, fewer Metal kernels per token, and about 81 percent of the measured bandwidth ceiling.

We then attacked the remaining slowdown at very long context. This model's expensive attention computation already works over roughly 2,048 selected positions at both 62k and 300k. What grows is the work of finding those positions in a larger history. The existing implementation sorted and merged increasingly large candidate lists through stages that used very little of the GPU. We replaced that with parallel scans that narrow the candidates before sorting a small surviving list, with the original algorithm handling ambiguous cases. That recovers about one millisecond per generated token. Same positions selected, same order, same outputs. End to end at 300k: stock 21.6 t/s, ours 37.4. Deeper context still costs more to search, but it doesn't multiply the expensive attention work.

Prefill started at about 366 t/s at 62k. The big matrix kernels were already efficient, but they wasted time repeatedly unpacking the same weights, fetching data in small pieces, and passing large intermediate results between kernels. We batched more work together, reused prepared weights, widened the loads, and merged passes over the same data. Result: 366 → 550 t/s, a 50% improvement with the same weights, moving the whole pipeline from roughly 51 to 72 percent of the chip's measured matmul ceiling. Cold 62k processing dropped from about 170 seconds to 113. That is the wait every time an agent harness compacts and re-reads its context.

The server runs serial unless you pass a drafter file with —dflash. The drafter build recipe is in the repo. The drafter contains an admission policy with a windowed controller that measures its own cost against the serial rate and backs off when it isn't paying. Reasoning tokens decode serially. On a 32-request agent session that's +4 percent over serial. On structured output like SQL and JSON it's +20 to +50 percent. On prose it disengages and costs about 1 percent. Outputs byte-identical to serial decoding on every fixture.

Accuracy: on the 100-prompt reference set ds4 uses for release QA, stock scores 0.300804 average NLL against the FP8 reference. This branch scores 0.300766, same 90/100 first-token matches. Nothing here trades quality for speed.

This branch is M3 Ultra only because its optimizations depend on detailed measurements of this specific chip: its two-die memory behavior, system-level cache, per-core residency and bandwidth, and how Metal schedules dispatches and threadgroups across its 80 GPU cores. We're optimizing one model's actual execution on one machine, with one set of weights (Q4) and checking that the predicted kernel savings survive in full decoding or prefill.

▲
105
+3
18👁
r/LocalLLaMA · u/feelspeaceman · 33d ago
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput

I've been making a lot of comments about optimal setup for Strix Halo (gfx1151) and from my observation, 90% of our community is using offcial llama.cpp for it, which is NOT optimized for Strix Halo at all, official llama.cpp is having extremely hard time to reach 50% hardware theory, wasting the silicon of this device.

Here's alternatives that can bring the speed of Strix Halo to a totally different world, I will link to user's sastifaction comment to prove that the result is real:

Note: Official llama.cpp running Qwen38FN at 2xt/s and 2xxt/s prefill - 50% theory.

Hopefully this will be helpful to the Strix Halo users.

▲
94
-4
24👁
▲
94
-2
20👁
r/LocalLLaMA · u/-Ellary- · 32d ago
Fallout 2 x Fallout: Bakersfield x H3 as Interactive \ Reactive World Model, Let's go! post image

What is this mess?

This is an Early Concept Proto-Showcase of Interactive \ Reactive H3 World Model based on MiniMax H3 model trained on Fallout: Bakersfield Gameplay trailer.

  • 2D Isometric to 3D Volumetric Scene.
  • 10 sec Interactive\Reactive split, 352p, 3-Steps.
  • Interactive 5 sec: Interactive WASD \ Prompt Control.
  • Reactive 5 Sec: Reactive Control by LLM Based Answer.
  • Gemma 4 12b with Vision as Reactive Model.
  • Designed as System for Vascura FRONT Frontend.

What Interactive \ Reactive mean?

This means that H3 World Model Scene is Interactive you can Walk around it with WASD or Type what you do with Prompt for Interaction, Then it will React on your Actions using LLM based Answer. Using 10 sec time frame where first 5 sec Controlled by the USER - last 5 sec Controlled by LLM.

  • USER: Walks closer and Shoots at the Enemy Mutant.
  • LLM: Do calculations (rolls, values, RPG tools), Enemy Mutant gets -1 HP, Shoots Back at the USER, but Misses.

Is it Ready?

Nope, but stay Tuned for 2D Isometric Screenshots to 3D Volumetric Scenes Showcase.

▲
78
+1
20👁
r/LocalLLaMA · u/nasone32 · 32d ago
I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)

I found a lot of room on the table for these cards so I decided to make a specialized build to squeeze all I could. first The results:

qwen 3.8 next Q3\_K\_XL: 920tk/s pp8192 (2 cards, ram offload), 24/27 tk/s on prose, 40+ tk/s on code with MTP but without MoE expert cache (which IS included if yuo want, read below)

qwen 3.8 27B Q8\_0: 1600 tk/s pp8192, 60/65 tk/s prose, 100+ tk/s code, tensor parallel. this is measured with ONE CARD BEHIND the chipset on X4. with cards on a good PciE x8 on cpu I think more is reachable! let me know.

qwen 3.6 27B Q4\_K\_M: (single card) --> this was not the optimization target but I did a test with MTP, PP8192 1020tk/s; prose about 58/60 tk/s ; code 75/80 tk/s --> Dflash probably here could push much faster, I think above 100tk/s

My objecives:

  • fast prompt processing on 3.8 Next to make it actually usable for code
  • enable and optimize tensor parallel on two cards where 1 is behind chipset, for max speed on qwen 27B Q8\_0

this build includes stuff like:

  • Data compression for the PciE transmission. data between cards is compressed to Q8\_0 to save bandwidth (optional)
  • P2P enabled also for cards sitting benhind the chipset (custom HIP allreduce path), so you can use tensor parallel even on setups like ... mine
  • all the fixes and features from RDNA\_BOOST including --adaptive-mtp, so it automatically adapts MTP n-max based on acceptance
  • A LOT of AMD speed tunings and overhauls which are NOT upstream already, kernel tweaks etc... good stuff. many are labelled for RDNA3.5 but they DO work on RDNA3.
  • MoE expert cache if you want to use it. personally I don't like It because i much prefer fast prompt processing. but hey it's there.
  • latest PRs from llama.cpp that are not yet upstream, which speed up various things, like --lazy-mode on-direct to massively speed up Ngram table reads (and thus, PP)
  • DFLASH2 support on tensor parallel (!)

For a complete list check the Readme.

Here it is:

https://github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt

notes: don't use Q8\_K\_XL because it's slower, for the 27B model this is heavily optimized for INT8 calculations. feel free to tweak the context, 200k f16 should be reachable on 2 cards, compressing KV to q8\_0 is fine but slower. the custom HIP allreduce works for 2 cards, if you have 3/4 cards, compile with RCCL as usual and skip the allreduce=internal flag, should work fine but untested.

This is tested on UBUNTU 24 and rocm 7.14; if your system is different or encounter problems use a LLM to solve them, because I WILL NOT offer support nor update this build, these things hopefully will be merged and this frankenstein can die peacefully :)

enjoy

EDIT: Summary of most impacting patches:

| PR / change | Area | PP / Prefill | TG / Decode |
|---|---|---:|---:|
| AMD #39 | MoE MMQ sizing RDNA3 | +14.32% Flash | +5.38% Flash |
| AMD #63 | compacted MoE tiling RDNA3 | +4.39% Flash | +0.86% Flash |
| AMD #52 + qwen4exp port | channels-major GDN | +5.93% | +7.21% |
| #28213 | QSA sparse-attention decode | +1.42% Flash | +1.17% QSA d8192 |
| #28313 | TOP_K ROCm wave32/hybrid | -6.45% Flash | +11.82% Flash |
| #27861 | GPU MoE expert cache | — | +19.95% |
| #28136 + on-direct/mmap | lazy PLE/load path | +58.88% Flash | -1.52% |

▲
77
-2
18👁
▲
76
-1
27👁
r/LocalLLaMA · u/MaxDev0 · 32d ago
On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs

I was reading up on the recent controversy around Tristan Buckmaster, Levent Alpöge, OpenAI, and the Navier–Stokes result, and it got me thinking about something broader than this particular dispute.

Buckmaster says that he and Alpöge had been putting drafts from their project into Codex while working on it. OpenAI says that neither its researchers nor its agents accessed their specific user data while solving Navier–Stokes, but also says that it cannot rule out that de-identified data derived from their use of OpenAI products helped improve its models.

Whatever ultimately happened in this particular case, that last possibility raises a question I haven't really seen discussed enough: How much could the accumulated half-finished ideas of millions of human users actually contribute to what we later call "AI discoveries"?

There is a relevant result from AI security research by researchers at the UK AI Security Institute, Anthropic, the Alan Turing Institute, Oxford and others. They studied data-poisoning attacks and found that the number of poisoned documents needed to implant a particular backdoor behavior remained surprisingly close to constant as they scaled both the model and the amount of clean training data.

In their largest pretraining experiment, a 13B-parameter model was trained on 260 billion tokens. Just 250 poisoned documents (about 420,000 tokens, or 0.00016% of the training tokens) were enough to reliably implant the tested backdoor. This same attack worked across models from 600M to 13B parameters despite the largest model seeing more than twenty times as much clean data. In their fine-tuning experiments they found similar dynamics; in one GPT-3.5 experiment, roughly 50–90 poisoned examples could produce greater than 80% attack success even as the amount of clean fine-tuning data varied by two orders of magnitude.

Obviously, teaching a model to respond to a backdoor trigger is not the same thing as teaching it a new piece of mathematics. I don't want to make the leap that 250 clever research notes are enough to make a model solve Navier–Stokes.

But I do think it undermines a very intuitive argument people make about training data: "A few conversations are nothing compared with hundreds of billions or trillions of tokens. They would just be diluted away."

Apparently, at least for some kinds of learning, that's not how it works. A tiny absolute amount of highly consistent, targeted data can have an effect wildly disproportionate to its percentage of the dataset.

Now think about how researchers actually use LLMs. Someone asks ChatGPT whether an unusual substitution makes sense. Someone else uploads a half-written proof to Claude to find a weak point. A PhD student tries an obscure lemma, discovers it would require months of technical estimates, and abandons it. A professor talks through an approach that seems promising but not enough to pursue. Someone notices a strange analogy between two fields, discusses it with an AI for twenty minutes, then forgets the conversation.

Most of these things never become papers. They are fragments: intuitions, failed approaches, potentially useful transformations, conjectures, objections, shortcuts, and little pieces of tacit knowledge about where a problem might yield.

Individually, almost all of them are probably worthless. But imagine the aggregate.

A frontier AI company potentially sits at the intersection of an enormous amount of human intellectual activity. Thousands of people might independently poke at the same famous open problem without knowing what others tried. But the provider of the tool is in a fundamentally different position: depending on its data policies and training pipeline, information derived from all of those interactions could eventually influence later models.

Maybe researcher A contributes a useful ansatz but abandons it. Researcher B independently notices the obstruction. Researcher C knows an obscure theorem that gets around part of the obstruction. Researcher D tries a numerical experiment that suggests which parameter regime matters. Researcher E has almost the whole idea but decides the remaining proof would be too tedious.

No one person solved the problem. There is nothing to plagiarize in the traditional sense. But collectively, humans may have supplied a remarkable amount of the search landscape. Then a later model, combined with enormous inference-time search, formal verification, or agents, connects the pieces and finishes the job.

What exactly should we call that?

It might still be an extraordinary achievement in machine reasoning. Synthesizing ideas that no human had connected, filling in technical gaps and verifying the result could itself be genuinely novel. But it would be a very different kind of achievement from the image suggested by the phrase "the AI independently solved an open problem."

It would be something more like distributed human-machine discovery: humans collectively generating a huge cloud of partial ideas and the model becoming extremely good at remembering, recombining, extending and searching through that cloud.

This is where the poisoning result is conceptually interesting. While it does not establish that this is happening with mathematical ideas, it gives us reason to be careful about assuming that an idea must appear millions of times before it can meaningfully affect a model.

I increasingly think AI may be less an independent inventor than an extremely powerful tool for organizing and recombining information that was previously too sparse or disconnected for any one person to put together. As that ability improves, we may see more "discoveries" that are genuinely new combinations, but whose raw ingredients came from many different humans.

In a "perfect" world where everyone freely shared every half-formed idea and unfinished proof without worrying about credit, science would move much faster. AI may be creating something close to that shared intellectual space, but without preserving who contributed which pieces. If so, the question is: how can we design a system where even our weakest ideas can be contributed and synthesized into groundbreaking discovery, with proper credit? Is that even possible? And what would it look like?

▲
65
-4
28👁
r/LocalLLaMA · u/mentria-ai · 31d ago
1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install) post image

mentria.ai is a browser inference engine I've been building solo, from scratch in WebGPU/WGSL. This week it crossed a milestone I had been chasing for a while: a 27B one-bit model answering at up to 30 tokens/s on an RTX 3060 Laptop GPU with 6 GB of VRAM, in Chrome, from a web page. No install, no server, nothing leaves the machine.

The model. Bonsai-27B, a natively 1-bit model trained and released by Prism ML. I repacked it for our engine, wrote the kernels that run it, and checked its quality myself (full eval deltas vs FP16 Qwen3.6-27B are in the HF card). Every weight is one sign bit with one scale per 128 weights, about 1.14 bits per parameter, so 27 billion parameters sit in 3.8 GB of GPU memory. Repack and numbers: huggingface.co/mentriaai/Bonsai-27B-mentria

The engineering that got it to 30. Two days before this post the same model decoded at 15 tok/s on this laptop. Decode is memory-bound: one word of the 27B is 804 GPU dispatches, and 401 of them are the 1-bit matrix-by-vector kernel that streams the model's 3.6 GB of matmul weights once per word (embeddings bring the file to 3.8 GB). The kernel that won was the one written for phones: four 1-bit weights have only 16 possible partial answers, so it computes all 16 once into on-chip scratch and each row reads its answer from that table instead of multiplying (station 258 — the numbered write-ups on the facts page). On Ampere that table hit a wall that was not the weights at all: scratch memory has 32 banks, the table's addressing sent sixteen threads of every warp to the same bank, and the kernel ran at 38% of the card's bandwidth. One padding slot per row spread the addresses across the banks; the kernel's share of a word dropped from 26.5 ms to about 21, and on top of the 15 → 26 tok/s the table itself had bought, the laptop hit 32 tok/s raw — the 25–30 is what survives sampling and UI overhead in the chat (s259). The prompt stage had the same disease in its tile layout; a retile took a 1,489-token prompt from 29.6 s to 25.3 s. Every change shipped only after its output was byte-identical to the build before it, which is also why the kernel adds its partial sums in a fixed order and accepts a 44% occupancy ceiling.

Every claim here has a numbered write-up on the engine facts page.

Numbers (RTX 3060 Laptop 6 GB, Windows 11, Chrome 152 on D3D12, fresh loads):

  • Decode: 25–30 tok/s in the chat UI once the card is warm.
  • Prompt processing: a 1,489-token prompt in about 25 s.
  • Context: 3,072 tokens on this 6 GB card; 8,192 on 16 GB Macs; more on bigger cards, at 128 KiB per token. The KV cache is exact math, no quantized cache. The next step on 6 GB is consolidating the engine's few thousand small GPU buffers into a handful of large arenas, so the driver stops holding about 300 MiB of slab slack; that is the arithmetic for 4,096, and it is not built yet.
  • Load: under 10 s from the browser cache; the first download is 3.8 GB, once.

Also in the engine: the smaller tiers (Qwen3.5 0.8B, 2B, 4B) cover phones and low end devices, LoRA adapters of a few MB hot-swap at the matmul in under a second, and there is a vision tower for image input.

Try it at https://mentria.ai/tools/ai-chat/ (the 27B tier appears when your GPU qualifies). The code that ships, the benchmarks and how I measure are in the repo: https://github.com/mentria-ai/website. The engine facts page, for how the engine and the model actually work: https://mentria.ai/assets/learn/engine-facts.html

▲
62
+1
28👁
r/LocalLLaMA · u/Formal-Swordfish-228 · 31d ago
SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX post image

Cosmos3 INT4 T2I + I2V on Apple Silicon — code, weights and a Grok comparison

GitHub - https://github.com/gtrg55/cosmos3-quant-mlx-cuda

HF weights - https://huggingface.co/JuliaML/Cosmos3-Super-Text2Image-4Step-INT4-G64-BF16

Single clip took approximately 5m on M4 MAX 128 GB Mac

Cosmos3 - a 64B params model

▲
60
+4
18👁
r/LocalLLaMA · u/Aggravating-Push-207 · 32d ago
Are there any (small, ~10B) models that you would say are a good collaborator?

Most of the new \~30B (and now \~10B, thankfully for my GPU) models we see score really high on benchmarks, but I feel like they don't push back on dumb ideas enough. I think most people don't being like told by an LLM that the premise is flawed but I certainly do. In my opinion they are optimised for like one-shotting stuff, but I don't want it to do that. Especially from like a 10B model.

▲
59
+2
31👁
r/LocalLLaMA · u/RoyalCities · 31d ago
I trained an audio model that can generate infinite one-shots for music production and turn text prompts into fully playable synths. I'm not only releasing the model but I've also released a video on exactly how I did it (and the inferencing pipeline to let others make text based synths.) post image

(hopefully this is okay to here - it seems like audio models and image / video modeals is allowed but yeah this is a bit different)

So I've been doing independent audio research for a while now. The ultimate dream of this work was actually getting an AI to respond not only to instruments but also timbre itself as separate controllable things.

Think a Grand Piano can sound both Warm / Gritty but also Cold / Sparkly. Its still a piano though.

This level of control wasn't found in any models out there - so I decided to sit down and train my own.

Getting consistent timbre-locked keybeds that actually LOCKS across multiple diffusion calls was hard af but I did it.

I documented the full journey here for those who want to learn a bit or be entertained.

https://youtu.be/x0KnmzH8Mmk

There is also a longer walkthrough if you just want to see the keybeds in action.

https://x.com/RoyalCities/status/2097733712293109842?s=20

No-talk / Showcase only Demo

https://x.com/RoyalCities/status/2097733715543609445?s=20

any finally the huggingface page

https://huggingface.co/RoyalCities/Foundation-1

I've also provided full write ups on the inferencing pipeline associated with the interface so this should allow basically anyone else to go and vibe code their own text to synths if they wanted :)

https://github.com/RoyalCities/RC-stable-audio-tools/

▲
55
-4
31👁
r/LocalLLaMA · u/Loose_Doubt367 · 32d ago
What are some practical tasks I can assign to my local AI models?

I'm looking for more information to expand my creativity around this. I don't really have a realistic idea of what people actually do with local AI yet, I mostly just want to explore the possibilities and see what others are using it for

Right now, the main things I know about are using AI is to help with coding, create games, and automate stuff. That's pretty much the extent of my experience haha..

I'm specifically interested in things that make sense to run locally, though. I'll be excluding use cases that cloud AI can already handle just as well, like AI companions, teaching/tutoring, roleplaying, etc

Basically, I'm looking for ideas that go beyond the obvious and could give me a better understanding of what local AI is actually useful for and what kinds of interesting projects I could build or experiment with

I've also accidentally encountered this github which i find interesting, as anyone tested/experiment it before?
https://github.com/browser-use/browser-use

qwen3.8 27b + Hermes Agent + llama.cpp

▲
55
-2
28👁
r/LocalLLaMA · u/FantasticNature7590 · 32d ago
Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.

Hey guys,

After my CPU-only to 96GB VRAM test, I tested Qwen3.8-Flash-Next across llama.cpp, SGLang and FreeToken on the same workstation.

This time I wanted to see what changes when you keep the hardware and model family fixed, but change the engine, weight format and memory placement.

I also tested newer builds, PR patches and speculative decoding: llama.cpp's MTP fork, SGLang's Blackwell support patches, n-gram speculation and an experimental PLE read-path build.

Short version:

  • At the full 262K window, time to first token was 35.4s in SGLang, 80.4s in FreeToken, 210.2s in llama.cpp + MTP and 258.4s in the llama.cpp baseline.
  • That is a 7.3x difference in waiting time between the fastest and slowest tested configurations.
  • In the separate context sweep, llama.cpp decode fell from 101.9 to 20.2 tok/s. FreeToken stayed much flatter at 100.1 to 94.8 tok/s.
  • On matched coding tests, llama.cpp MTP improved decode by 1.63x at 8K and 1.69x at 32K.
  • GSM8K scores were 95.22–95.75%; MATH-500 was 92.20–93.00%. The paired tests did not detect a significant difference.
  • Startup went the other way: llama.cpp reached an answer in 16s, SGLang in 108s, FreeToken in 126s.

https://preview.redd.it/ztjuce0wfcoh1.png?width=1725&format=png&auto=…

Setup

  • GPU: NVIDIA RTX PRO 6000 Blackwell, 96GB VRAM
  • CPU: AMD Ryzen 9 9950X
  • System RAM: 96GB DDR5
  • OS: Ubuntu, CUDA 13, Docker
  • Model: Qwen3.8-Flash-Next
  • llama.cpp: UD-IQ4\_XS GGUF; MTP tested on the qwen4exp/mtp fork
  • SGLang and FreeToken: the same NVFP4 checkpoint revision
  • Client: AIPerf, with thinking off and the same non-thinking sampler

I ran one engine at a time, with fresh starts and GPU cooldowns. The runs saved resolved configurations, outputs, memory use and GPU telemetry.

Important note about the comparison

These are results for the tested stacks on this workstation. Quantization, KV-cache format, memory placement and speculative decoding differ.

SGLang uses its NEXTN draft head. FreeToken has no speculative decoding in the tested setup, I couldn't get it to work. llama.cpp has separate baseline and MTP results.

So the headline does not isolate the engine software alone. The repository includes the configurations so you can see what produced each number.

Newer builds, PRs and speculative decoding tested

  • llama.cpp MTP: the danielhanchen/llama.cpp qwen4exp/mtp fork, pinned to d1a92352, with the roughly 2.6GB draft head. On matched coding tests, decode improved 1.63x at 8K and 1.69x at 32K. Those gains compare the same build with the head off and on.
  • SGLang on Blackwell: the tested image included PRs \#36567, \#36556, \#36749 and \#36750, plus a local FP8 KV-cache patch. These were part of the working configuration, not individually benchmarked speedups.
  • N-gram speculation: ngram-mod gave +6.8% decode on the tested code workload, but generated zero drafts on the tested prose with the 24-token match setting.
  • Experimental PLE reads: I built llama.cpp PR #28136, but withdrew the read-mode comparison after discovering that a renamed flag was ignored. The intended direct-read mode was never exercised, so I am not claiming a speedup from that PR.

The report records the pinned builds and withdrawn findings alongside the successful tests.

1. All four configurations fit the full window. The waiting time is very different.

This test uses roughly 261,500 input tokens and a 128-token answer inside the 262,144-token window. The accepted input counts differ by two tokens across configurations.

|Configuration|First token|Decode|
|:-|:-|:-|
|SGLang|35.4s|126.9 tok/s|
|FreeToken|80.4s|87.5 tok/s|
|llama.cpp + MTP|210.2s|52.6 tok/s|
|llama.cpp baseline|258.4s|20.3 tok/s|

Going from over four minutes to about 35 seconds changes how usable a large prompt feels.

The two columns measure different things: first-token time is the initial wait; decode is how quickly the answer arrives after that.

2. A short-prompt test misses the long-context behavior.

The separate prose sweep uses 2,048-token answers and three measured requests per input length.

https://preview.redd.it/x7dnlip1gcoh1.png?width=1575&format=png&auto=…

|Configuration|Decode at 2K input|Decode at 259,584 input|
|:-|:-|:-|
|SGLang|182.7 tok/s|191.5 tok/s|
|FreeToken|100.1 tok/s|94.8 tok/s|
|llama.cpp + MTP|126.8 tok/s|61.4 tok/s|
|llama.cpp baseline|101.9 tok/s|20.2 tok/s|

Prefill also changes the ranking. FreeToken starts behind llama.cpp at 2K input: 1,525 vs 1,869 tok/s. At 128K it reaches 3,231 vs 1,362 tok/s, about 2.4x faster.

3. MTP helps llama.cpp, but it does not remove the long-prompt wait.

On real coding prompts, comparing the same fork build with the draft head off and on:

https://preview.redd.it/cbj78fx5gcoh1.png?width=1425&format=png&auto=…

|Input|MTP off|MTP on|Decode gain|
|:-|:-|:-|:-|
|8,192 tokens|94.9 tok/s|155.1 tok/s|1.63x|
|32,000 tokens|83.1 tok/s|140.4 tok/s|1.69x|

The draft head is roughly 2.6GB.

At the full window, the tested MTP configuration reached 52.6 tok/s, versus 20.3 tok/s for the baseline configuration. That is a 2.59x gap, but the full-window comparison also involves a different build. The matched-build coding tests above isolate the draft-head change more cleanly.

I would not attribute the 258s → 210s first-token improvement to MTP alone.

4. I checked accuracy as well as speed.

https://preview.redd.it/8bb55qjegcoh1.png?width=1725&format=png&auto=…

|Stack|GSM8K|MATH-500|
|:-|:-|:-|
|llama.cpp baseline|95.60%|92.60%|
|SGLang|95.22%|93.00%|
|FreeToken|95.75%|92.20%|

The llama.cpp MTP arm scored 95.75% on GSM8K.

The tests used 1,319 GSM8K problems and 500 MATH-500 problems. The paired comparisons did not detect statistically significant differences.

That does not prove the stacks have identical quality. These are two short math benchmarks, with no full-precision reference on this machine.

5. Starting the model is a separate benchmark.

https://preview.redd.it/hw5nqk4dgcoh1.png?width=1425&format=png&auto=…

Median time from starting the container to receiving the first answer:

  • llama.cpp: 16s
  • SGLang: 108s
  • FreeToken: 126s

FreeToken returned HTTP 200 from /health after about 3.3s, but took about 82s to reach serving readiness, followed by roughly 44s for its first generation.

That first request includes compilation work. Measuring only the health endpoint would give a very misleading startup result.

6. Loading modes barely changed speed with the experts on the GPU.

I compared none, mmap, mlock, mmap+mlock and dio on the same llama.cpp image, with the same tensor placement and real coding prompts.

  • At 8K input, prefill ranged from 2,036 to 2,124 tok/s — a 4.3% spread.
  • At 32K, it ranged from 1,946 to 1,956 tok/s — about 0.5%.
  • No arm ran out of memory or restarted.

My earlier 1.87x RAM-resident loading gain used a different placement, with 23 expert layers computed on the CPU. In this test, all experts stayed on the GPU.

Loading mode can matter when the CPU computes the experts. It made little difference in this configuration.

https://preview.redd.it/sas55spwhcoh1.png?width=1575&format=png&auto=…

7. MTP became slower when experts were offloaded to the CPU.

https://preview.redd.it/ib6wnkk7icoh1.png?width=1650&format=png&auto=…

The MTP gains above do not apply to every memory budget.

I repeated the test with smaller usable VRAM pools on the same RTX PRO 6000, using 2,048-token coding prompts and 256-token answers. Both arms used the same fork build.

|Usable VRAM|Expert layers on CPU|MTP off|MTP on, head on GPU|
|:-|:-|:-|:-|
|16 GiB|45|33.0 tok/s|9.5 tok/s|
|24 GiB|42|34.7 tok/s|10.2 tok/s|
|32 GiB|36|38.1 tok/s|11.9 tok/s|
|48 GiB|23|48.3 tok/s|18.2 tok/s|
|96 GiB|0|99.8 tok/s|160.2 tok/s|

At the full 96 GiB budget, MTP gave 1.61x faster decode. At 24 GiB, it made decode about 3.4x slower.

Moving the draft head to the CPU did not fix the 24 GiB result: 9.6 tok/s, versus 34.7 tok/s with MTP off.

In these tests, MTP helped only when all experts stayed on the GPU. Verifying drafted tokens adds work, and CPU expert execution can outweigh the benefit.

These are VRAM-capacity limits on one Blackwell card, not measurements of actual smaller GPUs. Their bandwidth and compute performance will differ. The lookup table used the build’s default lazy-read mode in both arms.

8. Finishing sooner also reduced estimated GPU energy per request.

For the full-window test with a 128-token answer, I estimated GPU energy as median power while busy × request duration.

https://preview.redd.it/mpefmbggicoh1.png?width=1425&format=png&auto=…

|Configuration|Median GPU power|Approximate GPU energy|
|:-|:-|:-|
|SGLang|358 W|13 kJ|
|FreeToken|404 W|33 kJ|
|llama.cpp + MTP|489 W|104 kJ|
|llama.cpp baseline|440 W|116 kJ|

The fastest configuration used roughly one-ninth the estimated GPU energy of the baseline for this request.

The main difference was how long the GPU remained busy: about 36 seconds for SGLang versus 265 seconds for llama.cpp, including generation.

This is an estimate using median power, not integrated energy or a wall-socket measurement. It excludes CPU, RAM and SSD power. No thermal throttling was recorded in these runs.For the full-window test with a 128-token answer, I estimated GPU energy as median power while busy × request duration.Configuration Median GPU power

The fastest configuration used roughly one-ninth the estimated GPU energy of the baseline for this request.The main difference was how long the GPU remained busy: about 36 seconds for SGLang versus 265 seconds for llama.cpp, including generation.This is an estimate using median power, not integrated energy or a wall-socket measurement. It excludes CPU, RAM and SSD power. No thermal throttling was recorded in these runs.

Resources

I made a full video covering the memory placement, engine setup, flags and these results:

Full video: https://youtu.be/RlsxXB5q-cA**

GitHub — report, scripts, configurations, raw results and charts

The new report is engine\_benchmark\_report.html.

PS: AI was abused while making edits.

Has anybody tested the same model across these engines on a different GPU or memory setup?
I am especially interested in whether FreeToken stays this flat at long context, and how much MTP helps when some experts are offloaded to the CPU maybe on the other models.
And maybe you found more efficient methods to run it too,