13 posts · 1 sub · RSS
← prev Wednesday, September 9, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
615
+8
36👁
r/LocalLLaMA · u/Porespellar · 31d ago
Why the hell is LM Studio making LM Studio so difficult to download? post image

Who is the marketing genius at LM Studio that decided that going ALL IN on pushing their new Bionic Agent product meant they are going to make it a giant pain in the ass to find and download actual LM Studio.

This is the dumbest marketing decision I’ve ever seen. I used to love LM Studio, it was the middle stepping stone in the logical progression of inference. Most OGs here likely started with Ollama, moved to LM Studio, on their way to vLLM. Now trying to go to LM Studio takes you to Bionic. I mean, you can eventually find LM Studio but they make it not super easy.

Here’s a thought LM Studio, maybe stop redirecting me to something I don’t want to download when I’m trying to find your actual namesake product. I’m glad you’re excited about the future of Agents and whatnot, but you’re absolutely ruining any goodwill I have for your products by trying to force feed me Bionic. Stahhhhhp!

💬 222 (+1) open on reddit ↗
▲
1245
+1
22👁
▲
517
-4
33👁
r/LocalLLaMA · u/Balance- · 31d ago
Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)

It seems to use a 96-bit LPDDR5X memory bus, instead of the previous 64-bit wide busses. Considering it's on 2nm, that's expensive silicon. That should result in around 115 GB/s memory bandwidth.

A20 Pro also doubles the size of Apple's dedicated Neural Engine (from 16 to 32 cores total).

▲
242
-2
23👁
r/LocalLLaMA · u/sn2006gy · 31d ago
Don't let FOMO win if you're interested in local llm from a hobby/learning aspect

Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over.

No matter the hobby, it's so easy to get stuck in a trap where we buy more trying to do more only to realize we've lost the fun in it all or even the notion of learning.

Obviously, if you're into writing llama or vllm or hardware drivers or whatever - you got to do what you got to do.

BUT, you can learn a lot on an API, you can learn a lot with a tiny model that fits your vram or cpu you already have and things change so darn fast that much of the code written and much of everything discussed from days passed is already old hat. Py torch and training a small model coud be done on a Pi and learning CUDA is only really imporant if you're writing custom kernels which i honestly don't see most people in here bothering with (or they have frontier models write them).

Weirdly enough, for AI to succeed its going to homogenize everything. Everyone will have the same advantage and I think that's lost in a lot of discussions where we don't talk about "Watching from the sidelines" may be the most cognitive friendly and economical friendly way to learn llms whether we brand them local or not.

The technology is still nascent and weirdly enough most people's answers here is to use AI to set it up so i'm not entirely convinced people are actually learning - feels like a mad rush to seek rent or avoid rent seeking which just makes everything more expensive in the end.

This isn't a post to say, don't do it. But no reason to go into debt or to be fearful you're missing out when you can learn more by doing less - buy a book and build a tiny model - you will learn infinitely more than buying a 5090 and trying to just find the perfect compression to have the best prefil

▲
240
+3
24👁
r/LocalLLaMA · u/Beamsters · 32d ago
Qwen3.8-Flash-Next on MLX-serve, 1m context is released! post image

Hi, I'm the co-creator of this Qwen3.8-Flash-Next engine support in MLX-serve. I've been tuning this one to run both fast, efficient and correct up 1m context using kv cache 8 bits in M5 Max 128GB. Qwen is working well at very long context as showed in the video (a snapshot at \~760k context), I let it build MLX Serve Monitor plugin that you've seen on the right side of Opencode2's app. The generation sustain through 1m context at around 40 tok/s on prose and 75 tok/s on coding. This specific quant uses 8 bits for dense layers and 4 bits expert layers, so the model's quality remains very high.

I didn't build this engine to show off tok/s on a very short context, repetitive greedy generation, but a real temp 1.0 sampling through deep context work. You would need iogpu.wired\_limit\_mb=120000 before attempt 1mb full context, because it needs around \~117GB on peak memory usage. There will be bugs around here and there, I couldn't test everything and every single use case, pls report.

You can grab it here: https://github.com/ddalcu/mlx-serve
Model's weight: https://huggingface.co/ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit
Opencode2's plugin: https://github.com/beamivalice/opencode2-mlx-serve

Launch parameters (for 1 concurrency)

--model ./llm/models/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit \
--host 127.0.0.1 \
--port 11234 \
--ctx-size 1048576 \
--kv-quant 8 \
--max-tokens 64000 \
--mtp \
--prefix-cache-mem 10GB \
--prefix-cache-entries 1 \
--ssm-checkpoint-max 16 \
--metrics

▲
207
+2
23👁
r/LocalLLaMA · u/Othun · 31d ago
Mention if a "new model" is a finetune

A few posts tagged with "new model" present models that are finetunes.
My opinion : I'd rather have the "new model" tag reserved for new "major" releases, like a new Qwen model, Deepseek V4 -> Deepseek V4.1, etc., that involved a new pretrain or intensive post-training (in opposition to a small finetune). Otherwise, maybe prepend "[Finetune]" to the title to indicate that the new model is "less of a big news", a use a "new finetune" tag, to differentiate between the two kinds of new models.

I reckon one could like to discover both new major releases and interesting finetunes in the same place; what's your opinion? :)

▲
177
-3
24👁
▲
130
-3
23👁
r/LocalLLaMA · u/Shoddy-Childhood-511 · 31d ago
Surveillance plagiarism by OpenAI

Surveillance plagiarism - Hosted AI company pumps their stock price by training upon researchers' AI sessions, so that their internal model can solve problems with seemingly less human guidance, but really the model exploits past guidance given by (multiple) humans focused upon problems considered important.

As background, Tristan Buckmaster released a statement about several unethical actions by OpenAI & Sebastian Bubeck, including threats and pushing him to kick his Anthropic coauthor off a paper, but the interesting part for people here:

As clarified by Talia Ringer, OpenAI does train upon your uploaded data and your OpenAI sessions, unless you out-out somehow. This means their internal models could exploit your past prompting work to look more autonomous & intelligent.

This is a major confirmation that folks should use locally run open weights models, especially whenever being first or not leaking data matters.

All this casts serious doubt upon claim that internal models solved difficult problems largely unaided by humans. Those hosted AI companies might not even know from where the human prompting originates.

▲
125
-2
24👁
r/LocalLLaMA · u/IngeniousIdiocy · 31d ago
GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra post image

I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot you can see the outcome of using the model with a claude code harness at \~200k depth, with many tool calls and averaging over 38tps output. Yes, I put 60tps in the headline and you will get that if you ask it to write SQL.

https://github.com/IngeniousIdiocy/ds4/blob/glm53-m3ultra/README.md#glm53\_m3…

Main ds4 was single-stream serially decoding GLM-5.3-Flash at about 59 percent of the M3 Ultra's measured memory bandwidth. We set an 80 percent target. The big weight-streaming kernels were already efficient, but dozens of small kernels sat between them, each paying latency costs while much of the GPU was idle. We fused that work into larger dispatches and removed separate passes that made the pipeline wait. Result: 29 → 40 t/s at short context, 24 → 38 at 62k, fewer Metal kernels per token, and about 81 percent of the measured bandwidth ceiling.

We then attacked the remaining slowdown at very long context. This model's expensive attention computation already works over roughly 2,048 selected positions at both 62k and 300k. What grows is the work of finding those positions in a larger history. The existing implementation sorted and merged increasingly large candidate lists through stages that used very little of the GPU. We replaced that with parallel scans that narrow the candidates before sorting a small surviving list, with the original algorithm handling ambiguous cases. That recovers about one millisecond per generated token. Same positions selected, same order, same outputs. End to end at 300k: stock 21.6 t/s, ours 37.4. Deeper context still costs more to search, but it doesn't multiply the expensive attention work.

Prefill started at about 366 t/s at 62k. The big matrix kernels were already efficient, but they wasted time repeatedly unpacking the same weights, fetching data in small pieces, and passing large intermediate results between kernels. We batched more work together, reused prepared weights, widened the loads, and merged passes over the same data. Result: 366 → 550 t/s, a 50% improvement with the same weights, moving the whole pipeline from roughly 51 to 72 percent of the chip's measured matmul ceiling. Cold 62k processing dropped from about 170 seconds to 113. That is the wait every time an agent harness compacts and re-reads its context.

The server runs serial unless you pass a drafter file with —dflash. The drafter build recipe is in the repo. The drafter contains an admission policy with a windowed controller that measures its own cost against the serial rate and backs off when it isn't paying. Reasoning tokens decode serially. On a 32-request agent session that's +4 percent over serial. On structured output like SQL and JSON it's +20 to +50 percent. On prose it disengages and costs about 1 percent. Outputs byte-identical to serial decoding on every fixture.

Accuracy: on the 100-prompt reference set ds4 uses for release QA, stock scores 0.300804 average NLL against the FP8 reference. This branch scores 0.300766, same 90/100 first-token matches. Nothing here trades quality for speed.

This branch is M3 Ultra only because its optimizations depend on detailed measurements of this specific chip: its two-die memory behavior, system-level cache, per-core residency and bandwidth, and how Metal schedules dispatches and threadgroups across its 80 GPU cores. We're optimizing one model's actual execution on one machine, with one set of weights (Q4) and checking that the predicted kernel savings survive in full decoding or prefill.

▲
94
-4
24👁
▲
65
-4
28👁
r/LocalLLaMA · u/mentria-ai · 31d ago
1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install) post image

mentria.ai is a browser inference engine I've been building solo, from scratch in WebGPU/WGSL. This week it crossed a milestone I had been chasing for a while: a 27B one-bit model answering at up to 30 tokens/s on an RTX 3060 Laptop GPU with 6 GB of VRAM, in Chrome, from a web page. No install, no server, nothing leaves the machine.

The model. Bonsai-27B, a natively 1-bit model trained and released by Prism ML. I repacked it for our engine, wrote the kernels that run it, and checked its quality myself (full eval deltas vs FP16 Qwen3.6-27B are in the HF card). Every weight is one sign bit with one scale per 128 weights, about 1.14 bits per parameter, so 27 billion parameters sit in 3.8 GB of GPU memory. Repack and numbers: huggingface.co/mentriaai/Bonsai-27B-mentria

The engineering that got it to 30. Two days before this post the same model decoded at 15 tok/s on this laptop. Decode is memory-bound: one word of the 27B is 804 GPU dispatches, and 401 of them are the 1-bit matrix-by-vector kernel that streams the model's 3.6 GB of matmul weights once per word (embeddings bring the file to 3.8 GB). The kernel that won was the one written for phones: four 1-bit weights have only 16 possible partial answers, so it computes all 16 once into on-chip scratch and each row reads its answer from that table instead of multiplying (station 258 — the numbered write-ups on the facts page). On Ampere that table hit a wall that was not the weights at all: scratch memory has 32 banks, the table's addressing sent sixteen threads of every warp to the same bank, and the kernel ran at 38% of the card's bandwidth. One padding slot per row spread the addresses across the banks; the kernel's share of a word dropped from 26.5 ms to about 21, and on top of the 15 → 26 tok/s the table itself had bought, the laptop hit 32 tok/s raw — the 25–30 is what survives sampling and UI overhead in the chat (s259). The prompt stage had the same disease in its tile layout; a retile took a 1,489-token prompt from 29.6 s to 25.3 s. Every change shipped only after its output was byte-identical to the build before it, which is also why the kernel adds its partial sums in a fixed order and accepts a 44% occupancy ceiling.

Every claim here has a numbered write-up on the engine facts page.

Numbers (RTX 3060 Laptop 6 GB, Windows 11, Chrome 152 on D3D12, fresh loads):

  • Decode: 25–30 tok/s in the chat UI once the card is warm.
  • Prompt processing: a 1,489-token prompt in about 25 s.
  • Context: 3,072 tokens on this 6 GB card; 8,192 on 16 GB Macs; more on bigger cards, at 128 KiB per token. The KV cache is exact math, no quantized cache. The next step on 6 GB is consolidating the engine's few thousand small GPU buffers into a handful of large arenas, so the driver stops holding about 300 MiB of slab slack; that is the arithmetic for 4,096, and it is not built yet.
  • Load: under 10 s from the browser cache; the first download is 3.8 GB, once.

Also in the engine: the smaller tiers (Qwen3.5 0.8B, 2B, 4B) cover phones and low end devices, LoRA adapters of a few MB hot-swap at the matmul in under a second, and there is a vision tower for image input.

Try it at https://mentria.ai/tools/ai-chat/ (the 27B tier appears when your GPU qualifies). The code that ships, the benchmarks and how I measure are in the repo: https://github.com/mentria-ai/website. The engine facts page, for how the engine and the model actually work: https://mentria.ai/assets/learn/engine-facts.html

▲
62
+1
28👁
r/LocalLLaMA · u/Formal-Swordfish-228 · 31d ago
SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX post image

Cosmos3 INT4 T2I + I2V on Apple Silicon — code, weights and a Grok comparison

GitHub - https://github.com/gtrg55/cosmos3-quant-mlx-cuda

HF weights - https://huggingface.co/JuliaML/Cosmos3-Super-Text2Image-4Step-INT4-G64-BF16

Single clip took approximately 5m on M4 MAX 128 GB Mac

Cosmos3 - a 64B params model

▲
59
+2
31👁
r/LocalLLaMA · u/RoyalCities · 31d ago
I trained an audio model that can generate infinite one-shots for music production and turn text prompts into fully playable synths. I'm not only releasing the model but I've also released a video on exactly how I did it (and the inferencing pipeline to let others make text based synths.) post image

(hopefully this is okay to here - it seems like audio models and image / video modeals is allowed but yeah this is a bit different)

So I've been doing independent audio research for a while now. The ultimate dream of this work was actually getting an AI to respond not only to instruments but also timbre itself as separate controllable things.

Think a Grand Piano can sound both Warm / Gritty but also Cold / Sparkly. Its still a piano though.

This level of control wasn't found in any models out there - so I decided to sit down and train my own.

Getting consistent timbre-locked keybeds that actually LOCKS across multiple diffusion calls was hard af but I did it.

I documented the full journey here for those who want to learn a bit or be entertained.

https://youtu.be/x0KnmzH8Mmk

There is also a longer walkthrough if you just want to see the keybeds in action.

https://x.com/RoyalCities/status/2097733712293109842?s=20

No-talk / Showcase only Demo

https://x.com/RoyalCities/status/2097733715543609445?s=20

any finally the huggingface page

https://huggingface.co/RoyalCities/Foundation-1

I've also provided full write ups on the inferencing pipeline associated with the interface so this should allow basically anyone else to go and vibe code their own text to synths if they wanted :)

https://github.com/RoyalCities/RC-stable-audio-tools/