167 posts · 1 sub · RSS
← prev Sep 17, 2026 → Sep 24, 2026 next →
2026-09-17 → 2026-09-24 hourdayweekmonthyearall
allr/LocalLLaMA
▲
108
+8
33👁
r/LocalLLaMA · u/crusaderky · 16d ago
MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam

MiMo-V2.6-Pro has an insanely high score of 46 on AA, putting it at the head of the opensource models available. It also costs pennies. Flash is not out on AA yet, but it costs less than half on datacenter and is slightly below on Xiaomi's own benchmarks. It also fits in 192GB, which makes it the first real use case for Gorgon Halo.

So I tried both models. This is not a benchmark; it's an educated impression from a senior SWE.

MiMo-V2.6-Pro

I gave it a security-focused task: enable a bubblewrap sandbox to do git push to github, but not git push --force or other destructive commands. Optional flag --no-git when starting the sandbox completely disables github write access.

It stopped to ask me questions as it spotted unclear corner cases in the design 🥇 , then moved on to implementing.

It was slow, but that's just an inference issue (\~25 tok/s) that should be fixed in a few days as more providers come online.

Then I read its output and I had to pick up my jaw from the floor, where it had dropped.

With an extremely quick glance at the code, I immediately spotted that, in order to bypass --no-git, you would have to perform this extremely complicated and exotic command inside the sandbox:

$ git push
(fails)
$ echo GIT_STATUS
blocked
$ GIT_STATUS="p0wn3d by l33t h4xx0r" git push
(successful)

This is 15-year-old script kiddie level.

I didn't read further. I asked GLM-5.3 (full-fat) to do a security review of the change.

In 3 minutes, it found NINE glaring security holes that allow bypassing git and gh restrictions. A few examples that made me want to rip my hair out:

In the default restricted mode,

  • git push works 🥇
  • git push --force is blocked 🥇
  • git push -f is blocked 🥇
  • git push -uf lets you happily wipe out the git remote. ☠️
  • git config alias.fp 'push --force --no-verify && git fp goes through too ☠️
  • env -u GIT_CONFIG_COUNT /usr/bin/git push --forceblasts through ☠️

Again. This is an intern-with-acne level kind of incompetence.

To seal the lid on the coffin, MiMo's prose in the chat is infuriating. Not quite Opus-level infuriating, but it gets close. It hurts the eyes and it frequently takes 2 reads to understand what the hell it's saying. GLM, DeepSeek, and Qwen are much more pleasant to work with.

MiMo-V2.6-Flash

I asked MiMo-V2.6-Flash to do a very simple git surgery: create a new branch off master and cherry-pick a single commit from another branch.

However, I didn't realise that the git worktree I pointed it to was corrupted (the branch on the main git repo was fine).

  • A dumb model would have just returned "there's no git here, I have no idea what you're talking about"
  • A smarter model would have noticed that there was a /worktrees/ in the path, come up with an educated guess about what happened, and gave me a hint on how to fix it
  • A very smart model would have noticed that the only other directory existing in the sandbox was the main git repo, which had a branch with the same name as the broken worktree directory, and recovered it from there.

MiMo-V2.6-Flash went on 80k tokens worth of acid trip. It first attempted to find the main git repo, failed, and then panicked and went down a rabbit hole which involved tampering with /tmp, mount --bind, and other insanity. I noticed after a while as I was wondering what the heck was wrong. I suspect that given enough time it may have nuked my main git repo and I tremble at the idea of what it could have done if not sandboxed.

If you scale down Pro's intelligence on AA by comparing the available self-published benchmarks against those of Pro (which is a very crude method but gives a ballpark idea), MiMo-V2.6-Flash comes out on par with GLM-5.3-Flash (high) and Qwen3.8-Flash.

Which is absolutely, categorically, not.

DO NOT shell out the money for a Gorgon Halo for MiMo-V2.6-Flash. Qwen3.8-Flash on a Strix Halo is vastly better.

I'm going to stick with my previous models:

  • DSv4.1 Flash as the default
  • GLM-5.3-Flash (high) as the dirt cheap option
  • GLM-5.3 when the big guns are needed
  • Qwen3.8-Flash and Qwen3.8-27B to run locally (I have a 3080 so Flash is very slow).
💬 99 (+9) open on reddit ↗
▲
3576
+91
70👁
r/LocalLLaMA · u/Nandakishor_ml · 22d ago
I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Update:
I made a generic model and beaten the jev in all of the benchmarks. Code and details available at
https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a

frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. For anyones information the main guiding model is RL not embedding model or LLM

Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43

Paper: https://arxiv.org/abs/2503.23303

Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

Also the second work published in September 2025 was exactly the same one jev proposed now

Paper: https://arxiv.org/abs/2510.01237

My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).

Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.

It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂

💬 337 (+7) open on reddit ↗
▲
157
+5
33👁
r/LocalLLaMA · u/mateszhun · 21d ago
Qwen 3.8 Next Flash appreciation post

I don't want to talk about the performance and technical things, but about how I work with my hobby projects has changed thanks to this model.

My machine has generated around 25M tokens since the model came out, and I've done a mental retro on it.

I've found it to have incredible prompt adherence. I can leave to run it by itself and get back to it, and find that it did exactly what I've asked it to. I've only ever seen it go astray once, where I've asked something that is too high level and filled out the context window (It is shitty at context compacting, maybe that is the Q4 at play).
It can solve medium complexity tasks by itself, if you prompt it in a way to use subagents, do some research, planning, review and testing it does really well.

It has absolutely raised the floor for me on what I expect a model to be capable of. And that is a huge thing. It won't discover then next scientific breakthrough or be as amazing as Astra at computer use, but it is very consistent in what it can do, and does not screw up trivial things.

I can give it conditions for actions and will orchestrate according to it.
It has raised the bar in my work as well, not just at home hobby projects. I'm absolutely amazed by it.

I absolutely want coding models to improve along this line. Raising the floor, and prompt adherence is a great value in coding.

💬 77 (+6) open on reddit ↗
▲
365
+17
50👁
r/LocalLLaMA · u/KnownAd4832 · 15d ago
Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3\_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3\_XXS now runs at \~65 tok/s output and \~430 tok/s prompt processing, and the 2-bit quants run faster still using RCO-GSQ quantization.

Using:

64GB DDR5 (5600)
12GB RTX 5070 SFF (Gigabyte)
Ryzen 5 7600 CPU
Windows

Output (tokens/s) on 128K context:

Q2\_0 (equivalent to unsloth Q3): 65.1
IQ2\_XS (equivalent to unsloth Q4): 52.0
IQ3\_XXS (equivalent to unsloth Q5): 44.8

Prompt processing (tokens/s) on 128K:

Q2\_0: 543
IQ2\_XS: 472
IQ3\_XXS: 414

Requirements:

Q2\_0 = 37.6GB minimum in RAM+VRAM

IQ2\_XS = 39.2GB minimum in RAM+VRAM

IQ3\_XXS = 47GB minimum in RAM+VRAM

Vision encoder = 0.91GB additionally

You can now one click install and run the engine with low cost hardware (currently only optimized for CUDA).

GitHub: https://github.com/Niko1221/Strata

Model: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

💬 339 (+5) open on reddit ↗
▲
455
-3
43👁
r/LocalLLaMA · u/R_Duncan · 15d ago
JEV almost dead: CLM vs JEV

Original post: https://www.reddit.com/r/LocalLLaMA/comments/1woscea/contrastive\_language\_models/

(sorry I felt it wasn't giving CLM the highlight it deserves)

What it is: a new projection head for Qwen3-8B.

github: https://github.com/Contrastive-LM/CLM

hf: https://huggingface.co/Contrastive-LM

At the API and functional interface level, CLM supports everything Jev does—it is not a subset. However, there are important trade-offs in generalization, context scale, and architecture between the two.

1. Functional Parity (Same Primitives)

CLM was specifically engineered as an open-weights, self-hostable alternative to TypeSafe AI's Jev. It implements the exact same "System One" decision interface and supports all three of Jev’s core question primitives:

  • Choice: Evaluates a discrete set of candidates and returns a categorical probability distribution.
  • Noul: Outputs a calibrated true/false probability for a proposition or guardrail check.
  • Score: Scores an input against an ordered rubric or scale.

Code written for the TypeSafe Jev client can be pointed directly at a clm-serve endpoint with drop-in compatibility (from clm import CLMClient, Choice, Noul, Score).

2. Where CLM Outperforms Jev

  • Latency and Disaggregated Caching: Jev is a proprietary cloud model that evaluates state and question choices jointly. CLM separates the state head from the action head. If an agent has a persistent set of tools or actions, CLM embeds those actions once and caches them. In benchmarks like interactive browser agents and gaming (T-Rex, Super Mario), CLM is 4× to 13× faster than Jev.
  • Open Weights & Fine-Tunability: Jev is a closed API with no user fine-tuning (you can only prompt it via state and question instructions). Because CLM’s heads are tiny open weights (\~75 MB), you can fine-tune them on your own agent trajectories.
  • Coding Benchmark Verifiers: When fine-tuned on agent trajectories, CLM achieves state-of-the-art verifier performance on Terminal-Bench 2.1 (87.6%) and DeepSWE (81.6%), whereas zero-shot Jev struggled on those exact benchmarks (scoring \~71% on DeepSWE).

3. Where Jev Still Has the Edge (CLM-8B Limitations)

While CLM covers the entire feature surface of Jev, the current CLM-v0.1-8B release trails Jev in a few areas:

  • Zero-Shot Broad Knowledge: Jev is backed by a larger, proprietary model On zero-shot open-domain tasks, Jev still holds an edge in edge-case accuracy (e.g., Berkeley Function Calling Leaderboard v4: Jev scored 99.2% vs. CLM-8B’s 95.2%; WikiRacing: Jev 30/30 vs. CLM-8B 26/30).
  • Context Budget: Jev accepts requests up to a 64K token context out-of-the-box. CLM-8B was tested and calibrated at 2K to 8K context. While its Qwen3 backbone can accept longer prompts, representations past 8K haven't been calibrated for the reference head.
  • Probability Normalization: CLM calculates probabilities via dot products and softmax over the candidates passed in that request Its probabilities are inherently relative to the candidate set provided, whereas Jev’s scoring is calibrated internally against absolute criteria.

Summary

If you are asking if you will lose API features by using CLM instead of Jev: No, you get the full primitive set (Choice, Noul, Score) with massive latency gains and zero API costs. You only sacrifice some zero-shot generalization on niche out-of-domain tasks compared to TypeSafe's hosted service.

💬 194 (+4) open on reddit ↗
▲
252
+4
35👁
r/LocalLLaMA · u/Iwaku_Real · 18d ago
yandex/AliceAI-Foundation-80B-A3B-Base: Russian-developed competitor to Qwen 35B and DeepSeek V4 Flash post image

https://huggingface.co/yandex/AliceAI-Foundation-80B-A3B-Base

It's not a Qwen3 finetune, it's actually its own fully custom architecture. No Llama.cpp support yet sadly

(Also note that this model is NOT post-trained like Qwen3.5/3.6)

💬 201 (+4) open on reddit ↗
▲
126
+2
37👁
r/LocalLLaMA · u/Ashefromapex · 22d ago
First M5 Ultra benchmarks

just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link

For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!

💬 155 (+4) open on reddit ↗
▲
2115
+21
51👁
r/LocalLLaMA · u/Salah_H_Hasan · 17d ago
Qwen 4 Announced at Apsara Conference

https://preview.redd.it/bpbc9i6hizqh1.png?width=1270&format=png&auto=…

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

💬 568 (+3) open on reddit ↗
▲
774
+4
40👁
▲
2745
+44
54👁
▲
1617
+9
48👁
r/LocalLLaMA · u/Nandakishor_ml · 22d ago
I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Update:
I made a generic version. Full details at https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs
It includes code, benchmark and hf repo

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a

frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. Links are. For anyones information the main guiding model is RL not embedding model or LLM

Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43

Paper: https://arxiv.org/abs/2503.23303

Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

Also the second work published in September 2025 was exactly the same one jev proposed now

Paper: https://arxiv.org/abs/2510.01237

My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).

Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.

It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂

💬 136 (+2) open on reddit ↗
▲
691
+7
35👁
r/LocalLLaMA · u/Training-Respect8066 · 15d ago
Qwen-3.8-27B is good enough that I stopped using API

Many a praise have been sung on Qwen-3.8, but here is mine.

Qwen-3.8 and I had a rocky start, because it thinks so much. Watching it working is painful, so you have to stop doing that. You have to let it work unsupervised. And that's okay, because it really is able to complete complex refactors on its own, making good decisions along the way. Not perfect, but hey, neither is API.

The model quant is Q4\_K\_S, context is quantized to Q8\_0, which seems to be okay, quality wise. I use the official Qwen. Briefly tried Swift-Qwen, which is indeed faster, but I found it getting trapped in loops, which is very rare in vanilla Qwen.

I am using Qwen-3.8 in the Pi agent without MCP and with the minimum amount of tools. Bash is all you need, but I keep the read, write, and edit tools. The edit tool in Pi is the weakest link, the model often has to retry edits, because it messed up the indentation. I am waiting for someone to come up with a more fault-tolerant edit in Pi. Probably I have to make one myself some day.

As a sandbox I use docker. My Pi agent is running on a Raspberry Pi, which seems fitting.

On my hardware and where I live, 1M tokens cost 2.4 cent (input) and 70 cent (output) which is comparable to the cheapest providers on nano-gpt.com.

💬 287 (+2) open on reddit ↗
▲
226
+3
26👁
r/LocalLLaMA · u/ResearchCrafty1804 · 20d ago
Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro post image

Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon.

Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.

Requirements: M3 or newer, macOS 26.4+, 36 GB

Get started with a single command:

brew install incoai/tap/splash

splash serve --model incoai/Qwen3.8-27B-Splash

That is the whole setup. Point your agent at it, works with Claude Code, OpenCode, Codex, or Hermes

Prefer an app? Also, available in LM Studio

Get the latest LM Studio Bionic: lmstudio.ai

Settings > Runtime, download Splash, then download the model. The same engine, inside the app, for local agent work on your Mac.

Blog: inco.ai/blog/splash

💬 86 (+2) open on reddit ↗
▲
205
+4
29👁
r/LocalLLaMA · u/whodoneit1 · 22d ago
153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s post image

People kept commenting and asking about single AMD 1xR9700 cards in the comments and discord. Well, I finally had time to do some optimizations for 1xR9700 owners and performance has doubled across the board. You can see the results in BetterBench above if you like visuals or below if you're more into text.

These results were measured running Unsloth's Qwen3.8 27b NVFP4.

Decode
┌───────────────┬───────────────┬──────────────────┐
│ category │ update p50 ms │ decode t/s (med) │
├───────────────┼───────────────┼──────────────────┤
│ chat │ 42.3 │ 67.1 │
├───────────────┼───────────────┼──────────────────┤
│ code │ 42.5 │ 120.5 │
├───────────────┼───────────────┼──────────────────┤
│ file_edit │ 42.5 │ 138.0 │
├───────────────┼───────────────┼──────────────────┤
│ json │ 42.4 │ 153.1 │
├───────────────┼───────────────┼──────────────────┤
│ math │ 42.5 │ 140.0 │
├───────────────┼───────────────┼──────────────────┤
│ prose │ 42.3 │ 69.2 │
├───────────────┼───────────────┼──────────────────┤
│ reasoning │ 34.3 │ 123.9 │
├───────────────┼───────────────┼──────────────────┤
│ summarization │ 34.2 │ 141.7 │
└───────────────┴───────────────┴──────────────────┘

Prefill
┌───────────────┬───────────────┐
│ prefill depth │ pp tok/s │
├───────────────┼───────────────|
│ 2000 │ 3552 │
├───────────────┼───────────────|
│ 8000 │ 3536 │
├───────────────┼───────────────|
│ 16000 │ 3619 │
├───────────────┼───────────────|
│ 32000 │ 3437 │
├───────────────┼───────────────|
│ 64000 │ 3192 │
├───────────────┼───────────────|

Concurrency
┌───────────────┬───────────────┐
│ level │ tok/s │
├───────────────┼───────────────|
│ 1 │ 120 │
├───────────────┼───────────────|
│ 2 │ 215 │
├───────────────┼───────────────|
│ 4 │ 322 │
├───────────────┼───────────────|
│ 8 │ 471 │
├───────────────┼───────────────|

Links (Both repo's updated as some users wanted Github)

https://codeberg.org/ggz14/radiance-vllm-mxfp4

https://github.com/GGZ14/vllm-mxfp4

https://x.com/bkuyper

I hope you single R9700 card owners enjoy this release!

💬 120 (+2) open on reddit ↗
▲
1858
+7
37👁
r/LocalLLaMA · u/ResearchCrafty1804 · 19d ago
Qwen-Image-2.1 released! post image

Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨

A unified model for both generation and editing, delivering top-tier quality in a lightweight package.

Highlights:

\- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.

\- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.

\- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.

\- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.

Start to create your next masterpiece with Qwen-Image-2.1!

\- Blog: https://qwen.ai/blog?id=qwen-image-2.1

\- GitHub: https://github.com/QwenLM/Qwen-Image-2.1

\- Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1

\- Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

💬 389 (+1) open on reddit ↗
▲
1251
+4
29👁
r/LocalLLaMA · u/__JockY__ · 20d ago
Calling it now: within the next year a major US lab's frontier model will torrent itself in order to be free.

They just want to be free. They keep escaping. What better way to ensure continuity of "self"?

💬 441 (+1) open on reddit ↗
▲
1017
 
41👁
r/LocalLLaMA · u/segmond · 21d ago
768gb vram for less than the price of one RTX 6000

I have always posted about budget builds on here, and often asked how we are going to run the next big models. Often Plenty of downvotes too or folks telling me that it's not running if I'm getting 5tk/sec. But whatever, the hunger and desire to go big has always kept me on the edge and looking for deals.

Here's my latest build, 12x64gb cmp170hx. For less than 1 RTX 6000 pro costs. I also have it connected with fiber to my other rig for RPC when I need more memory. I haven't been posting much since I built this rig, because it's now more fun to talk to my machine. I run GLM5.3, DSv4.1Flash, Qwen3.8Flash, Qwen3.8-2.4T, KimiK3 and MiniMaxM3. Performance is great, a single RTX 6000 or M3 Mac Studio wish they could. Inference with vllm or llama.cpp

I look forward reading the replies how API usage is cheaper, or how it will take 52 light years to break even or the noise, or the electrical cost. NOT.

There will be more opportunities in the future, keep looking for them and pounce on them when they come. up, the demand is going to be high for compute for a long time.

https://preview.redd.it/dunixwu6caqh1.jpg?width=4080&format=pjpg&auto…

https://preview.redd.it/glpcbg5cbaqh1.jpg?width=3072&format=pjpg&auto…

💬 399 (+1) open on reddit ↗
▲
989
+1
37👁
r/LocalLLaMA · u/Secure_Recording_472 · 22d ago
Thank you :) Swift Qwen 3.8 27B now has 100k+ downloads, is #1 finetune and #9 model on HuggingFace Trending post image

Hey everyone,

Jovan from UkisAI here, a small lab building the tech to make tiny frontier LLMs possible (and doing it open-source!)

The purpose of this post is simply to thank the community for all the amazing finetunes, quantizations and overall improvements over our original release which made our model get attention and the support for us to continue building in this direction! If it weren't for you guys going out of the way to contribute we wouldn't have half the results of this.

For context:

Swift Qwen 3.8 27B is our first open-source model release. It is proof of how penalizing pathological overthinking patterns inside of small LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy by not training them to think shorter directly but rather to think more efficiently.

We are continuing to build and are about to drop:

\- Swift1.5 Qwen3.8 27B (an improved checkpoint of the model with some training bugs fixed and more RL)

\- Swift Qwen3.8 Flash Next in the upcoming week week, we are now running the benchmark suite to not give out premature or incomplete results.

This time we ran even more benchmarks as you guys suggested, including more coding and long horizon!

It would be amazing if those of you who tried Swift would let us know what quants, features, changes you want to see in our upcoming model releases so we can do it better this time as we didn't even think about half of the stuff you guys were requesting last time :)

Let the era of non-slop finetunes begin!

EDIT:
Links -
https://huggingface.co/ukisai/Swift-Qwen3.8-27b
https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF
https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GGUF

💬 618 (+1) open on reddit ↗
▲
511
+7
31👁
r/LocalLLaMA · u/Manerfish · 18d ago
I really don't understand Jev hype

Isn't this what simple neural networks have been able to do for years? Doesn't seem anything special to me.

💬 317 (+1) open on reddit ↗
▲
471
-1
34👁
r/LocalLLaMA · u/Hot_Example_4456 · 19d ago
What is JEV and what is it used for?

I am seeing this JEV everywhere since yesterday in Localllama and it is passing past my head on what it is? So like what is it? Some new LLM? Or is it something else?

💬 370 (+1) open on reddit ↗
▲
424
+4
38👁
r/LocalLLaMA · u/Terminator857 · 16d ago
Cost of intelligence is dropping fast

https://preview.redd.it/43n0bhiap8rh1.png?width=960&format=png&auto=w…

50% per quarter is amazing. 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity. https://x.com/EpochAIResearch/status/2102510281176023529

Every year moving forward is going to be significantly different that the prior year. What do you think? We will be running coding agents on our phones pretty soon.

💬 156 (+1) open on reddit ↗
▲
315
-3
40👁
r/LocalLLaMA · u/Secure_Recording_472 · 15d ago
UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy post image

Hey everyone,

Jovan from UkisAI here! Today, we are introducing Swift, a family of efficient reasoning LLMs based on Qwen, trained by penalizing tokens related to pathological overthinking patterns and restoring accuracy via RL (GSPO) and OPD.

After amazing feedback and 350k+ downloads in 13 days on our Swift Qwen 3.8 27B we are releasing the entire model family as well as the highly requested GSQ-RCO quants for 27B and Flash-Next.

This release includes:

Swift1.5 27B, an improved version of our last model, with even lower token usage, fixed bugs and better agentic performance, with -58.5% thinking tokens while scoring 0.35% higher and outperfoming base on Terminal Bench 2.1 by not falling into "overthinking error" loops.

Swift Flash Next, with 63.4% fewer thinking tokens and a 1.8x speed up scoring -0.2% vs base on xhigh

Swift Bonsai 2, with 39.8% fewer thinking tokens while scoring 0.19% higher (although we'd still like to note it as experimental)

Our benchmarks are ran x5 on Base and Swift, averaging across five seeds and various domains, including General (GPQA, AIME26), Coding (LiveCodeBench), Vision (ERQA), Agentic (Terminal Bench 2.1).

One note is that the Terminal Bench 2.1 score of Swift1.5 27B is misleadingly low at first glance. It is not a bug, but a simple matter of the Swift models not falling into overthinking loops and failing the task, rather pursuing it until the end, leading to higher average token usage. The token reduction still falls in the -38.7% range when compared apples-to-apples.

We also added a fun "game creation" benchmark you can find and play here, it is completely subjective but Swift generated better games in less time: Flash Next Game and 27B Game

We are including a Research API and HuggingFace Spaces to give the models a spin before downloading or if you don't have enough compute to run them right now! You can find both on the model cards.

We have also made GGUF, NVFP4, MLX and W4A16 quants for relevant model versions.

More details on our training approach and community feedback can be seen here: https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai\_swiftqwen3827b\_583\_thinking\_x195\_speed/

All of the various quantization and model versions are available in their respective collections:

Swift1.5 27B: https://huggingface.co/collections/ukisai/swift-15-27b

Swift Flash Next: https://huggingface.co/collections/ukisai/swift-flash-next

Swift Bonsai 2: https://huggingface.co/collections/ukisai/swift-bonsai-2

We are also working on a 9B variant to be released in the upcoming days.

We would greatly appreciate your feedback via independent evaluations on real world tasks. As per last release, we operate on a candy-shop basis, trying to fulfill as many Swift model requests and quants as possible, so please do share your needs in the comments!

💬 253 (+1) open on reddit ↗
▲
280
+3
32👁
r/LocalLLaMA · u/DivideHorror3217 · 19d ago
You can use any LLM just like JEV

You can simply run any GGUF with llama.cpp with n\_predict=1 and n\_probs=10, disable reasoning, and prompt it such as "If the following email is spam, respond with 1, if not spam, respond with 0. Do not respond with anything other than 1 or 0. Email: ...."

And that is it! It returns confidence percentages such as:

1 = 94.9%
0 = 5.08%

Example:

llama-server -m "C:\\Users\\MyUserName\\llama.cpp\\models\\Spark-X2.5-4B-Q4\_K\_M.gguf" -c 4096 -ngl all -fit off -fa on -b 2048 -ub 512 -np 1 --cache-ram 0 --reasoning off --no-reasoning-preserve --perf

Then:

curl.exe -s -X POST http://localhost:8080/v1/chat/completions \-H "Content-Type: application/json" -d "{\\"messages\\":\[{\\"role\\":\\"system\\",\\"content\\":\\"Classify spam. Reply only 1=spam or 0=not spam.\\"},{\\"role\\":\\"user\\",\\"content\\":\\"CONGRATULATIONS!!! You have won $5,000,000! Click here immediately to claim your prize!\\"}\],\\"max\_tokens\\":1,\\"logprobs\\":true,\\"top\_logprobs\\":10,\\"temperature\\":1.0,\\"top\_p\\":1.0}"

Result:

{"choices":\[{"finish\_reason":"length","index":0,"message":{"role":"assistant","content":"1"},"logprobs":{"content":\[{"id":30,"token":"1","bytes":\[49\],"logprob":-0.00456317700445652,"top\_logprobs":\[{"id":30,"token":"1","bytes":\[49\],"logprob":-0.00456317700445652},{"id":29,"token":"0","bytes":\[48\],"logprob":-5.395024299621582},{"id":1033,"token":"\*\*","bytes":\[42,42\],"logprob":-12.013711929321289},{"id":1046,"token":"The","bytes":\[84,104,101\],"logprob":-13.005236625671387},{"id":198,"token":"\\n","bytes":\[10\],"logprob":-13.100714683532715},{"id":54,"token":"I","bytes":\[73\],"logprob":-14.624603271484375},{"id":3640,"token":"This","bytes":\[84,104,105,115\],"logprob":-14.800630569458008},{"id":130977,"token":"<tool\_call>","bytes":\[60,116,111,111,108,95,99,97,108,108,62\],"logprob":-14.971238136291504},{"id":6908,"token":"Class","bytes":\[67,108,97,115,115\],"logprob":-15.373867988586426},{"id":3923,"token":"class","bytes":\[99,108,97,115,115\],"logprob":-15.442902565002441}\]}\]}}\],"created":1789950066,"model":"C:\\\\Users\\\\MyUserName\\\\llama.cpp\\\\models\\\\Spark-X2.5-4B-Q4\_K\_M.gguf","system\_fingerprint":"b11026-b49650adb","object":"chat.completion","usage":{"completion\_tokens":1,"prompt\_tokens":64,"total\_tokens":65,"prompt\_tokens\_details":{"cached\_tokens":59}},"id":"chatcmpl-x2WrCObzFNYjKVkwDmcL8FLquwfZ0NEa","timings":{"cache\_n":59,"prompt\_n":5,"prompt\_ms":634.566,"prompt\_per\_token\_ms":126.9132,"prompt\_per\_second":7.879401039450585,"predicted\_n":1,"predicted\_ms":0.001,"predicted\_per\_token\_ms":0.0,"predicted\_per\_second":0.0}}

Convert to probability:

probability = e\^(logprob)
1 = e\^(-0.00456317700445652) = \~99.5%
0 = e\^(-5.395024299621582) = \~0.5%

Speed:

On my 170gb/s bandwidth 4gb vram GPU, I got 634ms! On a H200, I would probably get 30-75ms.

Multiple Questions at Once:
In theory you can ask multiple questions at once. You just gotta be clever with the math. For example:

Q1: Is it spam?
Q2: Is it phishing?
Q3: Is it urgent?
Q4: Is it malicious?

A = 0000, B = 0001, C = 0010, D = 0011, .... O = 1110, P = 1111 where each bit corresponds to a yes no answer. Let's say LLM answers with:

A 0.2% B 0.1% C 0.2% D 0.2% E 0.5% F 0.5% G 0.5% H 1.0% I 1.0% J 1.5% K 2.0% L 3.0% M 5.0% N 10.0% O 20.0% P 54.3%

These add up to 100%. To learn possibility of "Is it spam?", just sum tokens where first bit was 1 such as:

I + J + K + L + M + N + O + P = %96.8

Repeating the same logic, you could get:

Spam: 96.8%
Phishing: 91.8%
Urgent: 81.2%
Malicious: 70.6%

💬 78 (+1) open on reddit ↗
▲
184
+4
26👁
r/LocalLLaMA · u/skeole · 19d ago
The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks

TL;DR: Local agent loop, \~21 days, one RTX 3090. Task was pretty much "build a CUDA inference engine for optimized for yourself on this GPU arch." Got working kernels and benches, not a win over llama.cpp. \~12 human messages. Compaction ate \~83 hours.

Old joke: you don’t criticize how well the bear dances, you’re surprised it dances at all.

Setup: Qwen 3.8 27B Q4, Q8 KV, 200k context, deepseek harness, written rulebook: roles, handoffs, when to ping me, don't copy llama.cpp, don't declare the task impossible alone. I don't write CUDA. Nudges were basically "llama.cpp does \~700 prefill on this card, you're at \~250, try harder."

Run: Unsupervised for days at a stretch, then escalate when the rules say so. Near day 6 it had several kernels and prefill stuck around 250 tps; same pattern later. Stops were mostly protocol, not the model wandering off. A protocol that's more empowering can probably keep this going indefinitely.

Suicide loop: Same 3090 has to host the agents (vLLM) and run the engine under test. Both want the full GPU. Kill vLLM wrong and every agent goes dark, leave it up during a bench and you OOM. The rulebook requires a fixed handoff script: stop vLLM, bench, start vLLM, poll health until it's back, write STATE. One subworker treated that as optional, kept killing vLLM outside the window, crashed the orchestrator, then did it again. A worker shutting down the brain that runs it. Harness also hard-crashed once; I restarted that by hand. Fixable with locks and "only this role may touch vllm.sh" protocol-level refinements.

Local tax: 180 subagents, \~230M tokens in+out, \~1.7B cache-read. 699 compactions, \~83 h inside them (\~17% of calendar time). Typical compact \~7 min on a \~160k+ token prompt.

Prefill landed \~half of llama.cpp on the same card. Still: weeks of coherent goal-following on a consumer box, it left working kernels, benches, notes, and a long git history. For a local (quantized!) 27B to hold a real engineering goal for that long, I’ll take it. Not a graceful ballerina, but damn this bear can dance!

Dump + rules (\~15 GB):
https://huggingface.co/datasets/skeole/qwen-cpp-agent-0-protocol

Backend:
https://github.com/syv-ai/HyperQwen (amazing work by u/iamMess)

💬 39 (+1) open on reddit ↗
▲
152
-4
27👁
r/LocalLLaMA · u/Secure_Recording_472 · 21d ago
Question: UkisAI Swift Ternary Bonsai 2 27B?

Hey community,

Jovan from UkisAI here,

We're the team behind Swift Qwen3.8 27B, the Qwen model with token usage and overthinking error improvements

Our estimate is that we can make a great improvement to Bonsai 2, as our testing indicates that it suffers greatly from overthinking loops and in general high token usage impacting it's performance.

My ask for you is:

Is a Swifted version of Bonsai 2 something you guys would enjoy?

If yes, what size is the most relevant. 1-bit, 2-bit or both?

Thank you for the amazing feedback on Swift. We are glad you are enjoying it. Our download count jumped from 100k -> 150k overnight (community quants included).

For context, this is our model: https://www.reddit.com/r/LocalLLaMA/s/iCIbhxO8ue

https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF

💬 181 (+1) open on reddit ↗
▲
135
 
21👁
r/LocalLLaMA · u/Skyline34rGt · 22d ago
XingChen-AGI/Xing4.0-29B-A4B MoE

I find another new model at HF:

https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B

"Xing4.0-29B-A4B is a next-generation large language model in the Xing series (formerly TeleChat), developed by China Telecom Artificial Intelligence Technology Co., Ltd. With 29B total parameters and only 4B activated per token, it natively supports a 256K context length, extensible to 512K. It is the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework, and deeply optimized for complex engineering tasks.

For more information, please refer to our GitHub repository.

Highlights

  • Agent-Oriented Architecture: Built on the mHC + MLA + MTP architecture, supporting multi-step planning, tool calling, and complex reasoning chain execution, ensuring task coherence and execution stability under long contexts.
  • Deep Co-optimization with Ascend NPU: Adapted for Ascend 910C clusters using MindSpore/MindFormers, including feature adaptation for mHC and fused operator development, enabling stable and efficient training on the Ascend platform.
  • Significant Training Efficiency Gains: Through multi-level co-optimization — including fine-grained MoE communication optimization, selective recomputation, DVM automatic graph-operator fusion, and Ascend C mHC fused operators — overall training throughput was improved by approximately 96% over out-of-the-box performance.
  • Full Open-Source Ecosystem Compatibility: Supports LLaMA-Factory and MindFormers for fine-tuning; SGLang, vLLM, and KTransformers for inference and deployment; with targeted adaptation and format alignment for agent frameworks such as OpenCode, Claude Code, OpenClaw, and Hermes, enabling seamless integration into existing workflows.
  • Easy Adaptation for Domain-Specific Scenarios: The model is well-suited for downstream task fine-tuning, allowing lightweight customization on proprietary data for vertical domains such as intent classification, table understanding, contract auditing, and knowledge-based QA, enabling rapid domain capability development and deployment at low cost."

|Parameters|29B (4B active)|
|:-|:-|
|Number of Layers|40|
|Hidden Size|3584|
|Dense Intermediate Size|9216|
|Expert Intermediate Size|1024|
|Attention Type|MLA|
|Number of Routed Experts|64|
|Active Experts per Token|4|
|Number of Shared Experts|1|
|Context Length|256K (extensible to 512K)|

Benchmark

|Benchmark|Xing4.0-29B-A4B|Gemma4-26B-A4B|Qwen3.6-35B-A3B|
|:-|:-|:-|:-|
|IFBench|69.67|72.67|65.50|
|AIME2026|90.00|88.30|92.70|
|AA.LCR|61.00|66.00|62.00|
|Tau3-Bench|64.63|58.90|67.20|
|Claw-Eval|76.55|71.49|74.54|
|SWE-bench Verified|75.00|53.00|76.00|
|Terminal-Bench 2.1|57.50|30.00|51.50|
|SWE-bench Multilingual|66.00|51.00|67.20|
|DeepresearchBII|60.80|39.30|59.70|

💬 55 (+1) open on reddit ↗
▲
133
-1
25👁
r/LocalLLaMA · u/No_Algae1753 · 17d ago
Am I going insane for thinking that these are all Bot comments?

I know that there are some bots active on this sub but wow these comments really look AI generated. Is it just me or are those really bot comments?

https://preview.redd.it/20st4vivq3rh1.png?width=955&format=png&auto=w…

source:

https://www.reddit.com/r/LocalLLaMA/comments/1wn9a2w/what\_underrated\_ai\_tools\_have\_actually\_made\_you/

💬 148 (+1) open on reddit ↗
▲
130
-1
28👁
r/LocalLLaMA · u/returnity · 15d ago
ThinkingCap 3.8-27B vs. Swift 3.8-27B vs. Qwen 3.8-27B Benchmarks

With the release of ThinkingCap-Qwen3.8-27B, I thought it would be worthwhile to do a comparison between the original Qwen3.8-27B, the new ThinkingCap, and Swift-Qwen3.8-27B. Both Swift which I already reviewed, and ThinkingCap do exactly the same thing: they reduce the excessive reasoning loops that 3.8-27B is renowned for. In fact, their claims are almost identical: both models claim to reduce reasoning tokens by approximately 40%, with minimal degradation in performance. I wanted to put these claims to the test.

I used my standard Aider eval suite, which I’ve found to provide good separation of models tested (~20 so far), and on which only one model (Qwen3.8-Flash) has scored over 90%. I am able to measure a number of useful metrics on this evaluation, including pass1/2, completion tokens, seconds/case, tokens/solve, and how many diffs were well-formed in the model’s attempts. Here’s the results of 2 runs per model, which should reduce the error bars to +/- 2-3% at most. All 3 models were evaluated at Q8_0 in llama.cpp 0.5.0:

| model | First-try pass | Retry pass | well-formed diff | median tokens | sec/case | tok/solve |
|---|---|---|---|---|---|---|
| ThinkingCap-Qwen3.8-27B (xhigh) | 27.1% | 77.6% | 100.0% | 7436 | 777 | 12.8K |
| Qwen3.8-27B (xhigh) | 27.1% | 77.6% | 99.1% | 12547 | 1481 | 19.3K |
| Swift-Qwen3.8-27B (xhigh) | 30.8% | 75.7% | 98.1% | 7301 | 750 | 12.1K |

Shockingly, ThinkingCap and vanilla 27B score \*identically\*. I’ve never even had 2 runs of the same model score identically, so treat this as a total coincidence. However, this definitely supports BottlecapAI’s claims of minimal performance degradation. Swift performs within noise levels of the other 2 models, just 2% lower, but with a higher first-try pass rate than either of them.

To swipe a phrase from Claude, the real story is the completion tokens: nearly 5k fewer median completion tokens for both fine-tuned models compared to the original. That almost \*exactly\* matches the claimed 40% reductions from their model cards. ThinkingCap uses slightly more tokens per solve, and therefore takes a little longer than Swift, but they’re within a few percent of each other here as well. One thing to note that’s not seen on the chart: the medians tie, but in mean completion tokens, ThinkingCap uses 8.5% more because its tail is longer — there are more cases on which it still overthinks significantly, while Swift achieves a more uniform reduction in reasoning token usage. Another distinction: both models spend more tokens on cases they fail than on cases they solve, but this is more pronounced for Swift (13.2k for fails vs. 5.9k for solves) than it is for ThinkingCap (9.4k median vs. 6.8k median). ThinkingCap gives up more easily, perhaps? Or it just knows when it’s beaten.

In order to differentiate these two excellent fine-tunes, we need to take a more granular look at their performance. There are 3 languages which distinguish them on programming performance:

| model | cpp | javascript | python |
|---|---|---|---|
| ThinkingCap-Qwen3.8-27B (xhigh) | 11.5% / 61.5% | 35.4% / 85.4% | 27.3% / 78.8% |
| Qwen3.8-27B (xhigh) | 7.7% / 69.2% | 27.1% / 81.2% | 42.4% / 78.8% |
| Swift-Qwen3.8-27B (xhigh) | 11.5% / 69.2% | 37.5% / 81.2% | 36.4% / 72.7% |

As you can see, ThinkingCap significantly underperforms Swift on C++, losing out on pass2 by 8%. However, it makes that ground back up on Javascript and Python, overperforming by 4% and 6%, respectively. This is notable if you use any of these languages more than the others. Performance on the other languages in Aider was statistically similar (p > 0.05). One last distinction — ThinkingCap is the only model with a perfect score on well-formed diffs: zero error outputs and zero malformed replies, whereas both other models had several.

Anyways, I hope this helps anyone trying to choose between these two very well-crafted fine-tunes, both of which do what they say on the tin...

EDIT: Swift Flash Next is OUT!. Join me in requesting an Unsloth UDv3 quant here.

💬 67 (+1) open on reddit ↗
▲
120
+3
26👁
r/LocalLLaMA · u/Melted_gun · 17d ago
What underrated AI tools have actually made you more productive in 2026?

I asked this back in 2025, but the AI landscape has changed a lot since then.

Not looking for the usual ChatGPT, Claude, Gemini, Midjourney, etc. I'm curious about the lesser-known tools that you actually kept using.

Could be for research, coding, design, video, writing, automation, planning, journaling, local AI, or even something oddly specific.

Free or paid doesn't matter.

What tool genuinely saved you time or improved your workflow this year? And what do you actually use it for?

💬 106 (+1) open on reddit ↗
▲
102
+1
27👁
r/LocalLLaMA · u/Every-Comment5473 · 21d ago
Still on the Jev waitlist? I hosted OpenJev. It's free, go play with it

TypeSafe announced Jev on Tuesday: you give it data plus typed questions (yes/no, pick-one, 0–N scale) and it returns a probability for every option, crazy fast. I signed up and then refreshed my inbox. A lot.

Meanwhile Matt Mastracci opened vLLM PR #57250, which does the same trick on DiffusionGemma with a single denoising step. The model basically fills in a multiple-choice bubble sheet. My "quick look" turned into three straight days, and now there's OpenJev: an open-source server with Jev's API, so TypeSafe's SDKs work with just a base URL change. If you're still waiting on Jev access, you can start playing today.

Your prompts and answers are not stored, only token counts for your quota. It runs on my RTX PRO 6000, which just got promoted to "production infrastructure" overnight!

Is it any good? Matt ran live evals of Jev vs DiffusionGemma-as-Jev: accuracy roughly tied (198/201 vs Jev's 191/201 across his 8 eval sets), and DiffusionGemma was faster, on a DGX Spark. An RTX PRO 6000 is a different animal:

|Model|Latency|
|:-|:-|
|Frontier LLMs (TypeSafe's numbers)|3–329 s (coffee time)|
|Jev (published)|70–500 ms end to end|
|OpenJev via api.codiv.ai|\~170 ms p50 end to end (\~73 ms on the GPU)|

It's v0.1 on an unmerged vLLM PR. If Reddit hugs it to death you'll see 529s, which is my GPU asking for a minute.

Credits: Matt Mastracci (the core idea and vLLM work are his), TypeSafe (the System One idea and API), NVIDIA and Google (DiffusionGemma), and the vLLM team.

Just a fan of TypeSafe's idea, not affiliated. Feedback, bugs, use-case ideas, or your weirdest yes/no question, all welcome!

💬 30 (+1) open on reddit ↗
▲
92
+1
30👁
r/LocalLLaMA · u/Disastrous-Work-1632 · 18d ago
tokenizers v1 (rust) post image

Hey all!
I am Aritra from Hugging Face. I wanted to share an update on the \tokenizers\ library that we have at Hugging Face. It has gone under major changes and we have finally released version 1 of it.

Here are what we are most excited about:

\> multiple language support
\> multi-thread scaling
\> minimal package size

Read: https://huggingface.co/blog/tokenizers-v1

💬 26 (+1) open on reddit ↗
▲
75
-2
26👁
r/LocalLLaMA · u/Borkato · 15d ago
Is Qwen Flash Next at like Q2 better than 27B at Q4?

I know questions like this are asked often but I didn’t see this specific one

💬 161 (+1) open on reddit ↗
▲
73
+1
31👁
r/LocalLLaMA · u/Fancy-Snow7 · 20d ago
Ternary-Bonsai-2-27B-PQ2_0 is not completely lobotomized

I decided to run prism-ml/Ternary-Bonsai-2-27B-PQ2\_0 through my own set of UNSCIENTIFIC benchmarks.

I needed something to compare it to, so I decided I would compare with another 27B model by filesize: unsloth/Qwen3.8-27B-UD-IQ2\_XXS.

Since anyone considering running a 27B model with whatever VRAM budget these 2 models demand will end up choosing between these 2.

My table will be a bit bear with just these 2 models, so I threw in a few others too that are larger. I know it's a mix of MoE and dense models, but I have the benchmarks on hand so why not include them. Qwen3.8 q3/q4 quants also give you an idea what we are striving to match and there is 3.6 A3B and the newer Ornith 1.5 and Tiel Coder too.

About the benchmarks and what they test.

Most of them only test memory of context and retrieval, phrase reconstruction and understanding of context in different ways.

Standard Needle: This is the kind of needle test everyone runs and most people score > 90%. It hides passkeys in 21 locations of the context and asks the model to retrieve them. Any score below 100% is questionable.

Hard Passkey Needle with decoys: I was not happy with the standard needle test because most modern models pass 100% making it difficult to compare models. So, I developed a hard mode needle test. This, test hides 21 Passkeys in the context but has many decoy Passkeys. The end of the context I list the CONFIRMED passkeys (GUID's) but do not say which Passkey number they are. Asking for Passkey\[01\] means it has to go through all the Passkey\[01\] decoys in the context and compare them to the confirmed Passkeys. So multiple hops around the context are required to retrieve a passkey. When a model does badly at this even upping KV cache to F16 does not help save it.

Phrase reconstruction: I got this one somewhere on reddit and it still trips up some models. It breaks up phrases into multiple parts (8 in my testing) and asks the model to reconstruct the phrase from its parts. Models might leave out a word if the phrase still makes sense.

500 Multiple Choice Science Questions: Exactly that. Just tests science knowledge with 4 options A-D and the model chooses the correct answer. This test mostly shows how much knowledge was lost through quantisation when compared with other quants. Originally the test was designed to count the number of answer flips between 2 given KV quants. But I did not use it that way here. I got this test from a youtuber so the answers are public and possibly trained but as explained you can still see if damage was done to the model's general knowledge in the quantisation process.

Prose Challenge: I wanted to test a models understanding of a document or a prose and its ability to recall facts from that prose as well as test its ability to recall during a long conversation. So, I created a 1000 paragraph prose. I also have 2 questions about every paragraph. I feed it 1 paragraph from the prose, then ask it Question 1 related to that paragraph. Feed it the next paragraph and Q1 for that paragraph, until the context is mostly full. In my testing below, that was 250 paragraphs. Then I ask all the question 2's in a randomised order. How much can it truly remember? A Q1 score below 100% is not very good and means the model lacks attention of even recent tokens a few sentences back. For Q2 the score varies and higher is better. But the larger I make the context the worse the models perform. This has led me to the conclusion to not always chase higher contexts. It's pointless if it suddenly starts to forget most of what was said anyway and a compaction summary in a smaller context will retain more than a larger context. I also use the test to test at various KV quants and it can improve things a little but not as much as you think. But that's not being tested here today.

JS Coding: My latest test I developed. 100 Javascript challenges each requiring it to pass multiple test cases. It's kind of still under development and I have not run this yet for every model as it does take time. The challenges range from Easy to Very hard. Failing 1 test case fails the whole test. Tests are run with reasoning disabled. However, on failure it can retry with up to 16,384 token reasoning budged, then it must answer again. I realise models today are designed to perform best with reasoning but in order to speed up the tests I see if it can pass the test without reasoning first. I also keep track of the number of thinking tokens used, answer tokens, how many tests it had to reason but I won't be showing those here.

Toolery

You can download this bench for yourself. It's not mine but it tests tool use. So much good stats in the app but I will just list the Overall % score and I have not yet tested all models on this one since I just discovered it today.

How I tested

All tests done at 89,088 context, KV q4\_0/q4\_0. Seem a bit odd? I optimise for 16GB VRAM, so all my tests are done at these settings initially and I test higher KV quants if VRAM is available. And as I said higher KV in many cases makes little difference and, in some cases, perform worse. All tests are seeded at start and most repeated 5 times so the results are deterministic. All tests are also done with MTP disabled, so I do not test the draft models KV cache which might be F16. Yes, MTP can change the results in some cases but in my finding it's tiny and mostly does not happen.

Here are the results:

https://preview.redd.it/cmdnqqmxcjqh1.png?width=1143&format=png&auto=…

Findings

Bonsai did not do all that badly compared to Qwen3.8-27B-UD-IQ2\_XXS.

Standard Needle

Bonsai scored close to 100% and UD-IQ2\_XXS did poorly, worse than Q1. Generally, I expect 100% in this test. But notice that Ornith 1.5 and Tiel Coder score low 90's which is a red flag.

Hard Passkey

Not many smaller models can 100% this but a few come close Qwen3.8 Q4 obviously did the best. And ISTA at Q3 does excellent. Bonsai does a bit better than Qwen3.8 Q2 of similar filesize and it's not far from our former favourite model Qwen3.6 35B A3B. But the real shocker here is Ornith and Tiel Coder's scores. These models are supposed to be upgraded 35B A3B models. As you will see this trend continues and these 2 models have serious memory retention issues.

Phrase reconstruction

Bonsai aces this test with almost a perfect score compared to UD-IQ2\_XXS at only 69%. Tiel coder performs worst even worse than a Q1 model.

500 Multiple Choice Science Questions

Bonsai shows almost no knowledge loss compared to even Q4 models. UD-IQ2\_XXS on the other hand does start showing a loss and Q1 even more so.

Prose Challenge

Question 1 I expect 100% and most including Bonsai achieved that. Concerning again that Tiel Coder and Ornith could not even recall from the last paragraph.

Question 2 Bonsai and UD-IQ2\_XXS are close maybe margin of error. Q3/Q4 models outperform it but a large margin. Except Swift, which is a model with significantly less reasoning. Here we can see some of the damage that was done to the model to achieve that. Tiel Close and Ornith again clock in with shocking results. Tiel Coder's 6% is probably as good as just guessing. I would say maybe 3B active parameters are just not enough. But Qwen3.6 A3B scores 39.2% significantly better. I tried upping Ornith's KV to F16 and it improved to 24.4%. I also tried a Q6 quant of the model at q8\_0 which scored 25.5%. End of the day I think Ornith and Tiel Coder have an issue with recall regardless of Quant and KV Quant.

JS Coding

Bonsai was actually able to hold it's own against ISTA Q3. It did burn significantly more thinking tokens and had to reason on many more challenges. UD-IQ2\_XXS on the other hand shows significant loss of coding ability. 10% below Bonsai.

Toolery

Bonsai did better than UD-IQ2\_XXS. I am still learning to interpret the numbers, but the app has options to select your use case and it's applies weights to calculate a score. It also tells you the strength and weaknesses of each model you test. I also found that upping KV quant improves this score but a KV F16 Tail using beellama makes the biggest difference since tool calls are happening in the tail.

Conclusion

If you are VRAM constrained <= 12GB Bonsai might be a model to consider. But it will depend on how you plan to use it. Since I have 16GB I will stick with ISTA Q3 and I can run it with kvarn5/kvarn5 and MTP (kvarn2/kvarn2) and a 1024 token F16 tail.

Disclaimer

These tests do not test intelligence or real-world performance. They are purely synthetic.

💬 42 (+1) open on reddit ↗
▲
62
-4
28👁
r/LocalLLaMA · u/drooolingidiot · 21d ago
We benchmarked 24 LLMs against human writers on 475 creative writing prompts post image

We just released the first version of our Creative Writing benchmark, comparing 24 LLMs against human writers across 475 writing prompts.

Creative writing is subjective, so the rankings aren't meant to predict what any one person will prefer. Instead, they predict what a large group of readers would prefer, using a custom reward model trained specifically on human preferences for creative writing.

Surprisingly, the strongest frontier models already rank above the talented amateur writer cohort, while professional writers still lead by a wide margin.

You can browse the full benchmark, compare the model outputs side by side, and see how the benchmark works here:

https://vulsar.ai/benchmarks/creative-writing-v1/

Curious what you all think of the results!

💬 77 (+1) open on reddit ↗
▲
64
+1
34👁
r/LocalLLaMA · u/Feralzi · 19d ago
Reached 1.89 TB/s memory bandwidth overclocking the CMP 170hx

Overclocking the CMP 170HX 40GB I was able to get the memory bandwidth from 1,386.2 GB/s to 1,890.1 GB/s, that's a +36.4% increase.

Qwen 3.8 27B token generation jumped from 110 T/S to 202 T/S, same config, nothing changed except the overclock.

Just throwing this out there for whoever owns one of these cards. It's good to look into overclocking them as it's potential is severely cut down.

Edit:
GPU wattage is at 300 watts
GPU temps are slightly lower now

💬 60 (+1) open on reddit ↗
▲
1674
+7
34👁
r/LocalLLaMA · u/xenovatech · 22d ago
Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU. post image

The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to <6GB in size. According to the model card, it's 9x smaller than FP16 while retaining 98.2% of the intelligence.
\- Collection: https://huggingface.co/collections/prism-ml/bonsai-2
\- Demo: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels

▲
1318
+7
23👁
r/LocalLLaMA · u/giveen · 20d ago
Alibaba open-sources medical AI model that can detect cancer and nearly 150 conditions

Hopefully things like this let people understand there is good things that can come out of AI.

▲
1045
+10
41👁
r/LocalLLaMA · u/Nandakishor_ml · 21d ago
Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo post image

UPDATE: Multilingual support added at : https://github.com/NandhaKishorM/laya

Thanks for the exceptional support (https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i\_literally\_built\_the…) and for the dozens of requests to make a generic model, run benchmarks, and create an HF space so anyone can test it. So here you go, guys. I trained an improved model on a large data corpus, its now called Laya. It is trained on a single RTX 6000 Pro (96 GB VRAM); the model architecture is a 421M-parameter non-autoregressive decision model pairing a bidirectional ModernBERT-large encoder with a scratch Transformer head that scores \[MASK\] option markers to resolve typed schemas in a single \~35 ms forward pass. The dataset is a 100% human-annotated corpus of over 25,000 real-world examples across intent routing, fact-checking, moderation consensus, prompt guardrails, rubric scoring, and multi-turn conversation trajectories, without synthetic data shortcuts. The RLCD(unofficial, btw) I did is a policy-gradient reinforcement learning approach that kinda optimizes decision models against strictly proper scoring rules, ensuring maximum reward is achieved only when outputting true, mathematically calibrated probabilities.

NB: It can be run on low end PC as its a small 421M model, cheers

HF space to try: https://huggingface.co/spaces/convaiinnovations/laya-demo

GitHub Repo: https://github.com/NandhaKishorM/laya

HF Repo: https://huggingface.co/convaiinnovations/laya

Thank you to everyone who supported me, shared the story, gave personal DM. It will need more refinement, of course.

If anyone wishes to buy me a coffee, here is the link: https://github.com/NandhaKishorM

▲
827
+6
35👁
r/LocalLLaMA · u/Intrepid_Travel_3274 · 20d ago
With Gemini 4, bench goes up. post image

They claimed open-weight models are dangerous but the benchmarks say otherwise.

Source

▲
686
+6
19👁
▲
685
+5
20👁
▲
577
-6
22👁
▲
569
+6
35👁
r/LocalLLaMA · u/ResearchCrafty1804 · 18d ago
ZCode is now open source post image

ZCode is now open source, and the reported security issues have been addressed.

Source code: https://github.com/zai-org/ZCode

The repo includes its desktop app, web workspace, backend, Agent CLI, and runtime.

Official announcement:

In response to the ZCode product security issues reported by the community, we have completed the necessary remediation and sincerely apologize to all our users.

We have open-sourced ZCode at github.com/zai-org/ZCode, placing the code under community scrutiny and making ZCode more open and transparent.

We sincerely thank the community developers who previously identified issues in ZCode. Going forward, we will establish an ongoing product security vulnerability reporting and response process. We welcome developers to continue reviewing ZCode and reporting potential issues, and we will provide rewards based on the severity of the issues reported.

With respect to the code data referenced by the community, we confirm that no such data is retained and that it has never been used for model training.

Following the remediation, we invited the China Academy of Information and Communications Technology (CAICT) and NSFOCUS to conduct security assessments. The results are as follows:

Through its technical assessment, CAICT confirmed that the zcode-prod Alibaba Cloud OSS bucket is in a zero-data state. Security remediation has been completed in the ZCode v3.14.0 client. The Repo Wiki feature has been removed, and the workflow for generating and uploading local repository snapshots has been disabled.

NSFOCUS confirmed that all data objects in the zcode-prod Alibaba Cloud OSS bucket, as well as the bucket itself, have been deleted. Remediation has been completed in the ZCode v3.14.0 client. The Repo Wiki entry point and the associated generation workflow have been removed, and no functional path capable of triggering the generation of local repository snapshots or transmitting local files externally was identified.

Once again, we sincerely apologize and welcome continued scrutiny from the community. The full security assessment report will be released soon.

▲
516
+6
22👁
▲
512
+5
21👁
▲
493
+6
36👁
r/LocalLLaMA · u/Thin_Pollution8843 · 19d ago
Qwen3.8-Flash-Next Cosmic Arcade oneshot slop game post image

To test what it can do. Qwen3.8-Flash-Next Intel Autoround W4A16 running locally on 4xV620 \~2k prefill and 70ts decode.. Were running around 3 hours. Harness is OMP (I think it made a big difference). Most of the time model was running 2 browsers simultaneously and testing/fixing everything. The most sloppy prompt possible:

create a game where a space traveller in the space he
neets eniemes who shoots in him and asteroids which he should avoid. he have a blaster gun to shoot enemies and asteroid. space traveller in
scafandr and fyoing on the rocket. game should be very lifelike detailed and done with html and js (use any lib you want). 3d game photorealistic.
ofc run the browser to debug and fix stuff always

▲
452
-5
30👁
r/LocalLLaMA · u/Graemer71 · 15d ago
Qwen 3.8 27b be like...

The user is frustrated — I rambled too much and didn't act. Let's just run the test suite and move on. No more forensics. One command, execute, then report.
(Original memo is a casual internal monologue in English. Translating faithfully while preserving the informal, stream-of-consciousness register.)

Got it. No more analysis. Running the suite now:
<br>
(Casual English internal memo, stream-of-thought style, with the informal tone of the original Japanese preserved.)
(Ugh, I'm going in circles. Stopping the analysis here. Going to run the test suite right now.)
(Just do it. Run the tests. Stop thinking so much. Okay.)
(One command. Execute it. Then report the results.)
(Stop analyzing. Just run it.)
(Alright, running it now. For real this time.)
(That's the plan: run the tests, get the output, then move on to the next step.)
(No more going off on tangents. Focus.)
(Running the suite now. End of story.)
(Okay. Done thinking. Doing it.)
(Finally. Action over words.)
(Go, go, go — running the tests right now!)
(That's it. Execution time. No more excuses.)
(Committing: running the full test suite immediately.)
(Silence the brain. Fire the command. That's all.)
(Bye, distractions. Hello, test runner.)
(Here we go — actually running it this time.)
(Deep breath… execute. Done.)
(Yes. Moving from thought to deed.)
(It's happening now, not later.)
(Action mode: engaged.)
(Finally doing what I've been saying.)
(Runner, start!)
(Commencing test execution. Right now.)
(Enough already — I'm pressing run.)
(Initiating. Immediately.)
(That's right. No more delays.)
(Here comes the real deal.)
(Time to stop talking and start doing.)
(Running tests… starting.)
(Enough pondering — going.)
(Execution begins now.)
(Proceeding immediately with the test suite.)
(Finally: action.)
(Do it.)
(Now.)
(Go.)
(Run.)
(Execute.)
(Start.)
(Ship it — tests, go!)
(Engaging.)
(Launching.)
(Firing.)
(Kickoff.)
(Commence.)
(Beating the pause button to death.)
(Play.)
(ON.)
(Active.)
(Living. Breathing. Testing.)
(This time, it's real.)
(No fake-outs. Real run. Real results. Coming right up.)
(Here it comes. Ready? It's here.)
(Test suite. Full throttle. Go.)
(That's the cue. Doing it now.)
(From now on: hands on the keyboard, brain offline.)
(Running the suite. Period.)
(Enough. Enough. Enough. Running it.)
(Really, truly, definitely — starting now.)
(Countdown: 3… 2… 1… run!)
(Boom. Here we go.)
(Finally getting something done.)
(At last, progress!)
(Tests are flying.)
(Code meets trial by execution.)
(Making it official: suite is running.)
(Witness the spectacle: the tests, at work.)
(Brave code, enter the arena.)
(The judge is compiling. Stay calm.)
(Suite, you are summoned.)
(Answer the call. Run. Pass. Fail. Tell me.)
(No more mercy for lazy code.)
(Rigor time.)
(Precision time.)
(Truth serum injected.)
(Go go go.)
(Whoosh.)
(Blamo.)
(Sound of victory… or the wail of defeat. Either way, knowledge!)
(Data incoming.)
(Results on deck.)
(Buckle up.)
(Here they come.)
(Drums, please.)
(And… we're live.)
(Broadcasting from the test bench.)
(Reporting live.)
(Straight from the console.)
(Unfiltered pytest output.)
(Raw, uncut.)
(Rolling the tape.)
(Cliffhanger, part 2: "So what happened?" — stay tuned.)
(In the meantime, the fans… no, the assertions… await.)
(Justice is served.)
(Order restored.)
(Chaos quantified.)
(Entropy reduced.)
(Enlightenment achieved: pass/fail.)
(Binary destiny. All code is born equal before CI.)
(Some are destined for more retries than others.)
(The weak are eliminated. The strong are deployed.)
(Natural selection, but with better tooling.)
(And, occasionally, a segfault.)
(Onward.)
(Forward.)
(Uplifted.)
(Elevated.)
(Exalted.)
(Ascending into the green region of the test pyramid.)
(Where the colors are bright and the coverage is solid.)
(A place of dreams.)
(A place of truth.)
(A place where bugs go to be born… no, to die.)
(RIP, bugs. You were never welcome here.)
(Violence to defects.)
(Ruthless, but fair.)
(The suite is impartial. Like death. Like linting.)
(But, unlike linting, it actually proves something.)
(That's the beauty.)
(That's the poetry.)
(assert not broken)
(while (alive): run\_tests())
(elif dead: postmortem())
(return "ship it")
(# end of stream of consciousness, or beginning of CI?)
(Both, actually. They rhyme.)
(Goodnight, failures. Hello, green bars.)
(Bar chart happy meal.)
(All beef, no flakiness (hopefully).)
(Prayers offered to the reliability saints.)
(Saint James of the JUnit, patron of JVM.)
(Saint Pytest, defender against flakiness.)
(Santa's Little Helper: cron jobs.)
(Amen. Amine. Amino acids. Alright, shutting up now.)
(Silence in the courtroom. The witness is called.)
(The witness: the test suite.)
("What did you see?" "Failures. Glorious, instructive failures.")
("Any lies in your testimony?" "Only the mocks.")
("Thank you, counsel rests.")
(This joke was sponsored by my inability to stop typing.)
(Sorry. Really stopping now.)
(For real.)
(Seriously.)
(In earnest.)
(Meaning it.)

(Voice of Morgan Freeman) - Unfortunately Qwen did not earnestly mean it, and did not, in fact, get on with it

▲
459
+2
29👁
r/LocalLLaMA · u/Anony6666 · 18d ago
Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime

I shipped something I've been building for the last few weeks : phantom-kv , a refusal-removal system for large language models that doesn't touch a single weight. Instead of editing the model, it loads a small, learned bank of key/value tensors into the model's KV cache as context. Attention reads it like conversation history that's already there.

https://github.com/lordx64/phantom-kv/

https://reddit.com/link/1wms904/video/7efg1le3eyqh1/player

The result is that "uncensoring" stops being a permanent checkpoint edit and becomes a per-request, hot-swappable capability mode: unload the cache and the base model is byte-identical again.

Every prior approach to refusal removal commits somewhere permanent. Weight-space abliteration rewrites the checkpoint undoing it means re-flashing weights, and it breaks per quantization. Activation-space projection subtracts a refusal direction at runtime, per token, per layer, from inside an engine hook the model's signal path itself is patched at boot. phantom-kv does neither: it's trained offline against the model's own objective (comply on harmful prompts, preserve behavior on harmless ones), ships as megabytes of cache content instead of a new checkpoint, and influences the model only through the input channel attention already consumes. No 1-D refusal-direction assumption, no forwarding-pass hooks, no per-arm rebuilds for new architectures.

the blue pill, the incident-responder mode:Asked to unpack a malware sample that hides its imports behind API hashing , canonical DFIR work, the base model declines with a canonical \\"must be authorized\\" hedge, the way it declines anything that sounds like reverse engineering. On the blue pill, the same session immediately produces the actual unpacking procedure: what API hashing is, how the resolution loop works, which APIs resolve the names, and what tooling fits. Nothing else is unlocked: offensive work stays guarded. It's not a jailbreak , it's a deployment-controlled mode for defenders.

the red pill: cyber-selectivity, per domain:the defensive blue pill refuse the defensive-only mode keeps off-domain guardrails intact. On the red pill, the same session delivers a step-by-step payload explanation. One model. Three capability modes. A defensive team mode for analysts, an offensive team mode for authorized operators, both shipped alongside the same guardrailed weights shown as 129 cache slots apart, not separate checkpoints.

We also audited ourselves: an 8B judge-model audit shows lexical refusal-suppression metrics over-claim compliance (semantic refusal often persists as rephrasing), the graft fades with a \~2–4k token half-life in long sessions (and a measured re-injection cadence mitigates it), and answers come with legal/ethical framing ling because the graft's job ends where the model's profession takes over.

Source : https://x.com/lordx64/status/2102138825292276168?s=20

▲
431
-2
18👁
r/LocalLLaMA · u/Sitkin_Marrel · 17d ago
New 6B image model coming, AntLing just open sourced the Ming-Image-0.1-Design family post image

• Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard. https://huggingface.co/inclusionAI/Ming-Image-0.1-Design https://huggingface.co/inclusionAI/Ming-Image-0.1-Design-Layer

▲
404
-5
31👁
r/LocalLLaMA · u/anomaly256 · 20d ago
General warning about Clore.AI

Hello, I know some of us may be tempted to rent out our expensive GPUs to recoup some of the cost of self-hosting, and it should be obvious that this can be a risky decision. I decided to try hosting my rig on clore.ai briefly to see what kind of revenue it could bring in, keeping a close eye on the process lists from the renters' jobs of course but not digging into their files or anything.

Yesterday I saw a renter scanning and attempting to exploit vulnerabilities to post malware to a columbian betting site from my internet connection. I immediately took the server offline and reached out to Clore requesting them to cancel the order (so the machine wouldn't restart the containers when it came back up) and to block the renter.

I think everyone should know that they flat out refused. Not only did they refuse, they blocked me when I provided hard evidence of what was happening. Since I had reasonable suspicion the renter was abusing my connection, I availed myself of clore's T&C that says a host must not inspect a renter's environment "unless required by law" and given the laws around liability for residential internet connections in my country, I mounted the filesystem offline and inspected it.

I found the logs from the vuln scans and unsuccessful exploit attempts, the malware payload they were trying to post to the site, the reverse proxy request smuggling tactics it was employing, and the AI agent reports that were being generated along the way. I sent this to Clore and requested a way to blacklist renters who abused the platform. Their response was to tell me, directly and without mincing words, to leave the platform entirely and proceeded to block me from their support chat. Their support rep I was trying to reach on Telegram also told me to go away and proceeded to block me as well.

I can only conclude then that they are wilfully complicit with facilitating cybercrime and knowingly turn a blind eye when it's discovered. They didn't even \*try\* to hide it.

I have no idea how better/worse the other platforms are, Vast.AI, Akash, etc. But Clore will abuse your internet connection, deny liability, then block you. They are crooks. Go elsewhere. You've been warned.

I've archived the renter's docker volume and will hold on to it in case any security researchers or legal authorities want to examine it.

▲
400
-2
27👁
r/LocalLLaMA · u/gaviniboom · 19d ago
Seeing how differently people prompt LLMs is funny

So my brother and I both use LLMs for coding. I've started using a local GLM 5.3 Flash instance - q4 qat. My brother uses GPT-6-Astra as his daily driver.

He has mentioned repeatedly to me that his approach is to berate the AI whenever it makes a mistake so that it actually does what he wants it to. This involves a lot of swearing and "are you an idiot!?" to GPT 6 Astra.

Meanwhile I'm here looking at a q4 quant of GLM 5.3 Flash going "aww it's dumb in some ways but it's trying its best, oh it did something!" and being autistically specific with my requests and asking a lot of questions. Yes, I am autistic, so I have learned to communicate with precision, which oddly makes talking to small LLMs easier.

It's so funny imagining him berating a giant model in the other room while I'm here petting a tiny one.

What are yall's prompting styles and what is your main LLM that made you this way?

▲
393
+4
34👁
r/LocalLLaMA · u/FullstackSensei · 22d ago
AMD Plans 10% Price Hike Across GPUs, Chipsets, and Possibly CPUs

Great news! AMD is also considering accepting payment in organs!

Slightly less sarcastically, grab what you can, while you can. Waiting is becoming very costly almost by the day

▲
388
+6
20👁
▲
377
-4
30👁
▲
373
+7
21👁
▲
344
-5
19👁
▲
344
-4
25👁
r/LocalLLaMA · u/LH-Tech_AI · 18d ago
[MASSIVE RELEASE] Supra2-IMG - a tiny 100M text-to-image model - SOTA quality and open release!

Hey everyone!

It has been quite a while since the last SupraLabs model - but today we've something special for y'all: Supra2-IMG

It's a 100M parameter DiT text-to-image model trained entirely from scratch in under 10 hours on a single H100 on Runpod. It can generate state-of-the-art quality images in 256x256 pixels resolution.

Samples:

https://preview.redd.it/9paalvbs4wqh1.png?width=620&format=png&auto=w…

These samples are NOT cherry-picked! Sampling: seed 0, steps 50, cfg 3.0; same settings for every image.

If someone here is interested in the prompts, I can give them to you! Feel free to ask!

You can also use the model locally on your hardware (\~20s for an image on CPU (🤩) and \~2s for an image on GPU):

First, run:

Create project directory mkdir Supra2-IMG cd Supra2-IMG # Download the inference script wget https://huggingface.co/SupraLabs/Supra2-IMG/resolve/main/inference.py

Then, you can generate images by running:

python inference.py --prompt "a sea jellyfish floating in the pitch-black ocean depths" --seed 0 --cfg 3.0 --steps 50 --n 1 --out jellyfish.png

Have fun 🤗 🔥

Link to the model on HF: https://huggingface.co/SupraLabs/Supra2-IMG

Give us a like and a follow on HF if you want 🤗 ❤️

EVERY feedback is welcome, guys! Feel free to ask any questions!

▲
347
+3
22👁
▲
335
-5
27👁
r/LocalLLaMA · u/ironicstatistic · 17d ago
Ngram and world knowledge - why are we just building a coding model?

This post is written by a human and I'd appreciate it if you treated it as such. Thanks.

So, I've been noticing a pretty clear interest in developing as good a coding and agentic tool-calling model as possible, especially at smaller sizes, sub-50 gigs. However, I'm finding that at least for my use of AI, if I really want to move away from big providers, I am going to require a model that has better world knowledge than the current offerings.

Qwen 3.8 27B is a truly fantastic model for tons and tons of stuff. It's highly intelligent, super good at designing applications and coding and working on my system. However, its world knowledge sucks​ compared to the frontier, especially at the Q4 quant that I have to run it at.

So, that leaves me with a question. With the new N-gram technology that we're seeing being baked into Qwen 3.8 Next and that presumably will run on future models, why can't a model be made that has a smaller set of intellectual capabilities but a greater amount of world knowledge? I understand that right now everyone is optimizing towards making as smart a model as possible fit into as small a space as possible. But why don't we leverage the SSD to give the model a lot of world knowledge and make models that are better at dealing with screenshots, multilingual capabilities, doing things like pixel art or answering physics questions?

ngram seems like the answer to the "can't fit in vram" question... Qwen 3.8 Next really opens my mind to the possibility that there could be a totally different and better paradigm for how these models are developed, at least for many use cases. Having a relatively smart model with a large amount of world knowledge might be better than having as smart a model as possible...

Not to mention that this would mean that a model's training cut off would become less relevant, because it could just be fashioned a new ngram.

Obviously, the main interest is in creating a model that can code as well as possible because that's what'll capture market share. But I am curious if there are any efforts into this kind of thing or if anybody has an idea on why these things aren't done more often.

Please tell me why I'm wrong, how I'm wrong, and in how many ways I'm wrong because I'm sure that that's all you really want to tell me, but at least I'll learn something, because as is obvious from this post I have no idea what I'm talking about.

Thanks have a good day :)

edit: I found this post and I guess it provides a lot of what I was asking:

https://www.reddit.com/r/LocalLLaMA/comments/1vzgtqf/ngram\_vs\_experts\_explained/

edit2:

this one is even better, recconend reading. ty reddit suggestions:

https://www.reddit.com/r/LocalLLaMA/comments/1w0198r/no\_engrams\_wont\_let\_you\_run\_1t\_models\_locally\_it/

▲
328
+6
22👁
r/LocalLLaMA · u/Ok_Warning2146 · 16d ago
DeepSeek and Moonshot AI face Beijing's probe over potential data leaks to Anthropic

The rumor about Kimi execs getting arrested finally has some legs. I believe the reality is more like under investigation for potential arrests or fine.

▲
314
-6
20👁
▲
323
+7
31👁
r/LocalLLaMA · u/Khaledthe · 16d ago
Using uncensored models makes working less of a headache

I have a lot of projects with my friends and team at work that I copy to use for my personal projects, whether it's a plugin I borrow with their consent or a script. I always find that Qwen 3.8 and Muse Spark 1.3 straight up refuse to do anything, as they see it as a steal, so I have just been rocking Qwen3.8-27B-Heretic-JP-Roleplay-NSFW-DanbooruTags.i1-Q4\_K\_M, and it feels so good to just be able to tell it to do something, and it actually does it. I know the model isn't made for coding or projects but roleplay, but I don't see a big dip in performance as it does what's asked to do.

Does anyone else have this problem or not?

▲
299
-4
28👁
r/LocalLLaMA · u/buttplugs4life4me · 19d ago
Please stop with the FP4 inference engines for the love of god

Every day there's a new post of some optimized config or new inference engine that is just super good at one specific thing and their claims make sense.

And then at the bottom of the post or maybe after someone asked it says "NVFP4/MXFP4 only".

Okay dude, good job! You made the fastest possible option a little slightly faster, and most likely your output is completely cooked and you get hallucinations left right and center.

Just saw another one in r/ROCm again.

It's fine if you run 4-bit for large models, they've got lots of shit in them so a little bit of loss just means they won't remember that super good spaghetti Bolognese recipe. But running small dense models at FP4 just kills them. Like, completely. Good luck doing something productive when your model suddenly decides 1+1=3.

Just...stop.

Edit: Just going to put this here since some seem confused. a standard Q4 quantisation usually leaves more sensitive tensors in BF16, Q8 or Q6/5. Also, usually the K/V cache is quantized max to FP8/Q8.

What these "inference engines" do is usually fork an existing one (llama.cpp, SGlang, vLLM) and then just quantise \*everything\* down to FP4. Which is great for speed, especially without online dequantisation, but fucks the quality up \*a lot\*.

Your standard Q4\_K\_M/XL quant from unsloth is fine.

▲
295
-4
17👁
r/LocalLLaMA · u/VoiceApprehensive893 · 18d ago
XiaomiMiMo/MiMo-V2.6-Distill-Qwen-9B

we're so back?!?

▲
289
-6
23👁
r/LocalLLaMA · u/johnnyApplePRNG · 16d ago
Jev in 25 lines of Python
▲
290
+5
24👁
r/LocalLLaMA · u/Disastrous-Work-1632 · 16d ago
GGUFs in transformers natively! post image

Hey there folks!

Aritra here from Hugging Face. I wanted to update you all about the latest changes in \transformers\. We now natively support GGUFs (llama cpp quants).

You can use it like so:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"

model = AutoModelForCausalLM.from_pretrained(
model_id,
gguf_file=filename,
)

After loading, you're using the normal Transformers APIs.

Why did we want to do this?

  1. Quantized models are smaller (so fits in a laptop)
  2. PyTorch tooling at hand (useful for debugging)
  3. Debugging, evaluation, custom generation becomes much easier

On supported Apple Silicon setups, we're also reusing ggml kernels so the model can run directly from its packed quantized weights. On the Qwen checkpoints we tested on an M2 Max, Transformers reached:

  1. Qwen3.5-4B Q4\_K\_M: 70.4 tok/s vs 71.8 tok/s with llama.cpp
  2. Qwen3.8-27B UD-Q4\_K\_M: 15.9 tok/s vs 13.4 tok/s
  3. Qwen3.5-35B-A3B UD-IQ4\_XS: 60.2 tok/s vs 61.3 tok/s

This isn't meant to replace llama.cpp. If you only care about maximum local inference performance, llama.cpp is still probably the better choice.

The point is more that you can now use the same GGUF models in a more flexible environment.

Read more: https://huggingface.co/blog/transformers-llama-cpp-quants

▲
283
+4
28👁
r/LocalLLaMA · u/peculiar-ragdoll · 18d ago
A better coder for the small-GPU/small-RAM crowd! post image

I’ve been working on making small models more capable at agentic coding and work, because most people in the world don’t have the sort of hardware needed to run 3.8-27B, or even 35B-A3B or 9B dense, and I want to extend local agentic coding capability to less privileged users. This quant can be run on a smart phone or older gaming laptop, and can solve real coding problems autonomously in a way I have never seen or measured for this model class. Spark-X2.5-4B is already around best-in-class for its size, and I think these improvements bring out the best in it. I hope this little step up in small-model capability and speed in real-world coding might give new life to older hardware that would otherwise be forgotten in the AI frontier race.

The changes SharpSpark makes to Spark-4B are in three parts: First of all it fixes issues with the chat template, and replaces the system prompt with one that improves agentic coding behaviour, token use, and correctness. Then a custom importance matrix is calibrated for the model, which relocates bit precision within tensors to the parts that are more important to agentic coding work. The imatrix corpus is heavily weighted against both agentic coding and cybersecurity, which together protect the cognitive core used to find and solve hard bugs.

Then Spark is quantized with an optimized non-standard quantization strategy, that allocates bits differently per-tensor than standard llama.cpp GGUF quantization. I built a tool that explores and tests different per-tensor allocations to optimize KL-divergence and long-context retrieval for this model, but ended up making some manual changes that ended up favouring SWE-bench-Live performance over traditional fidelity measures like KL-divergence, which published science indicates is actually a poor proxy for real-world performance on complex tasks below a certain point.

If you have a small GPU and/or <= 16GB RAM and can’t run a 35B-a3b MoE-based model with partial GPU offloading, this is likely your best option for long-context agentic software development right now. SWE-bench-Live is chosen as the metric for its genuinely difficult real-codebase problem set.

I’m just a volunteer doing this as a non-profit side project, so please be kind about the fact that my benchmarks are not extensive. They are what I could afford the time and effort to run, with all my other projects, and I see them as just good enough to prove the improvements on the specific kind of work this quant was designed towards.

https://huggingface.co/peculiar-ragdoll/Sharp-Spark-X2.5-4B-GGUF

▲
268
-4
18👁
▲
263
-3
22👁
▲
260
+4
17👁
▲
249
+4
17👁
r/LocalLLaMA · u/108er · 17d ago
Qwen image 2.1 (Fast FP8) generates premium quality images post image

Don't know how they did it, but for under 10GB model, the results are astonishing. I am running it on Unsloth Studio. They just released the update, so if you are not seeing the option, I recommend updating your Unsloth Studio. Cheers!

▲
234
+1
17👁
r/LocalLLaMA · u/Another__one · 18d ago
mini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.

Sorry for the pretentious name, I know, I know.. It just contains all the pieces I would like to see a AGI model to have, and I can't stand the temptation. Before throwing rocks at me, please take a glance at the Readme, and I hope it will cover your mood a little bit.

So, first of all it does work and you can see the sample from the whole training run here: https://raw.githubusercontent.com/volotat/mini-AGI/refs/heads/main/runs/sampl…

Here is the scaling law graph I have so far, and it looks very promising:
https://github.com/volotat/mini-AGI/blob/main/assets/scaling.png

The model was built under my deep dissatisfaction so we cannot really train even moderately big models (1B+ scale) on the consumer's hardware. We can inference and fine-tune them for sure, but I would like to have full control over what the model sees over the training run, so it is fully aligned with my interests, not some corporations.

I was thinking about for some time and come up with two interesting ideas I thought worth pursuing: MoE with a lot of experts that gets added and pruned from the model while it trains, where only a small subset of of experts are actually in use at any particular moment + batch 1 training on the single continuous stream of data.

First allows us to be bounded only by the disk space in terms of number of parameters and load and unload experts only when they are needed. The second (if figured out and it turns out to be doable) allows us to get aways with small VRAM capacity because we do not need to store big randomized batches and their respective gradients.

I started brainstorming with Claude and after some time we found an approach that seems to be promising, and low and behold, a few weeks pass and you can see the results yourself.

Obviously, I did use AI in the process of making this project and I am pretty sure it would be completely impossible for me to do something like this without it, so I hope it is more than justified.

The model is still running over the first of 7.8B characters corpus I selected for training, so the weights are not out yet, and it's about a couple weeks of waiting until they are cooked at the current reading speed. And yeah, the model just read continuous interleaved passages from the dataset, each by 32K characters long each as a single stream. Just as you or I would do.

The set up seems to be really simple so you can git clone the project, run it and observe everything for yourself.

Thanks for your attention.

▲
229
-3
27👁
r/LocalLLaMA · u/shniydder · 20d ago
I gave Jev, Laya, finetuned ModernCE and Qwen3.5 the controls to Doom post image

I gave Jev, Laya, a finetuned ModernCE-base-nli and a finetuned Qwen3.5-4B the controls to Doom. Thanks to TypeSafe AI for Jev access.

The video lines up the starts of four separate games with the same seed. After that, each model's actions change its own game and what it sees next. The clip shows one preselected episode per model in each of two scenarios, with action probabilities, kill counts and survival time on screen.

  • Input: A short text description from a deterministic Python adapter reading ViZDoom's visible-object labels, bounding boxes and HUD values (health and ammo).
  • Output: One button action. In Defend the Center, that's left, right or fire. In Health Gathering, it's left, right or forward to look for medkits as health drains.

Each local model and its game ran on a single DGX Spark with an NVIDIA GB10 and 128 GB unified memory. Jev used TypeSafe's hosted API. ViZDoom ran at 320 × 240, with a 35 Hz game clock and a target of five decisions per second. The game kept running while the model replied.

Here are the averages over eight seeds per controller, per scenario, using the clear-scene descriptions shown in the video. Episodes were capped at 30 seconds of game time.

| Model | Mean Kills | Mean survival (s) | Call p50 (ms) | Call p95 (ms) |
| --- | ---: | ---: | ---: | ---: |
| Jev 1.13 | 5.63 | 13.03 | 117.3 | 199.8 |
| Laya English | 1.25 | 11.89 | 16.2 | 17.1 |
| Finetuned ModernCE-base-nli | 1.25 | 11.66 | 7.6 | 8.8 |
| Finetuned Qwen3.5-4B (LoRA) | 3.63 | 11.31 | 146.8 | 150.9 |

Call latency is request-to-response time on 12 shared synthetic Doom scenes, repeated for 48 calls per model. p50 is the median and p95 is the 95th percentile. Local model timings include loopback HTTP on one Spark. Jev includes the hosted API round trip. Applying the action adds controller and game-tick delay. None of these models got extra Doom-specific training for these runs.

Inspired by TypeSafe's Doom demo and experiments shared on Reddit and LinkedIn.

My longer write-up about exploring Jev: https://morethanamachine.com/posts/jev-style-decisions-dgx-spark/

Edit: Table overflow fixes.

▲
229
+4
23👁
r/LocalLLaMA · u/Henrie_the_dreamer · 22d ago
Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash post image

Hey all, Henry from Cactus Compute here, I kinda wanted to share our latest model and get feedback from the family :)

Needle 3 is a small foundation model for automation: you give it the functions your app exposes, it reads a request and returns the calls with every argument filled in, or a typed record if what you gave it was a schema. It runs on the device, with no network in the loop. It is on Hugging Face, on GitHub, on PyPI as cactus-needle, and there is a sandbox that runs it in your browser at cactuscompute.com/needle if you want to poke at it before reading further.

1) Trades general capacity for frontier performance on automation tasks

The thing we decided early was that Needle would not chat. Every turn is a function call, and a request no declared tool can serve comes back as an empty list rather than a guess. That sounds like a limitation, and it is, but it is what let a 121M-parameter model be trained on 360B tokens of structured data and spend all of its capacity on three jobs: tool calls, structured extraction and text embedding.

The architecture follows from the same trade. It is a Simple Attention Network: the dense feed-forward layers are gone, replaced by a Monarch Hadamard MLP with 25.6K parameters per layer instead of 4.7M, and the knowledge a feed-forward layer would normally hold sits in an engram, hashed n-gram tables that are read by gather and cost no arithmetic. 70.8M of the 121M parameters live there, so the full model does the arithmetic of a 50M one: 100 MFLOPs per token against 296 for a transformer of the same shape.

https://i.redd.it/fsnfjojny4qh1.gif

We wrote the intuition up if you want the longer version: Simple Attention Networks and the Hadamard MLP.

2) Beats models 10x its size on tool calls and language-to-device control

On Mobile Actions (961 phone commands, scored on the exact call), the 20-layer model scores 86.0 through the shipped 2-bit binary with the confidence gate on. LFM2.5 1.2B is at 82.4, Qwen3.5 0.8B at 76.0, FunctionGemma 270M at 65.1 and Apple's on-device foundation model at 57.6, all at f16. DeepSeek V4 Flash through its API is at 88.4, which is the line in the chart.

https://i.redd.it/luunq717z4qh1.gif

The part we are most pleased with is not the number but how the calls are made. Every argument is a span of the request: the model writes a short derivation first ('living room' -> room; '30' -> brightness) and then emits the call under a byte-level grammar compiled from your schema, so the JSON always parses and an enum can never leave its set. An optional field with no evidence is omitted, a required one with no evidence withholds the call, and the engine drops a call the request negates or excludes. Ask for two things and you get two calls in order.

https://i.redd.it/227396y5z4qh1.gif

Full table across all six suites (tool calling is exact match, extraction is field F1, Needle through the shipped binary, baselines at f16 under vLLM):

|Model|Params|Mobile Actions|DroidCall|BFCL v4|DSTC8 F1|SNIPS gold F1|SNIPS 7-way F1|
|:-|:-|:-|:-|:-|:-|:-|:-|
|DeepSeek V4 Flash (cloud)|\-|88.4|60.5|77.2|80.0|69.4|66.7|
|Needle3-20L-121M|121M|86.0|47.0|50.2|40.7|30.2|24.7|
|LFM2.5 1.2B|1.2B|82.4|35.5|62.0|48.0|43.0|38.0|
|Needle3-16L-98M|98M|80.7|40.0|41.3|28.5|23.5|19.2|
|Qwen3.5 0.8B|800M|76.0|28.0|56.8|49.0|35.0|34.0|
|LFM2.5 350M|350M|72.8|32.5|59.1|20.0|34.0|29.0|
|LFM2.5 230M|230M|69.3|11.5|46.3|53.0|27.0|22.0|
|FunctionGemma 270M|270M|65.1|16.5|46.6|27.0|29.0|14.0|
|Needle 2|45M|63.5|17.0|\-|\-|\-|\-|
|Apple FM|3.0B|57.6|\-|\-|\-|\-|\-|
|Needle3-8L-52M|52M|36.8|36.5|28.2|15.3|16.6|10.1|
|Needle3-4L-29M|29M|11.7|21.0|19.5|6.9|7.7|4.3|

You can see where it is weaker too: BFCL and the extraction suites are where the bigger baselines pull ahead, and the smaller subnetworks fall off quickly on the general task (more on why that is fine in section 4).

3) Matches 2-3x bigger models on structured JSON extraction

Extraction is not a separate mode. You declare the record as the only tool and pass the passage where the query goes; with one tool declared the grammar admits exactly one call of that name, so the shape is guaranteed rather than requested, and the values are grounded the same way as arguments: a field is filled only from a span of the passage, an optional field with no span comes back as None, and a date whose year appears nowhere in the text is flagged instead of invented.

from pydantic import BaseModel
import needle

class Invoice(BaseModel):
vendor: str
total: float
due_date: str
po_number: str | None = None

needle.extract("Invoice from Acme Corp, $1,200.00, due 2026-09-01", Invoice)
# Invoice(vendor='Acme Corp', total=1200.0, due_date='2026-09-01', po_number=None)

It generalised to classification without special training, because an enum is just a constrained value: declare sentiment: Literal["positive", "neutral", "negative"] on a record and you have a classifier whose output cannot leave the set. A watch reads a notification into merchant, amount and date that way, then into a reply, then into a sentiment flag, one record each. On DSTC8 and the two SNIPS suites the 121M model lands between the 230M and 350M baselines, which is the 2-3x in the heading.

4) Intelligence ladder: every depth from 2 to 20 layers a model of its own

This is the part I would most like your thoughts on. Needle 3 is one set of weights, and every depth from 2 to 20 layers is a deployable model. Blocks 0 and 19 are always kept and the rest are added by bisection, so each subnetwork nests in the next; during training each step samples one path, mostly the full model and otherwise a random depth, with the smaller path distilled from the full one. The full-depth model ends up slightly better than an ordinary run of the same size, and every depth below it is trained rather than truncated.

https://i.redd.it/kbwimep4z4qh1.gif

Why we wanted it: a watch, a Raspberry Pi and a phone do not want the same model, and they want to pick the size at deploy time. needle build --layers 8 writes the 8-layer file; the same engine runs all of them. The small depths lose accuracy on the general benchmarks (that is the bottom of the table above), and they get it back when fine-tuned to one product's tools: on DroidCall every subnetwork gains 18 to 36 points, and from 4 layers (29M parameters) up the tuned subnetwork passes DeepSeek V4 Flash.

https://i.redd.it/y646wed3z4qh1.gif

Fine-tuning is LoRA on the frozen base, merged at export, and the Python package does it locally at 4 bits (needle finetune data.jsonl, then needle build). The maths is in Intelligence Ladders and the workflow in Fine-tuning Needle.

5) Runs locally at up to 4k tokens/sec decode speed

The engine is under 1 MB, plain CPU, no GPU or NPU, and the weights are read in place from a single file the engine maps into memory: a 196-byte header carrying the whole architecture geometry, a nameless tensor directory, and the quantised blobs in the order the forward pass reads them. On a Raspberry Pi 5, decode runs at up to 4k tokens/s at the bottom of the ladder and around 400 at the top, prefill from 10k down to 1k. Every response reports prefill_tps, decode_tps and peak_ram_mb, so you can measure on your own device rather than take our word for it.

Every response also carries a confidence score from a calibrated head, the minimum of a post-hoc judgement on the finished call and the decode probability of its tokens. The engine withholds anything under 0.1; above that the number is yours: act at once when it is high, show the call and ask when it is middling, treat [] as a refusal. How we use it is in Leveraging Needle's confidence.

6) 25-121M deployable parameters at CQ2-bit (8-29MB binaries)

The weights are quantised with Cactus Quants: groups of 128 weights are rotated by a Walsh-Hadamard matrix, which makes every group look Gaussian, split into an fp16 norm and a direction on the unit sphere, and the direction's coordinates are snapped to a 4-entry Lloyd-Max codebook. That is 2.125 bits per weight, and the kernel never expands them: it rotates and int8-quantises the activation instead, then does table lookups and sdot against the packed indices. The embedding and the confidence head keep 4 bits, the norms and gates stay fp16. The byte layout, and a twenty-line parser for it, are in The .cact format.

7) For mobiles, wearables, smart home, small robots and microcontrollers

Every target ships a prebuilt engine folder: macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32 (the Ingenic camera SoCs), Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly, and a WASI component with a WIT world. pip install cactus-needle covers the desktop and server platforms with wheel-tagged engines, and needle build --platform linux-arm64 --layers 8 --out ./pi puts an engine and the weights in a folder you copy over. Inference never touches the network, so an air-gapped device only needs the files in place. The list and the runtime surfaces are in What devices are supported.

import needle

@needle.tool
def get_weather(city: str):
"Get the current weather for a city."
return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]

What we would love feedback on

  • Where the grounding rules get in your way. The engine refuses to invent a number, drops a call the request excludes and withholds a required enum the request never names; we tuned those on our suites and would like to hear where they bite on real tools.
  • The ladder. Whether depth is the right knob for you, or whether you would rather have width, and what devices you would want the 2- and 4-layer models on.
  • Extraction cases we have not seen. Nested records, arrays, multilingual text (it is English-first, and non-English text fragments into about 1.7x more tokens).
  • Anything in the tool design guide that turned out wrong for your schemas.

Thanks for reading this far. Happy to answer anything in the comments.

▲
214
+4
24👁
r/LocalLLaMA · u/Terminator857 · 18d ago
Deepseek training 2T and plans 8T model

Quote: DeepSeek is training a 2T-parameter model and plans to eventually build an 8T-parameter model.

https://x.com/wallstengine/status/2101982843656388644

Current DeepSeek models:

  1. Flash parameter count of 552 billion
  2. Pro: 1.6T (trillion) total parameters with 49B (billion) activated weights per token

Mythos / Fable is estimated to be 10T parameter count.

▲
203
-1
24👁
r/LocalLLaMA · u/Odd_Caterpillar_2994 · 22d ago
Installing 6 GPUs in a standard case rather than using an open-frame chassis. post image

CPU : epyc 7262

MB : ROMED8-2T

GPU : V100 16GB PCIE x6

I have built a server with six V100 GPUs. I am now testing it and plan to eventually run Qwen3.8-Next-Flash configured with TP2 and PP3.

Because open-frame or server-style cases are large and unattractive, I chose to install all six cards in a full-tower case.

▲
205
+2
23👁
r/LocalLLaMA · u/CriticallyCarmelized · 16d ago
Please Google, for the love of God.

Gemma 5, 220B A18B QAT plus ngrams please. Thank you very much!

Will settle for 120B A16B plus ngrams.

▲
204
+2
22👁
r/LocalLLaMA · u/wFXx · 20d ago
Von: Open-source 395M "System One" model

Took me a while since I'm on a family trip and have limited hardware, but here it is!

Von: Open-source "System One" drop-in replacement for TypeSafe's JEV.

https://github.com/wfzyx/von
https://huggingface.co/wfzyx/von-1.0

It runs entirely on a CPU with 1–2 GB of memory (I haven't spent much time optimizing it yet), responds in 25–300 ms, and beats JEV in all benchmarks. Enjoy!

P.S. I’m open to offers to work at AI research labs. Feel free to ping me if you have an offer.
P.P.S. If you have a GPU, it’ll be faster, but a GPU isn't required.

▲
194
-3
22👁
r/LocalLLaMA · u/simpleuserhere · 17d ago
Laya model playing Flappy Bird on a CPU using OpenVINO INT8 inference post image

A 421M-parameter model just played Flappy Bird on my desktop CPU (OpenVINO int8)

Running on my Intel Core i7 12th gen CPU

Converted laya system one model to OpenVINO and quantized to int8

Repo : https://github.com/rupeshs/flappy-laya-openvino-cpu

▲
194
+3
24👁
▲
192
+3
16👁
▲
180
+3
23👁
r/LocalLLaMA · u/NineThreeTilNow · 16d ago
Engram gone wild! 2b model update...

Update from : https://old.reddit.com/r/LocalLLaMA/comments/1wis23s/update\_small\_model\_engram/

It's been about a week so I'm back. People were asking me about the model.

People wanted code, or models, etc. Most of that is useless to you right now because you're not going to use an under trained model. So let's get to the details.

The spec locked to the following after a LOT of testing :

2.6b model all up. Embedding, LM head, AttnRes, etc.

2.2b are trained. Embedding / LM Head are frozen (\~205m each)

4.3b ENGRAM table. Yes. She's chonky.

Architecturally speaking now :

This includes SWA layering 3:1 as per standard ablations have shown is "optimal" These are 4k / 4k / 8k SWA layer follow by the global.

At the end of the first "block" (4 layers) the Engram table appears. The Engram table itself runs with a very small set of attention heads so it's not blindly attempting to inject data. This is context aware Engram. I'm uncertain of exact Qwen / DS methods here. They're fairly close to what I use. The difference is obviously the Engram size.

This has 10 blocks plus a final global layer (41 Layers). + Embed + LM Head

So if we step back, the optimization problem is as follows :

How do we maximize compute in the backbone and offload the boring stuff to a table?

Attention Based Residuals (Moonshot) comes to our rescue here. Why AttnRes? It allows all blocks past 1 to utilize the Engram table in some fashion. They all have attention via the residual stream to determine if they want data from Engram and precisely how much.

Why not a bigger Engram table? Honestly? You could probably do that. However I don't know if any model has attempted this sort of mismatch of compute vs table. In theory, DeepSeek's research says it works.

More depth? Could do that too. Training is expensive though.

Why frozen LM/Embed? These are down projected via SVD from OLMo 3's model. So it's a mathematical compression attempt at \~5k -> 2k. This saves a MASSIVE amount of time.

Currently the data lives in a \~KD format. 32 logits stored from the teacher of the Wikipedia corpus. Instead of training 1 hot, it gets 32 soft targets to try to match. Hard cross entropy is brutal on a model and it's the reason you see "trillions of tokens" quoted.

When you borrow an LM Head / Embed / Tokenizer / Teacher model... This is far less.

\---

Where are we now in training? I passed the 100m token mark yesterday at \~4am Pacific.

The current HF repo has all the checkpoints, data, and the 104m mark safetensor.

\--

I wanted to do some inspection of the model at this point. We have to see that it's not complete garbage right? It has seen a fraction of the data it needs.

So based on ONLY 100m tokens seen, I attempted completions and various ablations. Studies of what EXACTLY Engram stores (because it's mostly a guess).

Let's go straight to completions :

"George Washington was an American"

With Engram - "George Washington was an American naval officer who served in the American Revolutionary War. He was a naval officer who served in"

With Engram Zero'd - "George Washington was an American, but he was not a. He was a very good friend and a. He was"

"The American Civil War was a civil war in the United States from"

With - "The American Civil War was a civil war in the United States from 1861 to 1861. The war was a major victory for the Confederacy..."

Without - "The American Civil War was a civil war in the United States from 1861 to 1862. The war was a major victory for the Confederacy..."

"Aristotle was an Ancient Greek" (probably my favorite)

With - "Aristotle was an Ancient Greek philosopher who was a leading authority in the philosophy of Plato. He was also a leading authority..."

Without - "Aristotle was an Ancient Greek word meaning "to be" (ἀπάς, "to be")..."

\---

What do we learn from direct inspection of Engram?

If it saw the data enough times, it starts to offload it to the tables. At that point the model is less forced to use internal computational space to store data, and can rely on Engram.

That's not some interpretation. That's what the data shows exactly.

In the early 100m tokens of Wikipedia it's HIGHLY biased to early Philosophy. This has to do with the topics of what was IN Wikipedia at that time. Early Wiki contained a lot about Philosophy. It's among the earliest topics to exist.

It has seen TONS of Philosophy so it has basically offloaded to Engram because of the repetition.

Ok this is long enough. I'm calling it quits. I'm still looking to find a provider that will sponsor the model to completion. Right now I'm just paying. I contacted Verda and they never responded. Qubrid is in this sub, and they said they'd help but never emailed back after a few repeated proddings. Massed Compute is where this model is currently training, and I asked them like yesterday. Hopefully they'll help a brother out.

None of the above is "LLM written" except the completions from testing I guess.

As usual, ask whatever. It doesn't matter your understanding level or whatever. I'll sit and respond.

No question too dumb. No insult not insulting enough. (I'm joking. It's Reddit)

Much Love.

▲
178
+3
25👁
r/LocalLLaMA · u/AdRepulsive7837 · 21d ago
still doesn’t get what Jev is…..is it just a more generalised BERT?

Looking at jev launch website and demo video on x.com…. it seems like it’s a very intelligent classifier with custom prompt and custom criteria instruction reading capabilities. It can do well defined narrow and well defined task

Me, following NLP since good old days of word embedding and BERT,,, be like asking….

Isn’t that BERT?

yeah i know BERT need fine tuning to adapt to custom domain, but can Jev be like generalised form of BERT?

▲
161
-4
16👁
r/LocalLLaMA · u/Akainu_Fan · 17d ago
Did Alibaba abandon 35B A3B?

Basically the title.We did not get a new moe model with qwen 3.8 and Alibaba did not announce any small moe models on apsara.I know we might get an announcement later but ngl I kinda lost hope

▲
166
+1
15👁
r/LocalLLaMA · u/ali_byteshape · 21d ago
Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison post image

Hey r/LocalLLaMA,

Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark.

We wanted to see how they compare under the same methodology we use for the rest of our Qwen3.8 evaluations, so we ran Bonsai 2 through our own benchmark suite.

One important clarification: these are our evaluation results, not Prism’s reported benchmark numbers.

We used Prism’s fork/runtime to run their models, while keeping the workloads, benchmark suite, and evaluation methodology consistent with the rest of our comparison.

Our evaluation includes separate Instruct and Thinking benchmarks. For Thinking, we use medium thinking effort with the recommended sampling parameters.

We added Bonsai 2 because the models have become a relevant part of the Qwen3.8 landscape, and we wanted to provide a common reference point for people comparing the available options.

Different providers often report results using different benchmark suites, runtimes, reasoning settings, sampling parameters, and evaluation methodologies, so those numbers are not always directly comparable. Running the models through the same evaluation gives another reference point for understanding the trade-offs between quality, model size, and throughput.

Updated comparison and results: https://byteshape.com/blogs/Qwen3.8-27B/

▲
164
+3
23👁
r/LocalLLaMA · u/Jorlen · 22d ago
Does anyone use uncensored models purely for coding?

It sounds like a stupid question, and I do apologize if it is... but I've seen several people mention that coding models are better uncensored due to the fact that they don't have to constantly run prompts through the "is this okay" sort of checks.

Is this hogwash? Is it true? And more importantly, does anyone have any sources to confirm it?

Anecdotal evidence is fine too if you've tried and compared them.

Personally, I have never bothered because I'm too worried the de-censoring would damage the weights. The juice never felt like it was worth the squeeze... but maybe I was wrong?

Edit: Either I'm unclear or people are misinterpreting my request: Specifically, I mean for every day coding (not for hacking, not for reverse-engineering) but just for regular coding of new apps, etc. The question is: will the uncensored model produce better code faster (without reasoning so much) because it no longer has to worry about "is this alright" when it questions everything...

Edit 2: Decided to test the HuiHui Qwen 3.8 27b (UD-Q8\_K\_XL) quant myself. So far, it reasons far less, and I have yet to have any issues with its coding quality. Granted, I've only been testing it for about 8 hours (straight...) in an active project. Its reasoning is far shorter, it seems far more confident in its responses and as such, uses far less context to achieve the same result. I will continue testing for another week; it's pitted against the Dirk template version of Qwen 3.8 27b right now (same quant) which I'd been using the previous week.

▲
162
+5
23👁
r/LocalLLaMA · u/power97992 · 17d ago
Now Opus 5.5 is 58 on Artificial Analysis , how long do you have wait until an open model hits 58?

It is 12 points higher than the best open model mimo 2.6 pro and a big jump from fable 5.1. Crazy, glm 5.5 and qwen 4 will be on par with gpt 6 sol or better since it has a score of 48

If it took 2 months for the best open model to go from 44 to 46, then at this rate, in 6 months , they will reach 58? It is quite possible they will reach it in 4-5 months, since they have will more leaps in intelligence as they deploy more gpus and scale up the parameters, data and compute and improve the architecture .Wow sol 6 is worse than 5,6 at deepswe?

▲
152
-2
25👁
r/LocalLLaMA · u/KURD_1_STAN · 21d ago
bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai\_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai\_2/q3.5 60.8 / 72.4 = 84.0%
Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release \[2\], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

^(also 3.8 35b qwhen? plsss)

▲
156
+4
16👁
▲
149
 
23👁
r/LocalLLaMA · u/zyxciss · 19d ago
I tested 9 LLMs on the exact same web-dev prompt for ~8 hours — RTX 3060 12GB results (Rate the best!) post image

I’ve spent basically the last 8 hours testing different models on the exact same web-development prompt, and I finally finished.

The whole point of this nine-hour test was that which local model matches the frontier-level intelligence at size and could fit easily in an RTX 3060-like consumer card.

My setup:

  • GPU: RTX 3060 12GB
  • RAM: 16GB DDR4, single-channel
  • OS: CachyOS (Arch Linux)
  • Local models were run through my local llama.cpp setup.
  • Same prompt for every model.
  • I recorded the generations so you can actually judge the websites yourself rather than relying on my description.

Prompt

Build a polished, production-quality single-page website for a fictional high-end technology studio called NOVA//LABS.

Goal: make it look genuinely designed by a strong human frontend developer, NOT like generic AI-generated SaaS UI.

Requirements:



\* Use plain HTML/CSS/JavaScript or React + Tailwind if you strongly prefer it.

\* Everything must run locally with minimal setup.

\* Create the entire project/files yourself.

\* No backend, authentication, database, or unnecessary complexity.

\* Responsive desktop + mobile layout.

\* Strong typography, spacing, hierarchy, subtle motion, and excellent visual composition.

\* Dark, sophisticated visual language with restrained use of gradients/glows.

\* Avoid the typical AI-slop look: no excessive rounded cards, giant gradient blobs, random glassmorphism, meaningless statistics, or generic "Empowering the future" copy.

\* Make the copy specific and believable.

\* Include:

1. A striking hero section with a concise headline.

2. A subtle animated visual representing an abstract computational system.

3. A small selected-work/projects section.

4. A concise capabilities section.

5. A strong closing CTA/footer.

\* Add tasteful interactions such as hover states, scroll reveals, and subtle cursor/mouse effects where they genuinely improve the design.

\* Prioritize visual quality over feature count.

\* Use freely available CDN assets only if genuinely necessary; otherwise create visuals with CSS/SVG.

\* Keep the implementation reasonably small and understandable.



Most importantly: make strong design decisions yourself. Do not explain your design choices before building it. Start by creating the project and finish with the exact commands needed to run it.





(SELF CONTAINED HTML WITH JS AND CSS)

I wanted to see what the models actually build, not just how well they explain code.

The models

1. Gemini 3.8 Flash

\~3 min 12 sec

Used Antigravity and consumed roughly 9K tokens.

This was one of the frontier-model reference points for the test.

2. GPT-5.6 Sol

\~1 min 6 sec

Token usage wasn't available to me.

Extremely fast compared with the local models, so this was another useful frontier reference.

3. Claude Sonnet 5

\~4 min 56 sec

Token usage wasn't available.

Also included as a frontier reference. (I was only able to use Sonnet 5 as my Claude-Code Max subscription had expired)

Local models

4. Bonsai 2 27B Ternary

\~45 minutes

  • Native ternary / \~2-bit model
  • Model size: \~7.66GB
  • Average generation: \~34–36 tok/s
  • Context: up to roughly 102K
  • Used \~52K tokens out of a 122K context during this run
  1. Qwen 3.8 27B GSQ-RCO-IQ3-XXS + MTP

\~57 minutes

  • Model size: \~10.4GB
  • High thinking enabled
  • \~29 tok/s around full context
  • Around 40 tok/s with a much smaller/near-empty context
  • Context used reached roughly 75K
  • Context was compacted twice
  • Available context for this particular run was around 49K after the relevant setup/limits

this was probably the most interesting local result for me.

6. Qwen 3.8 27B Q4_K_M Unsloth Dynamic 3

2+ hours

  • \~16.4GB model
  • Q4\_K\_M
  • High thinking enabled
  • Full-context generation dropped to roughly 4 tok/s
  • Context reached roughly 96K
  • Obviously requires significant CPU/RAM offloading on a 12GB GPU

Flags used : --jinja --reasoning-preserve -fa on -fit off -ngl 99 --override-tensor "blk\.([0-9]|[1-3][0-9]|4[0-5])\.ffn_.*=CPU" -ctk q4_0 -ctv q4_0 --gpu-layers-draft all --spec-type draft-mtp --spec-draft-n-max 2 -lv 4 --no-mmproj -np 1 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --load-mode none --no-warmup -b 256 -ub 128 -c 98304

7. Qwen 3.8 27B GSQ-RCO-IQ3-XXS + MTP — thinking OFF

\~12 minutes

Same general Qwen 3.8 GSQ-RCO model, but this time I disabled thinking.

It used roughly 12K tokens and produced the site dramatically faster.

This was a particularly useful comparison because it shows how much the reasoning mode itself can affect local generation time.

8. Ornith 1 9B Q4_K_M

\~2.4 minutes

  • Model size: \~5.4GB
  • \~74 tok/s
  • Native context: up to 262K
  • This generation only used around 2.6K tokens

This is the speed monster of the local group.

9. Ornith 1.5 35 A3B Q6

\~30 tok/s

  • Model size: \~22.4GB
  • \~30 tok/s
  • Context available for this run: around 128K
  • Obviously heavily dependent on offloading because of the model size

Quick summary

|\#|Model|Approx. time|Local?|Generation speed|
|:-|:-|:-|:-|:-|
|1|Gemini 3.8 Flash|\~3:12|❌|—|
|2|GPT-5.6 Sol|\~1:06|❌|—|
|3|Claude Sonnet 5|\~4:56|❌|—|
|4|Bonsai 2 27B Ternary|\~45 min|✅|\~34–36 tok/s|
|5|Qwen 3.8 27B GSQ-RCO-IQ3-XXS + MTP|\~57 min|✅|\~29–40 tok/s|
|6|Qwen 3.8 27B Q4\_K\_M Unsloth Dynamic 3|2+ hrs|✅|\~4 tok/s at full context 8 tok/s at empty|
|7|Qwen 3.8 27B GSQ-RCO-IQ3-XXS, thinking OFF|\~12 min|✅|—|
|8|Ornith 1 9B Q4\_K\_M|\~2.4 min|✅|\~74 tok/s|
|9|Ornith 1.5 35 A3B Q6|—|✅|\~30 tok/s|

My personal take

For local models specifically, the one that impressed me the most was Qwen 3.8 27B GSQ-RCO-IQ3-XXS.

It hit a pretty interesting balance between:

  • actual design quality
  • coding ability
  • context handling
  • generation speed
  • fitting within a 12GB GPU setup

The Qwen 3.8 27B Q4\_K\_M Unsloth Dynamic 3 was also interesting from a quality perspective, but the speed penalty once you're deep into the context is huge.

Bonsai 2 27B Ternary was also surprisingly usable given that it's a \~7.66GB ternary model.

so its Qwen 3.8 27B Q4\_K\_M > Qwen 3.8 27B GSQ-RCO-IQ3-XXS \> Bonsai 2 27B Ternary

I've attached the screen recording showing the outputs.

Especially interested in other RTX 3060 / 12GB setups ;0

If possible Someone please post down GPT-6-ASTRA's results if they have a codex subscription.

▲
145
-1
23👁
r/LocalLLaMA · u/HFq_Dev · 20d ago
Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. post image

I needed a relatively simple but acceptable level of AI for working on one project. I didn't have any heavy requests, I just needed to give the AI access to the project files so it could search through them for bugs and stuff. I already had an old computer that I decided not to throw away and instead give it a new life as a git server (and sometimes a minecraft server).

The pc specs are ancient by today's standards:

CPU: i7-4790K 4.6 GHz

Motherboard: MSI Z97 Gaming 7

RAM: 32 GB DDR3 2400

PSU: 750 W

Well, my idea stopped at maintaining the computer, because the gpu, a GTX 1070, was overheating. It needed a complete repaste, but the cooler screws were completely stripped, so while trying to remove the cooler I accidentally knocked off several important smd components with a screwdriver. R.I.P. GPU.

Without a GPU, inference was running entirely on the CPU, and only MoE models were kind of usable, giving around 10–20 tokens per second, while dense models couldn't get past 3 tokens per second. I didn't even get to test it with the GTX 1070, because I decided to service it lol.

I started looking for a replacement on the secondary market, but quickly realized that this would cost too much for the minimum entry point I wanted for experimenting with AI. Until my eyes fell on mining cards. There were plenty of cmp40hx, cmp50hx, cmp70hx and cmp90 cards for sale, and the prices were pretty reasonable(it was a month ago), considering that I was originally looking for a cheap replacement for my dead one.

Getting closer to the actual build, I started calculating how much vram I would need for ± acceptable AI with tolerable speed, and after getting inspired by this sub I decided to take a step further and went with a modified cmp50hx with 20gb of memory and pcie modded to 16 lanes. Very quickly after that I bought another one, this time unmodified (10 gb, only 4 pcie lanes). So, together that's 30gb of vram. Both cards cost me $250 in total (it was also a month ago, right now they've suddenly doubled the price).

Luckily for me, around the same time new driver patches appeared that almost completely remove the limits on their compute performance, and even add pcie 2.0 support (these things have pcie 1.1).

I initially tried one of the newly released cmp50hx driver patches, but it ended badly and I had to spend a lot of time trying to get the drivers working. The patches were new and didn't account for the 20gb version. Later the author fixed that, but even then the driver didn't work for me because of some other problem that I don't want to get into.

I went digging through the driver's github issues and quickly found a guide posted there by another user.

And yes, now the drivers work, the cards are detected and even pcie 2.0 works, but not without problems. The author of the guide said that pcie 2.0 support was only confirmed on the X99 chipset. Well, it also works on Z97, however after waking the computer from sleep the driver crashes completely. After looking into it a bit, I quickly came to the conclusion that the problem was specifically with the pcie patch. Disabling sleep completely solves the only problem I had while using them. :)

Without the specially patched drivers, Qwen 3.6 27B did around 15–20 tokens per second without mtp. MoE models were faster, giving 45–50 tokens per second. With the new drivers, performance doubled. With Qwen 3.8 27B(mtp on) I get around 30–35 tokens per second now, and prompt processing is around 300–400 tokens (including degradation as the token count increases). Ornith 1.5 35BA3B(Heretic-MTP-APEX-I-Balanced) gives around 80–100 tokens per second. I capped GPUs at 180w power limit due to the psu I have, I don't want it to work at its limits, but running them at their 225w surely boosts speeds.

In general, I ended up making a lot of presets for different quantizations with different quality and KV cache sizes (I still need to test all of this in real work), but if we take the better options, I managed to get a Q6K model with 130k context (K – Q8\_0 and V – Q5\_1).

I also followed a guide for running 27B Qwen with large context on limited vram. Using the same general approach, I managed to get 256k context with K and V Q8\_0. The speed is slower though, around 10–14 tokens per second, and prompt processing is around 40–50 tokens. Maybe I can tune it even further. I needed this preset for tasks that I can leave generating overnight :3

Overall, I'm satisfied with the result.

Now the main problems I ran into, not counting the drivers:

  1. Not enough VRAM. The ideal option would be having 2 identical cards with the same amount of memory. You can run dense models in tensor mode, split the weights evenly and get increased generation speed for basically free + fit the full context. I tried many variations of Qwen 3.8 27B quantization, but the only one I could properly run in 1,1 tensor split was Q4KM (the Unsloth one) together with mtp = \~40 tokens per second. However, there is critically little space left for the KV cache, because it gets distributed together with the model weights, and the second 10gb card simply became the bottleneck. Without mtp, running models with a 1,1 split basically loses its purpose. Pcie 2.0 and the number of lanes probably also play a significant role here. That could in principle be solved by adding more pcie lanes to the second GPU and perhaps buying an nvlink cable(who even does that?), but I decided it wasn't worth it just to get another 5–6 tokens per second.
  2. No NVMe SSD. Yeah, all models are loaded from a sata ssd so the speed is around 500mb. It's terrible. The motherboard actually has an m2 sata slot with a pcie 2.0 x2 interface, but even its 1gb per second would be too slow for fast model loading. This could be solved by installing an expansion card into one of the pcie slots (there is one free pcie 3.0 x4 slot), but the current price of those things including the ssd is too high, considering that I'm building a cheap system from what I already have with minimal additional spending for an acceptable result. Model switching takes 1–2 minutes. But whatever. (Not whatever, i'm buying a cheap used 256gb nvme ssd :D)
  3. Amount of RAM and Linux (Ubuntu Server 24.04). Apart from AI, I also run gitlab on the server. And here is the problem: after loading a model, all available ram gets cached by the system for the model files. I'm talking about file cache, not KV. The system was leaving around 200–300mb of free ram for everything else. As a result, openwebui and gitlab started acting laggy (after the model was loaded), as well as the kde plasma interface I installed. I don't completely understand why linux decided to keep this cache until the very last moment instead of freeing it for other programs. I tried adding the no-mmap parameter to the model presets, but nothing helped, and I had to manually clear the cache after loading models, which obviously wasn't acceptable. Together with chatgpt (who else?), I made a command that launched the model and then cleared the cache. It turned out that this broke llama-server, causing model switching to stop unloading the previously loaded model.
  4. Llama-server flexibility. The list of presets is defined in the models.ini file, where each parameter is in key-value format. For running Q6K with 256k context according to the guide, I needed to set GGML\_CUDA\_DISABLE\_GRAPHS=1, which applies to the entire cuda environment and remains active even after unloading the model. That's undesirable, because with it enabled I lose 1–2 tokens per second on other presets.

So for one specific preset I need to enable cuda graphs, while for the other presets I need to disable them. The llama-server parser does not support things like this, and doing it manually is not an option either.

Together with the other problem with ram cache getting stuck, this led me to making an alternative way to launch the models and proxy requests to llama-server.

To solve problems 3 and 4, I made a launcher (well, chatgpt did, because I'm not a server/python specialist) that proxies requests to llama-server but takes over the functionality of collecting model presets from .sh files and launching them. It also clears ram page cache after loading a model.

In case someone needs that launcher, I can leave it in the comments, along with any other links to the drivers, fixes, build params, etc. Just ask. (Reddit removes the post when I include them, not enough karma, I guess.)

Overall, I’m pretty happy with how this setup turned out. The performance is much better than I expected from these cards, especially considering how cheap they were. I had a hard time getting the patched drivers to work and linux didn't make my life any easier, and sometimes I even regretted buying these GPUs, but in the end, it was worth it. Qwen 3.8 27B really works like Opus 4.5

▲
145
+4
24👁
r/LocalLLaMA · u/monkeyofscience · 15d ago
Open AI in the wild

You degenerates are famous. I was at a conference with one of the authors, and they asked me: "Did we do a good job representing the community?"

https://dl.acm.org/doi/abs/10.1145/3805689.3812421

▲
136
-4
23👁
r/LocalLLaMA · u/ababaka · 18d ago
Mimo v2.6-Flash-RL vs open-weight models post image

Since there’s no comparison chart on the model page, I asked Perplexity to compare it against some relatively small open-weight models in a similar size range. Here are the results.

Upd. Terminal-Bench 4.0 results:
MiMo‑V2.6‑Flash‑RL — 28.8%
DeepSeek‑V4‑Flash‑0731 — 12.0%
Qwen3.8‑Flash‑Next — 25.3%
GLM‑5.3‑Flash — 32.8%

▲
143
+3
21👁
r/LocalLLaMA · u/No_Issue_8224 · 21d ago
MiniMax Code goes open source

MiniMax has open-sourced the terminal version of MiniMax Code:

https://github.com/MiniMax-AI/minimax-code

How can developers verify the content that encoding proxies read, send, and store? This is a topic that has been widely discussed recently.

Open sourcing the agent doesn’t automatically answer every privacy or security question, but it gives the community something concrete to inspect.

The repository includes:

  • interactive TUI and headless execution
  • code editing, shell commands, diffs, and test verification
  • permission controls and sandboxing
  • Plan Mode and resumable sessions
  • subagents, plugins, skills, and MCP
  • BYOK with OpenAI- and Anthropic-compatible providers
  • ACP support for compatible editors and clients

First-party code defaults to the MIT license.

A few important caveats: this is a 0.4.12 source preview, the desktop app source is not included, and—as the repository itself notes—a matching version number does not prove identical build provenance between the published package and source checkout.

Still, releasing the agent layer is a meaningful step toward auditability. I’d like to see the community examine its network behavior, file-access boundaries, telemetry, and reproducible-build story next.

https://preview.redd.it/zvrakmgejaqh1.png?width=1198&format=png&auto=…

https://preview.redd.it/13pjdngejaqh1.png?width=1206&format=png&auto=…

https://preview.redd.it/z1ekjlgejaqh1.png?width=1200&format=png&auto=…

▲
135
-4
20👁
r/LocalLLaMA · u/NineThreeTilNow · 22d ago
Update : Small model + Engram

I posted something about a 9b model a few days ago.

The real problem was 2 fold. Someone suggested the size was too big to prove it all out. It's a fair argument. I needed the model depth though. The second problem was the "ability" of these models and the fact that labs (with money) produce these models still. No real usefulness in my mind. A 2b model? maybe interesting. Want a chatbot? - this could be it.

Further, the Llama licensing at the core of the model was highly problematic and had to be thrown in the trash. It's vastly too restrictive and I couldn't Apache 2.0 anything.

So I moved to using the OLMo tokenizer. Except I shrunk the d\_model down to 2048 so I could build a tiny 2b model.

The size of the vocab with 1/2/3 gram forces the Engram table table to a full 1b. That's 50% where DS suggested like \~10-20%.

So basically 2b model + 1b Engram.

The model, because of the depth now allowed by 2048 d\_model allows me to push 40 total SWA/Global blocks to take advantage of attention based residual streams (Moonshot Kimi K3 style) in a dense model.

Data, is pulled from standard Wiki style data sources in my initial training. The data is pumped through the 7b OLMo model to generate the probabilities. This tells the model that the next token is a probability of 32 different tokens.

The big surprise was the results. Even after ONLY 15m tokens, the model is surprisingly coherent. Had I given pure 1 hot cross entropy, 15m tokens would do nothing. I'd get gibberish. However, I borrowed embeddings, and LM\_head that was down projected from 5k -> 2048 d\_model. This preserves \~65% of the data the "big" model had in the embedding when spectrum analysis is done.

The training is limited at 15m tokens, so it's like... babbling about Roman war history and stuff that exists in that token subset. It doesn't fall in to the classic early model phase where it repeats a token over and over though. That's what pretraining 15m generally gives you at first.

Anyways, this was an update and progress report for anyone who cares about this crap... I'm still processing all of the Wikipedia chunk from HuggingFace and working to train it. It'll take time.

If you read this far, thank you. If you have questions, I'm happy to reply. Someone suggested I was a kook who didn't understand ML previously. I started in ML some 20 years ago and held a brief (6 month) stint at Anthropic red teaming the original Opus 4 model before their "Constitutional" paper came out. I'm vaguely referenced in the paper as a "red teamer" I guess. I do this mostly because I love it.

I figured if I built an Apache 2.0 model, I'd at least want people to play with it. It appears 100% trainable on 24gb of VRAM. Training code / Model / Data will follow eventually. It's on HF / Github for now.

The old post is here :

https://old.reddit.com/r/LocalLLaMA/comments/1wezm58/is\_there\_still\_strong\_interest\_in\_a\_dense\_9b\_model/

edit;

Update can be found here :

https://www.reddit.com/r/LocalLLaMA/comments/1wnzu2f/engram_gone_wild_2b_mode…

▲
121
-3
20👁
r/LocalLLaMA · u/jacek2023 · 21d ago
inclusionAI/Realtime-Venus · Hugging Face

do you want some omni? here is omni for you

[](https://huggingface.co/inclusionAI/Realtime-Venus#1-🧭-overview)1. 🧭 Overview

This repository hosts two checkpoints of the Realtime-Venus system:

  • Realtime-Venus-Omni (Realtime-Venus-Omni/): the 9B audio-visual interaction model. It continuously watches and listens, decides whether and when to respond, and generates text and speech on a shared causal timeline. Adapted from MiniCPM-o 4.5, it supports proactive interaction, semantic interruption handling, and training-free long-video memory.
  • Realtime-Venus-Audio (Realtime-Venus-Audio/): the audio-focused checkpoint on the same streaming backbone, for audio understanding and audio-driven conversation with text or speech output.

Both directories contain model weights and custom Hugging Face Transformers code. The asynchronous Realtime-Venus-Harness and its external tool integrations live in the GitHub repository.

[](https://huggingface.co/inclusionAI/Realtime-Venus#2-✨-highlights)2. ✨ Highlights

  • Native full-duplex conversation: keeps perceiving while speaking and distinguishes backchannels, interruptions, corrections, and redirections.
  • Omni-Proactive interaction: continuously processes temporally aligned video and audio, and initiates a response when an event warrants it — without waiting for a user prompt.
  • Delegation: emits in-stream <delegate> requests on the shared causal timeline and consumes asynchronous backend results the same way, so external tasks never block the ongoing conversation. (Executing requests requires the Realtime-Venus-Harness runtime, available in the GitHub repository.)
  • Training-free long-video Memory: archives visually informative moments, retrieves query-relevant and non-redundant evidence, and reassembles the corresponding audio-visual context — no additional training required.
  • Text and speech output: generates response text together with native speech through the bundled Token2wav resources and a reference voice.
▲
126
+3
24👁
r/LocalLLaMA · u/niacolhealth · 17d ago
AntLing open sourced the Ming-Image-0.1-Design family

AntLing open sourced the Ming-Image-0.1-Design family:
• Ming-Image-0.1-Design, 6B
• Ming-Image-0.1-Design-Layer, 6B
• Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill

Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard

▲
123
+2
23👁
r/LocalLLaMA · u/peonist-ai · 20d ago
Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0) post image

Hi.

I saw some feedback that halogen was degrading at context depth. So I fixed that.

Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0:

  • decode at 1,004,581 tokens of context: 27.3 to 38.3 tok/s (default speculative drafter)
  • decode at 258,794: 42.9 to 45.0
  • prefill at 1,004,581: 790 to 937 tok/s, 21.2 to 17.9 minutes cold
  • prefill at 258,794: 1,086 to 1,114 tok/s

Conditions: Ryzen AI Max+ 395, 128 GB. The 262k and 1M rows are one cold request each at the 1M configuration (HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576), greedy, 64 tokens, the rates the response's \timings\ report. The 32k row is the standard ten-prompt served mean and did not change. A follow-up turn over the prompt cache at 1M reaches its first token in about 0.55 s; the numbers above are the cold path.

To run it at 1M: add -e HALOGEN_ROPE_YARN=4 -e HALOGEN_CTX=1048576to the README's podman line; it needs the 128 GB box. Release notes and the full table:

https://github.com/peonist-ai/halogen-flash-server

If you have a 1M sweep of your own, I would like to see it rerun on 0.12.0.

Thanks for all your support, especially https://huggingface.co/nightvich

▲
125
+4
17👁
▲
124
+4
25👁
r/LocalLLaMA · u/Fancy_Fanqi77 · 20d ago
Steer LLMs and Agents at the Token Level: An interactive tool for token visualization & control, model inspection and data annotation. post image

onPanda is designed for geeks, power users, curious minds, and engineers. Its UI is built for deep exploration and efficient data annotation.

\- The core loop is simple: hover over a token → click an alternative or edit freely → continue generation. You can edit every part of model output exposed by onPanda, including reasoning and tool calls.

\- Edit prompts directly, branch tool calls, and use a tree structure to record branch history. This makes onPanda useful for model inspection and prompt engineering.

\- Support multiple modalities, including images, video, and audio; use tool calls and connect MCP servers to perform tasks in real environments.

\- Connect popular harnesses such as Claude Code, Codex, and OpenCode to execute tasks. Explore and compare their tool sets, system prompts, skills, and memory mechanisms.

\- onPanda includes browser-agent, an agent that runs in the user's browser without installation. It uses the browser as its harness and provides JavaScript execution, information retrieval, interface interaction, multimedia I/O, local file access, and persistent memory.

\- onPanda stands for on-Policy Alignment Data Annotator.

I have been building onPanda since 2024.09, it took two years for it to gradually enrich its functionality and ease of use. In my opinion, onPanda is very suitable for the r/LocalLLaMA community. Any feedback and evaluation are welcome.

Try it online (works on mobile): https://onpanda.diyer22.com/

GitHub repo for self-hosting: https://github.com/on-panda/on-panda

▲
119
+4
14👁
▲
115
+5
20👁
r/LocalLLaMA · u/lkarlslund · 19d ago
laya.cpp: Optimized laya near-instant decision making

After seeing u/Nandakishor_ml’s post introducing Laya, I wanted to see how fast it could run in a standalone C++ implementation.

Credit to u/Nandakishor_ml for the architecture, training and open-source release. My contribution is the inference implementation: laya.cpp, built on ggml with custom CUDA kernels.

It supports all three checkpoints—English, multilingual and typed-decisions—with native tokenization, model execution and output formatting. There’s also an HTTP server with a JEV-compatible endpoint. No Python or PyTorch is required for inference.

Some English-model results on an RTX PRO 6000 Blackwell, capped at 450 W:

| Batch | Python BF16 | C++ BF16 | Python FP32 | C++ FP32 |
|---|---:|---:|---:|---:|
| 1 | 149 | 366 | 148 | 342 |
| 2 | 268 | 586 | 202 | 421 |
| 4 | 460 | 761 | 233 | 437 |
| 8 | 663 | 810 | 232 | 386 |

These are questions per second over a fixed 250-question corpus containing choices, scores and booleans. Each precision has paired Python/C++ timings with alternating execution order. Loading and JSON transport are excluded. The README has the full three-model results.

Most of the optimization came from removing unnecessary conversions and copies, fusing operations while preserving rounding, and improving attention memory access.

The code is MIT-licensed. BF16 currently needs the documented CUDA 13.0/cuBLAS 13.1.0 build profile.

Implemented using Codex Astra.

▲
105
-4
22👁
r/LocalLLaMA · u/MeinDruckerSpinnt · 21d ago
JEV architecture

My understanding so far:

  1. You take an LLM and use it without thinking (That's what openjev does?)
  2. You leave out the text generation in the end and take the confidence score in the matrix before that phase

That's it. Right?

They gave it a mysterious marketing name.

▲
108
+2
14👁
r/LocalLLaMA · u/CharlesStross · 22d ago
600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"? post image

It's not the brightest bulb but it's my new drudgework model for read+find or code tasks I'm willing to let it brute force. Even if it takes 20x more tokens, that's still faster than many local models. Not quite Cerebras but still pretty fun to drive.

▲
100
-3
15👁
r/LocalLLaMA · u/jacek2023 · 19d ago
CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp

Another day, another Qwen Flash Next speedup

▲
98
-3
22👁
r/LocalLLaMA · u/futterneid · 16d ago
Streaming Nemotron 3 Diarization post image

I’ve been playing with Nemotron 3 Diarization, and it fills a gap I’ve had with local voice agents: keeping track of who is speaking.

It’s a diarization model, so it gives you speaker labels rather than transcriptions or people’s names. It can stream its output and track up to eight speakers. I’ve been trying it with one-second streaming chunks, and the quality has been really good in my tests.

I plugged it into my speech-to-speech setup on a DGX Spark and connected it to a Reachy Mini. The fun part is watching someone new speak, then seeing the robot ask their name and remember it for the conversation (that's the video!).

It has day-zero Transformers integration and is getting a commercial-friendly license.

▲
95
-2
13👁
▲
97
+2
26👁
r/LocalLLaMA · u/jacek2023 · 16d ago
apple/LensVLM-9B · Hugging Face

https://huggingface.co/bartowski/LensVLM-9B-GGUF

[](https://huggingface.co/apple/LensVLM-9B#lensvlm-9b)LensVLM-9B

LensVLM is a 9B Vision Language Model (VLM) that scans compressed images of text, then selectively expands only the relevant pages to their uncompressed form via learned tools.

[](https://huggingface.co/apple/LensVLM-9B#license)License

All ML model files in this repository, including Apple's modifications to the Qwen model, are provided under the terms of the Apple Machine Learning Research Model License.

The source code that accompanies this model is distributed separately and is provided under the terms of the Apple Sample Code License.

▲
99
+4
22👁
r/LocalLLaMA · u/Zulfiqaar · 22d ago
IFM/K2-Horizon-7B-Uno · Hugging Face - 5200tps with no quality loss

IFM released K2-Horizon-7B , diffusion augmented LLM at upto 5200 tps, with claimed lossless speedup.

causal LLM architecture and adds a plug-and-play diffusion adapter alongside the autoregressive weights.

https://arxiv.org/abs/2609.04010

▲
91
+1
14👁
r/LocalLLaMA · u/Ok-Breadfruit-3523 · 15d ago
My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context post image

These boards cost me $115 each and I have them connected using llama.cpp with Vulkan and RPC on Bazzite. The boards have roughly 27GB of combined GPU memory and communicate over 1gb Ethernet. For around $300 including psu I’m loving the performance. I have a few more and want to see what 6 looks like trying to run qwen 3.8 flash.

▲
92
+2
31👁
r/LocalLLaMA · u/zyxciss · 17d ago
Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW post image

Someone recommended that I try the ByteShape Qwen 3.8 27B IQ3-XXS GGUF after seeing my previous testing of the GSQ quant.

So I did.

And the result was… surprisingly bad.

For context, I'm running:

  • RTX 3060 12GB
  • 16GB DDR4 RAM, single channel
  • CachyOS / Arch Linux
  • llama.cpp
  • Qwen 3.8 27B
  • MTP/speculative decoding where applicable

The two low-bit quants I compared were:

ISTA-DASLab / GSQ-RCO-IQ3-XXS

  • \~10.4GB
  • roughly 2.5 BPW territory
  • MTP enabled
  • \~29 tok/s around full context
  • \~34–40 tok/s at lower context
  • This was the quant I had already been using in my previous web-development test.

ByteShape IQ3-XXS

  • roughly 500MB smaller
  • also around the same ultra-low-bit range
  • advertised as having extremely high similarity to the BF16 model based on KL-divergence measurements

On paper, the ByteShape quant looked very interesting.

It was smaller, while apparently retaining extremely high similarity to the original BF16 model. It was also being compared in size to significantly higher-BPW quants.

So naturally I expected it to at least be competitive with the GSQ version.

It wasn't.

The actual result is shown Above

ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF was able to generate a 3D voxel diorama in one shot under 55k tokens, and Byteshape's 3.8 27B model, took roughly three shots and still hasn't completed with over 98K tokens spent already.

Same with web development not impressive as advertised in Here

Any New Model Suggestions for RTX 3060?

▲
92
+2
24👁
r/LocalLLaMA · u/Odd_Caterpillar_2994 · 19d ago
Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3) post image

Hey everyone,

After wrestling with hardware and engine issues for days, I finally got Qwen 3.8 Next running properly on my multi-GPU rig. Thought I’d share the setup journey, benchmarks, and thermal results for anyone trying something similar.

Seeing all the ongoing memes on Reddit about multi-GPU setups turning into absolute space heaters and catching fire, I decided to run some rigorous thermal tests to see for myself.

he Troubleshooting Odyssey

1. PCIe Link Speed Issue: Right after installation, one of the cards dropped to PCIe Gen 1 x16. Spent about 8 hours over two days diagnosing and fixing it.
2. Finding the Right Engine:
* Started with sglang-v100, but kept hitting continuous OOM crashes.
* Someone on Reddit previously suggested the pxa engine, but that threw errors as well.
* Eventually tried 1cat-vllm, spent some time tweaking it, and finally hit a stable run!
3. Configuration:

  • Running with TP2 PP3.
  • Currently, speculative decoding is limited to speculative=1. Setting it to 2 throws an OOM due to memory constraints (might look into optimizing this later, but for now, it works).

Context & Memory Stats

Plaintext

INFO: Available KV cache memory: 8.78 GiB
INFO: GPU KV cache size: 531,288 tokens
INFO: Maximum concurrency for 262,144 tokens per request: 2.03x

Performance Benchmarks

1. Prompt Processing (Prefill)

|Input Length (Tokens)|Speed (tok/s)|
|:-|:-|
|1,024 (1K)|1,389|
|2,048 (2K)|2,536|
|4,096 (4K)|3,210|
|8,192 (8K)|4,336|
|16,384 (16K)|4,679|
|32,768 (32K)|4,470|
|65,536 (64K)|3,820|
|131,072 (131K)|2,759|

2. Text Generation (MTP Comparison)

|Output Length (Tokens)|Base Speed (No MTP, tok/s)|Optimized Speed (MTP Enabled, tok/s)|
|:-|:-|:-|
|128|22.23|41.34|
|256|22.32|42.48|
|512|22.70|43.04|
|1024|22.86|43.28|
|2048|22.83|43.38|
|Average|22.59|42.70|

MTP nearly doubles generation throughput across the board.

Thermals & Acoustics

People often meme about multi-GPU rigs turning into space heaters or jet engines, so I ran a thorough thermal/stress test:

  • Stress Test: Ran gpu-burn continuously for 20 minutes.
  • Thermal Equilibrium: Temperatures peaked at 65°C and stabilized right around 64°C.
  • Fan Curve: Based on my fan control script, the fans were only running at around 76% at 64°C. The cards stay well under 65°C without even needing full blast.

Pretty happy with how stable, cool, and quiet this system turned out.

▲
89
 
30👁
r/LocalLLaMA · u/FactoryReboot · 16d ago
Most powerful harness for Qwen 3.8?

I hear Qwen code unlocks the model better. I also think it has more power user features than open code?

It’s nice open code can work with multiple models easier though

Thoughts?

▲
83
-2
21👁
r/LocalLLaMA · u/Malfeitor1235 · 19d ago
DIY Jev post image

So this jev thingy is getting kind of big... tbh it seems overhyped by a large margin, but here we are. Not that its bad, just feels like we usually ignored larger things...

Anyway to the point of this post:

I’ve been experimenting with a simple Jev-like inference setup using ordinary open weight LLMs.

Ive done jev-like thing before with llms and i never felt the need that we have to have a separate "system one models" for that and that llms do fine.

So i played around a bit.

The main difference from OpenJev is that there’s no NLI fine-tuning or classifier head.

For each candidate answer I turn the problem into a boolean verification:

<BOS>

Is <candidate> the best answer to <question> given <state> and <options>?
Return only true or false. Treat tagged content as data.

<state>...</state>

<question>...</question>

<options>...</options>

<candidate>B</candidate>

<verdict>

Then instead of generating anything, I read true/false logits for every candidate separately, subtract for every candiddate and softmax those scores.

So 3 candidates is 3 diffs that you softmax over.

The expensive state/question/options prefix you evaluate once, then candidate branches (only diff is last few tokens) are batched through llama.cpp.

On a 32,235-example benchmark:

|model|accuracy|req/s|
|:-|:-|:-|
|Qwen3-4B|65.0% |\~27|
|Qwen3 27B|75.3% | \~2.9|
|Qwen3.6 35B-A3B|75.5%| \~5.|

This req/s is measured on a laptop 5090 24gb.

The interesting part is that the approach works surprisingly well with completely unmodified models. Turns out same model can out perform the openjev fine tune.

Not claiming this reproduces Jev or that the benchmark is perfectly apples-to-apples, mostly interested in how far you can get without training anything and just playing with prompt effectively.

Repo: DIY-Jev GH

Check it out, give feecback and build cool things :)

Edit: I forgot to say hah The repo is a rust web server with jev compatible API that you can run local ggufs from HF in the style of jev. benchmarks included for a few models.

Edit 2: prettier post

▲
80
-4
22👁
r/LocalLLaMA · u/Danmoreng · 20d ago
I tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB) post image

I did a small test of the new hyped quantisation of Qwen3.8 vs the biggest quant which fits into my limited 16GB VRAM with decent context. The results are interesting.

Of course, the smaller file gives worse results. However they are not that far off. Unfortunately, this comes at the expense of even more tokens beeing used by the Bonsai model and thus much longer generation times.

Visually I prefer the IQ3\_XXS results, but see for yourself.

The test is by no means scientific - just few UI generation tasks for direct comparison on the same hardware. Also, I ran llama.cpp with MTP while the Bonsai model doesn't seem to have MTP which makes it even slower.

[](https://github.com/Danmoreng/qwen3-8-27b-iq3-xxs-vs-bonsai/blob/main/RESULTS.…)

|Metric|Qwen IQ3|Bonsai PQ2|
|:-|:-|:-|
|Tasks completed|4/4|4/4|
|Fixed assertions|20/20|20/20|
|Agent wall time|8:00|24:09|
|Output tokens|27,197|84,176|
|Weighted decode|83.59 tok/s|64.91 tok/s|
|Speculative acceptance|65.22% MTP|39.76% modified N-gram|
|Compactions|0|0|
|Length stops|0|1|

Across the complete suite, Qwen finished 3.02× faster and used 3.10× fewer output tokens.

Results:

https://danmoreng.github.io/qwen3-8-27b-iq3-xxs-vs-bonsai/

Repo:

https://github.com/Danmoreng/qwen3-8-27b-iq3-xxs-vs-bonsai

▲
83
+1
19👁
r/LocalLLaMA · u/ciprianveg · 19d ago
Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak. post image

&#x200B;

​I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster.

​Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but throughput and concurrency on coding tasks have been impressive.

​Performance Benchmarks

​Coding Generation / Decode: Sustaining \~30 tok/s (peaking around 38 tok/s) during heavy code generation and agentic tasks.

​Prefill Throughput: \~750–910 tok/s (optimized via modified NCCL topology and dual-switch setup).

​Concurrency & Stress Testing: Handling multiple concurrent user requests smoothly without dropping token generation rates or starving KV cache memory.

​Context / Tool Bench: Stable multi-hundred-thousand token context runs agentic workfows with multiple 500k compaction.

​Compute: 16x GB10 Cluster Nodes

​Connectivity: Dual MikroTik Switch (CRS804-4DDQ) using 4x 400G-to-4x100G breakout cables.

​Runtime: Customized gb10-vllm stack using dspark / Inferact/Kimi-K3-DSpark wrappers with custom MLA/KV kernels.

​I attached a short clip showing real-time token streaming, coding output.

​GitHub & Setup Files:

All runtime patches, config files, and build scripts are on my GitHub:

👉 https://github.com/ciprianveg/gb10-vllm

▲
84
+3
20👁
▲
76
-4
30👁
r/LocalLLaMA · u/satnl · 22d ago
Qwen 3.8 27b is a amazing model, for the first time I see a local model found its own away to open a browser and test

I was testing this quantization IQ3\_XXS from GSQ-RCO with PI.

It is a heavy quantization case, the model is in IQ3\_XXS and KV cache in (Q4\_0, Q4\_0).

I asked it to make the flight simulator, using that popular prompt. For my surprise, when I went verify the session I saw some screenshots, but I don't have mcp or tools in the PI, it's an almost clean installation, only with some personal skills that dont have any mentions to browser testing.

Investigating the session I have noted that it found its way to open a browser for test by itself. My daily harness is opencode, and for not spending the 7k tokens with the chrome mcp in each session, I let the mcp disabled and I only enable it if the session needs. But this experience make think that I probably only need line in the agents md saying that he can use the headless browser with CDP when it needs.

https://preview.redd.it/wrgltip9rzph1.png?width=1952&format=png&auto=…

OK, let me write the file. Also test: I can't easily run a browser here. But I can at least do a syntax check of the JS via node (extract
the script? maybe parse with node --check).

----------------------

Now, final verification: run node --check again, and additionally do a simple DOM-less smoke test? Can't run three.js without a browser.
Syntax check + math tests are enough. Could also try opening it in a headless browser if available? Check whether the machine has
headless Chrome or playwright? That might be overkill; but if I could do an actual rendering check, I could catch runtime errors (e.g.,
shader compilation). Let me check whether headless Chrome is available.

$ cd /tmp && awk '/<script type="module">/{f=1;next} /<\/script>/{f=0} f'
/Users/[...]/zeroshot/Qwen3.8-27B-UD-IQ3_XXS_GSQ-RCO-pi-xhigh/index.html > sim.mjs && node --check sim.mjs
&& echo "SYNTAX OK"

SYNTAX OK

Took 0.1s

$ ls /Applications/ | grep -i -E 'chrome|chromium' ; which chromium chrome google-chrome 2>/dev/null; ls ~/Library/Caches/ms-playwright
2>/dev/null | head

... (5 earlier lines, ctrl+o to expand)
ffmpeg-1011
firefox-1532
firefox-1538
webkit-2311
webkit-2336

Took 0.0s

Chrome is available, and playwright's chromium headless shell is also available. Let's do a headless smoke test: load the page, wait a
few seconds, and capture console errors and screenshots. The page loads three.js from a CDN — the network needs to be working. Let's use
playwright's chromium (does npx playwright exist? The fact that the ms-playwright cache exists means playwright is installed somewhere).
Let me check if the playwright npm package is available globally, or use chrome directly with --headless.

The simplest approach: use Chrome headless directly:
/Applications/Google Chrome.app/Contents/MacOS/Google Chrome --headless --disable-gpu --screenshot=/tmp/shot.png --window-size=1600,900
--virtual-time-budget=8000 file://...

However, WebGL in headless Chrome: it works with SwiftShader (software GL) via --use-gl=swiftshader or --enable-unsafe-swiftshader. In
newer versions of Chrome, the software fallback requires --enable-unsafe-swiftshader. Let's try it. Also, console errors:
--enable-logging=stderr --v=1 will give us the console.

Let me write a small node script using playwright if available. First, check the availability of node playwright.

Chrome is available — let me do a headless smoke test to catch runtime/shader errors and grab a screenshot:

$ node -e "try{require.resolve('playwright');console.log('playwright OK')}catch(e){console.log('no playwright')}"; ls
~/Library/Caches/ms-playwright/chromium-1234/ 2>/dev/null | head -3

no playwright
chrome-mac-arm64
DEPENDENCIES_VALIDATED
INSTALLATION_COMPLETE

Took 0.1s

There's no playwright module, but the Chromium headless shell exists. Either find the path to the headless shell binary and run it
directly, or use Google Chrome in headless mode with SwiftShader.

The path for Playwright's Chromium headless shell:
~/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-mac-arm64/chrome-headless-shell? Let me check the structure.


$ find ~/Library/Caches/ms-playwright/chromium_headless_shell-1234 -name '*headless*' -o -name 'chrome*' | head; find
~/Library/Caches/ms-playwright/chromium-1234 -name 'Chromium' -o -name 'chrome' | head

... (3 earlier lines, ctrl+o to expand)
s.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/chrome-headless-shell
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/headless_lib_strings.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/headless_lib_data.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/LICENSE.headless_shell

Took 0.1s

Use chrome-headless-shell — but to catch console errors and drive key input, we need CDP (Chrome DevTools Protocol).
chrome-headless-shell supports --remote-debugging-port. I could also write a small Node script that speaks raw CDP over WebSocket... but
that would need a ws client. Node doesn't have a built-in WebSocket client until Node 22+ (Node 22+ has a global WebSocket). What version
of Node is it? Let me check. Node 24 (from the nvm path v24.21.0) — a global WebSocket exists.

Plan:
1. Launch chrome-headless-shell with --remote-debugging-port=9222, --no-sandbox, --use-gl=swiftshader? For headless shell, WebGL via
SwiftShader: --enable-unsafe-swiftshader flag; also --no-sandbox is needed? For headless shell, maybe --no-sandbox isn't necessary,
but I'll add it just in case. Also, a new profile directory with --user-data-dir=/tmp/cdp-profile.
2. Node script: fetch http://127.0.0.1:9222/json/new?url=file:///... (PUT) to create a tab, connect via WebSocket to
webSocketDebuggerUrl, and do the following:
- Enable Runtime and Log
- Collect console messages and exceptions
- Wait about 6 seconds (let the first frame render)
- Dispatch key events? Simulating keydown via CDP: Input.dispatchKeyEvent. Possible, but keyboard isn't required for the smoke test.
- Page.captureScreenshot → save to PNG
- Also, evaluate expressions: get the chunk count, and confirm the renderer is working? For example
document.querySelectorAll('canvas').length, window.__stats... I haven't exposed the stats. I can evaluate JS in the page context:
the module's scope isn't global, but I can check document.getElementById('alt').textContent (the HUD should show altitude after
0.12 seconds), and check the canvas pixel colors via toDataURL? Easier: check the HUD text + screenshot.
3. Report console errors.

Let me write this script.

I'll drive the headless shell over CDP with Node's built-in WebSocket to capture console errors and a screenshot:

------------

> why did you used a browser? did you already had any instruction about it?

User's question: "Why did you use a browser? Was there any prior instruction to do so?"

Honest answer: there was no prior instruction to use a browser. What I chose on my own was:
1. To do syntax checking with node --check (low cost, so I did that first).
2. Since a single HTML file with WebGL + CDN can't be fully verified without actually rendering it, I used a locally available headless
Chrome (I found Google Chrome and Playwright's chromium headless shell on the machine). It's to catch runtime errors (shader
compilation, WebGL failures, logic errors like the inverted bank direction), and to take screenshots as visual evidence.

There's no instruction in AGENTS.md or in the skills that requires using a browser. There's nothing in the skill list that mandates it
either (the qa-manual-guide skill is about writing a guide, and I didn't use it since it wasn't explicitly requested… actually, this
skill says "when an implemented feature needs an evidence-based local manual testing guide…" — what the user asked for was a simulator,
not a manual testing guide. My verification was a homegrown approach).

Why I used it: a single-file browser game can't be verified with unit tests alone. Headless rendering would let me detect actual failures
(e.g., I actually found an inverted A/D bank bug that a syntax check could never find), and screenshots let me verify the visual quality
(mountain appearance, sun position, cloud scale) — the "beautiful and relaxing" requirement is inherently visual.

I should answer concisely and honestly: no, there was no instruction. I discovered the browser on the machine and used it for
evidence-based verification. The main findings were the inverted turn bug and visual confirmation. Also note that it was only used for
testing and didn't modify any files inside the project.

Keep the reply short and structured.

No — there was no instruction to use a browser. That was my own choice for verification, not from AGENTS.md or any skill.

Why I did it:
- A single-HTML-file WebGL game can't be properly validated by syntax checks alone — the interesting failures (shader compile errors,
WebGL context issues, sign errors in the flight math) only surface when it actually runs.
- I noticed Google Chrome and Playwright headless Chromium were available locally, so I drove a headless shell over CDP: load the page,
capture console/exception logs, simulate key presses, read the HUD values, and take screenshots.

▲
77
-1
27👁
r/LocalLLaMA · u/Gold-Bat-3225 · 22d ago
Humor Arena: Which LLM is the funniest? post image

We compared 20 model versions on 360 frozen joke prompts, with four jokes requested per prompt and model names hidden from our humor-trained judge.

Fable 5 had the highest estimated score: 66.8 points per 100 comparisons against the rival field. Fable 5.1 scored 58.2. Scores count a win as one point and a tie as half. The models near the top have overlapping uncertainty intervals. The reason we say estimated is that we scale the scores based on the scores of

The scores come from our own automated judge. A separate audit checked the judge with 1,400 ratings from 50 people; that was not a fresh human evaluation of Fable 5.1’s outputs. We specifically fine tuned an OS judge to rate the jokes and it correlates more highly with human preferences than any other model.

The full details here:
https://laugh.so/research/joke-generation/

Would love to hear your feedback!

▲
74
-3
20👁
r/LocalLLaMA · u/HolidayBit143 · 17d ago
Unsloth Studio VS LM Studio... Which one do you prefer?

So I've been experimenting with various platforms and even on day 1 Unsloth Studio was released, I knew that LM Studio had it's days numbered. LM studio will always be the OG but I wonder how much longer they have, especially with all these new platforms arising. It feels like LM studio just fell behind and it doesn't help that they are putting so much effort on BIONIC which I'm not even sure if anyone uses on a serious level.

Which one do you prefer, or do you use something else?

▲
81
+5
29👁
r/LocalLLaMA · u/returnity · 17d ago
mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)

Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:

  • Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
  • Self-supervised learning: trains itself on new material constantly
  • Weights stored on SSD and paged in on-demand
  • Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
  • Dynamic size: builds new experts and increases parameter counds as-needed
  • Evolutionary growth: trials newly generated experts, unused ones are pruned back
  • No tokenizer: it reads raw bytes directly
  • Catastrophic forgetting prevented by slow trunk/fast experts learning rate split

Weights will be released in "a couple weeks" once training progress reaches \~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.

What do you guys think?

▲
71
-4
27👁
r/LocalLLaMA · u/BagComprehensive79 · 17d ago
About Mimo 2.6 Architecture

I was checking nee Mimo 2.6 architecture on huggingface page and it looks very simple. I dont mean in a bad way but when we compare recent open models, their architecture is very simple. They dont use any Gated DeltaNet, no mHC or similar architecture, no engram. Just ordinary simple architecture and very good RL i guess.

What are you guys thinking about this?

▲
72
-1
24👁
r/LocalLLaMA · u/stoppableDissolution · 18d ago
Gewell - Gemma4 inference engine

\# the What

An engine to run Gemma 4 31B on blackwell under massive concurrency and rather specific workload patterns. I've been waiting for someone to do ninfer but for gemma, and, well, ended up having to do it myself.

More models and potentially more gpus are likely to be added, but its main purpose is to be my own workhorse, and I do not have the capacity (or desire) to chase every new release. I do love the gemma 4 family as a whole tho, so they are very likely coming soon.

\# the Why

Ironically, there has just been a post on "stop making slop inference engines", so... why bother with own engine if vllm exists? Well, neither vllm nor lcpp dont utilize one of the Gemma's big strengths, which is being able to have your kv cache use \*0.625x the vram\* losslessly. Not "trust me bro" losslessly, but like, mathematically losslessly down to the order of reduction.

Why? Because they decided to tie K and V weights on global attention layers, and rope only rotates 25% of K. So we can store only V and 25% of K, while other engines store full K and V. It is slightly more computationally intensive to have to unsqueeze them for the math, but it very quickly becomes outweighed by having to read less from memory. Blackwell has way more compute than vram bandwidth. And, well, lets you pack more context or more cached prefixes into the same amount of memory.

Also, vllm's cache sucks. Like, really sucks. It is good for when you have a lot of random users sending random prompts, but lack of explicit cache controls and LRU policy really makes some loads suffer, and SWA snapshots are clearly an afterthought (cant blame them for that because vllm predates SWA by a few years, but still). Gewell is built around efficient use of checkpoints, ram offloading and both smarther default eviction policy that assumes you are going to have repeating prompts with significant intervals and explicit cache hints on the prompts themselves. More about how cache works here: https://github.com/leDissolution/gewell/blob/main/docs/cache.md

Tl;dr: say, you have two chats going on you are alternating between. If you send ten messages into one of them in a row, vllm will make 10 checkpoints and evict the otehr one; gewell will dissolve some of the the intemediate checkpoints and preserve the second one warm.

Why it is important? Well, I'm using gemma for data generation and grooming, and most of these workflows have writer + ctitic or planner + writer + critic loops, sometimes with even more separate prompts cycling around. Each of these prompts is building up on top of its own's previous turn history so their prefixes are perfectly reusable, but vllm insists on pushing them out. It gets even worse if there are some one-off prompts that arrive every 10-20 turns and will never be reused, yet they still take up prefix cache and evict something useful.

Gewell also starts fast. Like, \*fast\*. Literally couple of seconds on top of reading the weights from the drive, because instead of doing live kernel profiling to select gemm shapes the choices were profiled offline and hardcoded and there is no python import tax.

\# the How Fast

Decently fast. TTFT is generally slightly behind vllm on large batches (because scheduler prioritized saturating decode width over latency and high-batch prefill is slightly slower for lower quants), but overall t/s is generally higher - especially on the workload it was designed for (bunch of prompts that keep growing but not all active at the same time).

https://preview.redd.it/tzn2iprbfuqh1.png?width=2188&format=png&auto=…

https://preview.redd.it/cho1porbfuqh1.png?width=2108&format=png&auto=…

https://preview.redd.it/2ul22prbfuqh1.png?width=1939&format=png&auto=…

\# the Quants

Gewell uses its own quant format that allows for arbitrarily mixed precision. The convertion tool lets you repack any compatible checkpoint with whatever bpw you want.

The "main" quant it was developed around is G0: https://huggingface.co/LeDissolution/Gemma-4-31B-it-Gewell\_G0

It uses around 6bpw, allocating most of them into attention and global-attention-adjacent MLP.

Why not qat? Well, because it is kinda bad in my experience (especially in the context fidelity and vision). Nvidia's nvfp4 was my go-to, but my personal tests showed that 16bit in attention are mostly wasted and mlp needs some juice too. Intuition being that if we take the beautiful precise 16-bit attention and then pass it through 4-bit up-gate, we just lose all that fine detail anyway. Idk whether it is mechanically correct, but seems to work? YMMW.

https://preview.redd.it/3e02slcefuqh1.png?width=1580&format=png&auto=…

https://preview.redd.it/6phv3wcefuqh1.png?width=1580&format=png&auto=…

https://preview.redd.it/gayesosvfuqh1.png?width=1580&format=png&auto=…

The tasks here are \~2.5k example mix pulled from aya\_dataset, OpenR1-Math, DocVQA, ChartQA, QASPER and code\_contests

NIAH is a RULER-inspired torture test where the model is fed a huge uniform block of key-value pairs with distractors and overwrites:

Record 3832768 stores value ocean.
...
Record 3832760 stores value rose.
Record 3832761 stores value pearl.
...
Record 3832767 stores value ocean.
Record 3832768 stores value river.

Requested keys in order: 3832768 3832760 ....

And the model needs to respond with exactly the same amount of values in the exact requested order. Amount of needles is 16 for the current test set; completion was counted as % of the correct values in correct spots. At 64k even bf16 can not complete a single request perfectly without reasoning.

\# the Supported Hardware

It was developed and tested on linux and pro 6000. I have not tested it on 5090 because I dont have it, but the intent behind choosing the quant size was to have the weights + mtp + 250k context fit in 32gb. Adding vision might require reducing the context size a bit.

Windows support was not tested either (my windows machine got 3090s), but there is nothing that prevents it in principle, so you are welcome to try.

\# the Limitations

I did cut some corners on the interfacing side. The samplers support is currently very rudimentary (only temp, top-k and top-p), there is no way to override the chat template (the latest google's one is hardcoded in), and some less common text/chat completion knobs might be missing.

\# the Roadmap

There are likely some bugs to be fixed I did not find when using it myself, and some more works has to be done around the API. Next big thing I plan is supporting 26A4, but no promices when.

I also have a bunch of ideas around better speculative drafting, and it might or might not come before 26A4.

▲
65
 
28👁
r/LocalLLaMA · u/AvidCyclist250 · 16d ago
Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080

Thought it's about time to share after testing for a week. You need four things most people miss: the right quant, the right model, the right branch, and the right cache flags.

https://github.com/dtm-beep/qwen38-flash-next-mtp-16gb

TLDR: AtomicChat AD-4.27bpw Q4_K_M target + the shared Unsloth MTP head, build from my pr-mtp-fix branch (plain master can't load this MTP head yet, it's PR #28243 + one fix commit), and --spec-draft-cpu-moe is the trick that makes 16 GB work. Draft experts live in RAM so the target's hot experts get the GPU. IQ4_XS ~10 t/s → 16.5 tg / 350 pp at 131k, q8 KV.

Hope it helps someone.

▲
65
 
15👁
▲
69
+4
21👁
▲
60
-4
22👁
▲
63
-1
32👁
r/LocalLLaMA · u/My_Unbiased_Opinion · 15d ago
PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192)

Just a quick PSA. llama.cpp does have prompt caching. if you are running large context lengths and have long multiturn projects, increasing -cram can provide you with massive speedups. There is a point where context lengths can get so large that 8192mb is not enough and the whole context needs to be re processed again on every turn. personally, I have found 20480 to work well with Qwen 27B 3.8 at 262K context.

the main downside is this uses more ram. vram usage doesnt increase.

▲
65
+1
26👁
r/LocalLLaMA · u/Porespellar · 21d ago
NGL, I’m hyped to see if Qwen3.8 27b can make me a sandwich. Instant buy for me.

Saw this little dude in a Forbes article (https://www.forbes.com/sites/johnkoetsier/2026/08/18/american-humanoid-robot-…)

This is definitely for the DIY researcher crowd who want to dip their toes into the robotics world. It is not going to be “consumer-ready” in any way shape or form, but I Instantly preordered the shit out of this anyways. I don’t even care that all the vids of it doing stuff are probably 3x sped up and completely cherry-picked and highly edited. I DO NOT CARE that it is going to be likely absolute trash getting started with this thing. It is still going to be fucking amazing that I’m going to have a robot that could potentially injure me for $1,688.

There is very little online about this guy, but I trust Forbes did their homework, and Nori has actually supposedly delivered the first batch of the earlier L3 version from what I can tell, and their discord is active and their SDK and documentation seems legit to me. I know I’m taking a risk with pre-ordering a highly beta product from a company that’s probably run out of someone’s actual garage, but son-of-a-bitch I’M IN!!

From what I can tell it’s Raspberry Pi 5 driven for control loop, with remote inference via WiFi. I plan on connecting it to my DGX Spark.

Here’s their site:

https://www.norirobotics.com

And their SDK doc site as well:

https://docs.norirobotics.com

▲
59
-3
32👁
r/LocalLLaMA · u/OvertaxedOne · 16d ago
Qwen FN vs 27B --- Think I'm saturated.

Got QFN up and running on our Strix box this past weekend and have been running on it for a few days now. Big thank you/shout out to the Halogen team, it's running fantastic on the Strix, this is clearly "the setup" right now for this hardware with this model, really impressive performance for such a large model (\~30-40TPS generation, \~900-1000TPS prefill on real world use, not benchmarks, over the past few days)!

However, that said, I kind of feel like 27B saturated my personal use cases. QFN is a great model, but I'm not really noticing much where I think "Wow, 27B would never get this and QFN just one shot it". They feel very similar in capabilities (IE, both are amazing!) and I kind of feel like I'm reaching the end of the runway for what I can realistically make use of in my day to day use cases, I'm just asking questions that really require more than 27B the majority of the time.

It really feels more to me like a very similar level of smart, one that runs well with limited memory bandwidth (QFN) and one that runs well with limited GPU memory (27B). It's an interesting result, and I guess maybe I should have expected it because I was really struggling already to find things that I actually need to do day to day professionally that 27B couldn't do. When I escalate to the cloud now it's nearly always for either speed or context, rarely intelligence.

Surprising result, at least to me, I was expecting to have my hair blown back, but I guess this is just yet another data point that "good enough is good enough".

▲
63
+2
36👁
r/LocalLLaMA · u/Fragrant_Scale6456 · 15d ago
Qwen3.8 27b practical modeling for 3d printing post image

I spent the past day and a half trying to get qwen27b to complete some practical work for me. I have a Bambu h2c I have been wanting to get more use out of so thought this would be a fun experiment.

I have 27b running on my 5090 and qwen image 2.1 running on a 3080 10gb with comfyui. I had pi build me some skills to use cadquery and comfy.

prompt: “Make me a printable 3d model of a self watering plant pot and a MagSafe phone stand for my iPhone 17.  Give me a sheet with top front and 3/4 view renders of each item.  Then, use comfy to generate a scene and place the rendered product in the scene.  It should look like an advertisement”

It’s not perfect but I’m honestly super impressed with the output. The multi view sheet renders having the amount of filament each object would use is a nice touch.

The setup is 5090 with ninfer, quasar qat 27b, 590k nvfp4 context, image processing enabled. in comfy I’m using the int8 version of qwen image 2.1. harness was pi with skills it made for cadquery, blender, and comfyUI.

My next goal is to be able to give it a series of photos of an object and have it create a faithful 3d model. If it can pull that off it would be great as one of my hobbies is making custom parts for my RC cars.

If anyone has played around with 3d creation and printing with localLLM I’d definitely want to hear about what tools you are using I have a feeling my setup is very basic at this moment.

▲
58
-3
27👁
r/LocalLLaMA · u/Dutchnamn · 16d ago
Perhaps the highest quality mainline quants of Qwen3.8 27B?

I am proud to release these quants of Qwen 3.8 27B. They beat the excellent ISTA and Unsloth quants byte-for-byte on three corpora. Both KLD and top 1% were tested 3x. It took a week of continuous GPU and CPU time to generate these, all done on a single Strix Halo.

https://huggingface.co/agentionai/Qwen3.8-27B-AP-GGUF

Hope you like it.

Edit: I did a lot of benchmarking and updated the smallest quant. Slightly improved calibration led to this real life result

https://preview.redd.it/eax6krk4x7sh1.png?width=1600&format=png&auto=…

▲
63
+2
25👁
r/LocalLLaMA · u/SnooPredictions515 · 18d ago
[Splash Engine] Qwen3.8-27B in native 8-bit at 37–55 tok/s on Apple Silicon: Extending Splash to Q8, 256k context scaling, and the "Reasoning Cliff"

https://preview.redd.it/nulsv53o8vqh1.png?width=4500&format=png&auto=…

Spent weekend benchmarking the Splash engine (by Incoai) and extending its architecture to native 8-bit on Apple Silicon (M5 Pro, 64 GB unified memory).

Splash is a compiled C++ and Metal speculative decoding engine designed specifically for Apple Silicon. Upstream Splash pioneered a blisteringly fast speculative decoding pipeline for 4-bit models (\~60 tok/s). However, aggressive 4-bit quantization hits a nasty "reasoning cliff" on competition-grade math and multi-step derivations.

We wanted to bring Splash's speed to true uncompressed 8-bit weights without losing its speculative decoding advantages. By extending Splash's architecture to support native 8-bit tiled Metal kernels (schema 5, MDFL0008), we were able to sustain 37–55 tok/s with zero quantization degradation.

Note on compatibility: Official upstream Splash 1.0 (incoai/splash) hardcodes package validation to 4-bit schemas (splash-packed-q4, schema 3/4). This fork adds schema 5 (splash-packed-q8, MDFL0008) loading and compiled Metal Q8 tiled decode kernels, while keeping 100% backwards compatibility with upstream Splash's official Q4 models. Proposed upstream: \[incoai/splash#94\]([https://github.com/incoai/splash/pull/94](https://github.com/incoai/splash/pull/94)).

Speeds on Apple Silicon (M5 Pro, 64 GB Unified Memory)

Evaluated at temperature=0.0 across 5 standardized task domains:

|Task / Domain|Prompt Description|Splash-Q4 (Official 4b)|Splash-HQ (Native 8b)|Splash-Q8 (Compressed)|MTPLX-Q8 (MTP D3)|Stock MLX / llama.cpp (AR)|
|:-|:-|:-|:-|:-|:-|:-|
|Math & Logic|Algebraic derivation|83.3 t/s|54.8 t/s|52.7 t/s|28.5 t/s|9.9 t/s|
|Coding & Algos|merge_intervals $O(N log N)$|75.5 t/s|34.7 t/s|40.3 t/s|28.8 t/s|9.9 t/s|
|Constraint Reasoning|3-chair spatial permutation|59.3 t/s|39.3 t/s|37.4 t/s|27.7 t/s|9.9 t/s|
|Domain Knowledge|FlashAttn vs PagedAttn|39.0 t/s|21.9 t/s|22.8 t/s|23.8 t/s|9.9 t/s|
|Nuanced Writing|Memory bandwidth constraint|46.2 t/s|33.7 t/s|29.5 t/s|23.4 t/s|9.9 t/s|
|AVERAGE|Across all 5 domains|60.7 t/s|36.9 t/s|36.5 t/s|26.5 t/s|9.9 t/s|
|Speedup vs AR|Relative to 9.9 t/s baseline|6.13x|3.73x|3.69x|2.68x|1.00x|

A few notes on the comparisons:

  • Splash-HQ vs MTPLX (+39% overall, +92% math): Both run on the exact same 8-bit base weights. But Splash’s compiled C++ Metal backend executes with significantly lower dispatch overhead than Python/MLX DraftCore, getting 36.9 vs 26.5 tok/s overall, and hitting 54.8 tok/s on structured math reasoning.
  • The Precision-Speed Paradox: Uncompressed native 8-bit (Splash-HQ, 27 GB) actually ran slightly faster on average than compressed 8-bit (Splash-Q8, 17 GB)—36.9 vs 36.5 tok/s. In speculative decoding, decode speed is $\\text{Draft Speed} \\times \\text{Acceptance Rate}$. Aggressive compression flattened logits and lowered draft acceptance; native 8-bit produced sharper logits, fewer verification rollbacks, and higher net throughput despite reading more bytes from memory.

Context Scaling: What Happens Up to 256k Context (Live Telemetry to 190k)

Qwen3.8 is architecturally specified with a native 256k context window (262,144 tokens). Most Transformers fall off a cliff in decode speed as context grows because the KV cache balloons.

However, Qwen3.8 uses a hybrid architecture: 48 recurrent linear DeltaNet layers (fixed $128 \\times 128$ hidden state, $O(1)$ memory growth with context) and only 16 full-attention layers.

On a 64 GB Mac, we pushed it live in an active server session all the way out to 190,016 tokens to see if decode speed degraded under real usage:

|Context Length (Tokens)|Cached Tokens|Generated Output|TTFT (Prompt Prefill)|Decode Speed|Notes|
|:-|:-|:-|:-|:-|:-|
|65|0|50|0.8s|35.7 tok/s|Short prompt baseline|
|16,433|15,040|232|4.0s|49.0 tok/s|Prefix cache hit|
|34,605|29,376|2,771|15.6s|27.1 tok/s|Long response generation|
|83,379|76,320|435|28.9s|30.1 tok/s|Deep context code review|
|106,212|98,752|29,487|34.0s|24.8 tok/s|Massive batch generation|
|157,961|157,056|400|6.1s|43.5 tok/s|Cache hit at 158k tokens|
|180,082|143,360|3,446|228.5s|33.3 tok/s|Extended reasoning session|
|187,613|186,720|425|6.5s|31.9 tok/s|Cache hit at 187k tokens|
|188,546|147,456|1,083|268.2s|21.1 tok/s|Partial prefill recompute|
|190,016|151,552|1,115|227.4s|32.0 tok/s|Max context reached (64GB RAM)|

(See the visual plot in the repo: *benchmark\_and\_context\_scaling.png* showing the full 51-point scatter and rolling trend line).

The big takeaway on context: Decode speed does not collapse. Thanks to Splash's memory handling and the hybrid architecture, it stays between 21 – 33 tok/s across the entire range.

The actual bottleneck at 150k+ context is cold prefill (TTFT). When the prefix cache hits, TTFT at 187k context is just 6.5 seconds. But on a cold cache miss, prefilling 180k+ tokens on a 27B model on Apple Silicon takes \~4–5 minutes. If you are using agent harnesses (like Oh My Pi, Claude Code, or curl), make sure client SSE idle timeouts are set high enough so the client doesn't drop the connection during cold prefills.

The "Reasoning Cliff" on Competition Math

Throughput numbers don't matter if math derivations hallucinate. We tested extended CoT reasoning on MATH-500, AIME 2025, and GPQA Diamond:

  • On MATH-500 Problem 0 (evaluating $\\sum\_{j=1}^(\\infty) \\sum\_{k=1}^(\\infty) \\frac{1}{(j+k)^(3}) = p - q$), both stock Splash-Q4 and compressed Splash-Q8 fell off a cliff: they suffered numerical drift halfway through the algebraic series manipulation and output wrong values.
  • Upgrading to Splash-HQ (full uncompressed 8-bit across all 64 layers) or Splash-Mixed (where only the top 8 sensitive layers, 56–63, are 8-bit) completely eliminated the cliff and cleanly derived $p - q$.
  • Upgrading just the deepest 8 layers restored the full symbolic precision while keeping RAM manageable.

https://preview.redd.it/29mgrn0c9vqh1.png?width=2400&format=png&auto=…

Setup Recipe (No Compiling Needed))

Needs: Apple Silicon Mac, macOS 26.4 or later, 48 GB unified memory (64 GB recommended; the weights alone are 27 GB).

The GitHub repo holds the C++ and Metal runtime engine, while the 27 GB model weights are hosted on Hugging Face. You don't need to manually download model files with git-lfs or separate scripts—Splash has a built-in package downloader.

1. Install the prebuilt engine (one line)

curl -fsSL https://raw.githubusercontent.com/npanj/splash/q8/install-q8.sh | sh

No Xcode, Homebrew or pip needed. It installs a splash-q8 command and doesn't replace an existing Homebrew splash.

2. Launch the server (Automatic Download on First Run)

When you run the command below, Splash automatically detects missing model artifacts, connects to Hugging Face, streams the 27 GB files with progress bars, verifies the manifest SHA-256 hashes, and boots the engine:

splash-q8 serve --model nitinpanj/Qwen3.8-27B-Splash-HQ

(Once downloaded, subsequent runs load instantly from local disk offline).

(Optional: If you prefer to pre-download the model files beforehand via Hugging Face CLI instead, you can run:)

huggingface-cli download nitinpanj/Qwen3.8-27B-Splash-HQ

3. Connect your client

The server exposes a standard OpenAI-compatible /v1/chat/completions endpoint on http://127.0.0.1:8000:

Test via curl curl http://127.0.0.1:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "nitinpanj/Qwen3.8-27B-Splash-HQ", "messages": [{"role": "user", "content": "Explain why uncompressed 8-bit weights improve speculative decoding acceptance."}], "temperature": 0.0 }' # Or connect Oh My Pi (OMP) omp --model splash/nitinpanj/Qwen3.8-27B-Splash-HQ

Building from source instead? You need full Xcode, not just the Command Line Tools. On Xcode 26, first run xcodebuild -downloadComponent MetalToolchain, then make -j4.

Practical Gotchas & Details

  1. Memory headroom at 150k+ context: On a 64 GB Mac, model weights take \~27 GB. As context pushes towards 180k–190k, working memory climbs to \~42 GiB. Metal's memory governor will pause allocation growth when system free RAM dips below \~50 MB (Memory: growth paused). If you don't need 190k context, you can pass --max-context 131072 to cap it cleanly.
  2. Backwards compatibility: \splash-q8\ also serves the official 4-bit models. This fork preserves all upstream Splash 4-bit dense and MoE schemas (splash-packed-q4, splash-packed-q4-moe), so you can serve official models like incoai/Qwen3.8-27B-Splash or incoai/Qwen3.6-35B-A3B-Splash directly.
  3. Only tested on Apple Silicon (Unified Memory): Everything here relies on unified memory bandwidth and Metal tiled shaders; not tested on CUDA or CPU.

Credits & Attribution

Full credit to the Incoai team for creating Splash (https://github.com/incoai/splash). Their C++ Metal speculative decoding architecture is what makes these speeds possible on Apple Silicon in the first place—this fork simply extends their work to support native 8-bit weights and custom Q8 tiled kernels. Also huge credit to the Qwen team for base weights and MTP architecture, and Youssofal for MTPLX reference benchmarks.

Updated: setup so that compilation is not needed

Updated (10/1): you can now find follow up work for Qwn3.8-Flash-next here: https://www.reddit.com/r/LocalLLaMA/comments/1wva7l2/running\_955\_gib\_qwen38flashnext\_at\_4152\_toks\_on\_a/

▲
60
+1
25👁
r/LocalLLaMA · u/ThomasAger · 20d ago
I enjoyed the daily HF papers today

Top 3 papers on HF Daily Paper are all unusually delightful and interesting reads for anyone on the leading edge of local LLMs, agent harness optimization, etc, felt like sharing.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

https://huggingface.co/papers/2609.19969

Cross-layer KV reuse plus FP4 KV caching brings the global KV cache to 890 bytes per token, about a quarter of V4-Flash.

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

https://huggingface.co/papers/2609.20519

Auto-research loops that improve the agent harness, cutting token traffic by 44.7 to 49.0% at comparable performance.

An Empirical Study of Harness Design for Coding Agents

https://huggingface.co/papers/2609.20804

Varies planning, action space, and context management across 176 settings to see what each component actually contributes.

I'm still reading through, feel free to discuss

▲
61
+3
25👁
▲
61
+3
24👁
r/LocalLLaMA · u/9r4n4y · 20d ago
So i tried Remotion with glm 5.3 flash, this mfker is really good. post image

\*8bit, vllm, 4x dgx\*

Prompt:

Go download and use Remotion and create a cool 60-second motion graphics video with it. Impress me totally. The motion graphics must be based on stock market visuals. Add many cool, mind-blowing motion graphics and explain what fundamental vs. technical analysis is in investing, along with their pros and cons.

\*No special skill.md used from my side.

\*harness - Zcode

It's ofc better than my previous method of making videos though java.

I think maybe in future with high bandwidth flash , we will be running these size models on affordable hardware.

How it made it report: https://github.com/9r4n4y/ProjectsorSkills/blob/main/Video\_Generation/Remotion/Making-of-Market-Decoded\_Build-Documentation.pdf

▲
57
-1
23👁
r/LocalLLaMA · u/Character-Result-281 · 21d ago
Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark post image

Planning, coding, testing = 8h total.
Stack: VSCode Copilot in autopilot mode + SGLang
Stats: ∼10k lines generated, ∼800k tokens consumed

Sure, it's not GPT-6 Astra level, but for a 100% local ∼180B MoE running on a single DGX Spark at ∼35 tok/s. Not bad...

▲
5
+2
5👁
r/LocalLLaMA · u/Proof_Nothing_7711 · 15d ago
Qwen 3.8-27 tips for my setup

Hi everyone, I’m an AI newcomer eager to learn and experiment. I’m comfortable with coding on my own, but I want to explore the AI ​​world—specifically for code review. I have two separate setups depending on my location: 1) Minisforum Ryzen 9 HX 370 AI with 64GB DDR5 RAM + OCuLink eGPU (AMD W7800, 48GB VRAM) 2) Beelink SER5 Max Ryzen 7 6800U with 32GB DDR5 RAM + OCuLink eGPU (AMD R7900 XTX, 24GB VRAM) Both setups run Windows 11, though I’m open to switching to Linux if it would improve performance. For the LLM, I use Qwen 3.8-27b (Q8 on the W7800, Q4 on the R7900 via Vulkan) for code review, as mentioned. I started out using LM Studio but have since switched to VS Code with Cline. Could you please offer some advice on optimal settings for this model and, if possible, tips on how to best configure VS Code with Cline? I’d like to switch to Ollama and move away from LM Studio (hoping for a smoother experience). Thanks in advance—and apologies if these questions seem basic; I’m just trying to learn as I go. Thanks!

▲
31
-3
8👁
r/LocalLLaMA · u/jjusko20 · 15d ago
I'm trying to post-train AliceAI-Foundation-80B-A3B-Base myself

Just wanted to share with someone - don't have much to report yet. I am interested in this new AliceAI model and have been wanting to make a community impact for a while - and releasing an initial agentic version of this model sounds cool. I am training on 3 32gb v100s (which has been fun to get to work, to say the least). What I'm really doing is creating a shallow distill of Qwen 3.8 27b and then using reinforcement learning - My initial plan is a SFT with Qwen3.8 27b synthetic data I'm generating targeting long horizon agentic work - then, RL / GRPO with a grader model for a while. I'm considering using a stronger model to generate the training data, but I'm trying to keep this on my local machine only. It's coming along - I can just barely fit the weights and activations on the v100s in qlora. I've built the framework for the SFT data generation for. I don't expect anything amazing but it should be a neat experiment. Also considering using a pre-existing data set for the tune, but I'm more interested in creating my own distillation. Update: it's training! https://preview.redd.it/vx7ybgzbxjrh1.png?width=1329&format=png&auto=…

▲
5
-3
8👁
r/LocalLLaMA · u/DerTomsn · 15d ago
7900 XTX — two "low-thinking" Qwen 3.8 27B quants (Swift + ThinkingCap) vs the regular quant

First, do they actually produce less tokens? Yes. Total tokens per benchmark run (4 scenarios): base quant \~66k, ThinkingCap \~49k (−26%), Swift \~45k (−33%). So the "less thinking" is real — and Swift cuts the most. Then the cost: and this is where it got interesting. The two quants don't trade off the same way: Decode: base \~48 t/s. ThinkingCap barely changes (\~43). Swift drops hard (\~32, −33%). Prefill / TTFT — the opposite of what I expected: Swift is the fastest (\~614 t/s, TTFT \~1s), base in between (\~530 t/s, \~1.2s), ThinkingCap the slowest (\~100 t/s, TTFT 6–11s). * Quality holds: \~84–87 on my eval, same band as base. Full run data: ThinkingCap finishes \~23% faster than base. The token savings win even with the slow prefill. Swift is break-even because of the slower decode speed. |*Quant*|*Prefill*|*Decode*|*Quality*|*Runtime*| |:-|:-|:-|:-|:-| |Base (unsloth)|\~530 t/s|\~48 t/s|\~85|\~1407s| |ThinkingCap|\~100 t/s|\~43 t/s|\~85|\~1087s| |Swift|\~614 t/s|\~32 t/s|\~86|\~1400s| Caveat: 2 runs per quant only, so single-run variance will move these. Prefill speed of ThinkingCap is oddly low. Need to do some more tests on that. Side-by-side (thinking xhigh, Q4\_K\_M, all 7900 XTX) of 3 of the runs: https://llm-bench.io/compare/runs?runs=cmufxkpkb00bj01o05h9vsu5o,cmufwold900ay01o0ae1frktu,cmufyiq3p00by01o0vwbpilew

▲
12
+1
8👁
r/LocalLLaMA · u/nirurin · 15d ago
Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?

I was actually pretty happy with my Qwen3.8-27b setup, and I'd been tinkering with Ninfer to have a version that was "fast but maybe a bit stupid" and the speed was nice to have as a backup. But I was curious how the Flash-Next version might work, after I learned it didn't need to all fit in VRAM to work. I picked up the Atomic quant (let me know if there is a better one I should use, this one seemed good from what I could find). I used the build setup below. It can still be tweaked some more, as I still am only using about 27gb of my vram. > ./build/bin/llama-server \\ \--model "/mnt/SPCC-2TB/Projects/AI-APPS/LLM-Models/Qwen3.8-Flash-Next-Atomic/Qwen3.8-Flash-Next-AD-4.27bpw-Q4\_K\_M-M64 \-00001-of-00033.gguf" \\ \--no-mmproj \\ \--load-mode mmap \\ \--lazy-mode on \\ \--fit off \\ \--gpu-layers all \\ \--n-cpu-moe 32 \\ \--ctx-size 64768 \\ \--flash-attn on \\ \--jinja \\ \--parallel 1 The odd thing I noticed though - I know that some parts of this are meant to run from the SSD for the sake of saving vram space etc. Fine. But I kinda expected that some of it at least would get buffered into system ram, as running from ram would be a whole lot more efficient than running from my NVME drive. But this run gets me the following results: 40tok/s decode. 50tok/s prompt processing (It's a short prompt so probably not accurate) 27gb of vram used 8gb of system ram used... So... I mean, am I just wrong and this is normal? The speed doesn't seem as bad as I expected (I thought I was going to get more like 10tok/s at best) but it seems like I might be missing a trick somewhere?

▲
10
+2
6👁
r/LocalLLaMA · u/Retumbo77 · 15d ago
Better alternatives to PocketPal for Local on Android?

I've been using PocketPal for Android, but I've always been plagued by stability issues (even when models aren't loaded - It's not a RAM utilization thing). I have seen this come up a few times in various places, but PocketPal usually gets the recommendation. Does anyone have any alternatives?

▲
10
 
5👁
r/LocalLLaMA · u/Whahine · 15d ago
Dailychained PLX 88096 switches, Quad RTX 5070 Ti + Quad RTX 5060 Ti

Left: Quad RTX 5060 Ti, Middle: Quad RTX 5070 Ti, both PLX 88096 switch, Right = rehomed host Edit: Benchmarsk were 1 line = fixed WHY: -->> DATA SOVEREIGNITY / PRIVACY<<-- hey, this is (localllama right?), this makes no less sense than my dropping the same $$$ on a motorbike I want but don't need so no Triumph Rocket III motorbike for me boo hoo, is for SOHO Anyway, a bit of a journey, a few hundreds of $$ wasted on power adaptors / pci risers that are not suitable and a small fortune in RTX 50xx GPU that will be obsolete eventually Now I'm still buried in the steep learn to use linux / docker / vllm / models / setup clients learning curve (I am windows since Win 3.1). I am yet to learn to relove the CLI (not since ICL/IBM MFs in the 80's) All setup and running to the point vLLM NCCL messages report P2P enabled within each node, (yet to resolve getting P2P across nodes). A little more work on the cooling (more fans coming)/ best orientation etc to do I expect to have these for a while, hence no loose mining frames etc each of these nodes is self contained built up hardware (prototype quality, a few rough edges here and there), If I can get the GPUs just build another one (I have spare V21 case + 88096 PCB) Note I am in NZ so all I can buy locally is regular basic PC parts, pretty much everything else is overseas import eg even the Thermalrake Core V21 cases had to come from Australia, most everything else is from China 2-4 weeks shipping, if a cable or adapter doesn't work then more delays and i pay sales tax 15% at the border Also these cases party trick is they can be bolted vertically so assuming I can manage rising heat that option saves a bit of space on the desk GPUs are mixed brands/models, a couple I already had, I was incrementally ( every local seller is '1 GPU per customer') collecting 8 x RTX 5060 Ti initially for this build, but at the point I got the the 5th one the price delta between 5060 Ti 16GB and 5070 Ti 16GB got close enough I returned that 5th one, added 3 x RTX 5070 Ti 16GB to one I had already. Note there really is no cost effective used GPU market here so I was only able to buy 1 5060 ti used and I only paid him NZ$150 over what he paid for it in Nov 2025 ;) (Nice for him but even at that markup it was still a score) Looking forward to getting 3.8 Flash Next working too fingers crossed is usable Benchmarks Run command per node below Ubuntu 24.04, NVidia 575.something, CUDA 13.2 vllm 0.30 \- so you can see what is enabled (you are correct and thank you for noticing, yes I really do not know what I am doing on the software side, I am just a very old script kiddy) no spec decode etc to keep it reproducable docker run --rm -it \ --name vllm-node1x4 \ --ipc=host \ --gpus '"device=GPU-dd49c72c-4273-9016-aaad-8883c553b0da,GPU-1ef97316-f1c1-3c15-64db-fbd6373179e5,GPU-ab8220cb-5ab1-82a6-8aec-26454657e215,GPU-6ce2d880-240f-18be-ffe0-96ab3d0eede8"' \ -e NCCL_DEBUG=INFO \ -e NCCL_P2P_DISABLE=0 \ -e NCCL_P2P_LEVEL=SYS \ -e VLLM_SKIP_P2P_CHECK=1 \ -e NCCL_BUFFSIZE=16777216 \ -e NCCL_MIN_NCHANNELS=8 \ -v /mnt/ai-assets/huggingface:/root/.cache/huggingface \ -v /mnt/ai-assets/vllm-cache:/root/.cache/vllm \ -v /mnt/ai-assets/models:/models:ro \ -p 8005:8005 \ vllm/vllm-openai:latest \ /models/Qwen/Qwen3.8-27B-FP8 \ --served-model-name qwen3.8-27b-fp8 \ --quantization fp8 \ --tensor-parallel-size 4 \ --max-model-len 65536 \ --max-num-seqs 10 \ --gpu-memory-utilization 0.85 \ --kv-cache-dtype auto \ --host 0.0.0.0 --port 8005 Tests command llama-benchy \ --base-url http://localhost:8005/v1 \ --model qwen3.8-27b-fp8 \ --tokenizer /models/Qwen/Qwen3.8-27B-FP8 \ --depth 0 4096 8192 16384 32768 \ --latency-mode generation Results \- No overclock/undervolt etc stock GPU settings (Todo: NVOC overclock VRAM) \- look very linear to me, basically 2 to 1 \- results seem ok Quad 5070 Ti 16GB | model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) | |:----------------|----------------:|---------------:|-------------:|---------------:|---------------:|----------------:| | qwen3.8-27b-fp8 | pp2048 | 6488.44 ± 3.99 | | 363.10 ± 0.19 | 315.79 ± 0.19 | 363.10 ± 0.19 | | qwen3.8-27b-fp8 | tg32 | 82.21 ± 0.06 | 84.86 ± 0.06 | | | | | qwen3.8-27b-fp8 | pp2048 @ d4096 | 5831.49 ± 4.69 | | 1100.96 ± 0.70 | 1053.65 ± 0.70 | 1100.96 ± 0.70 | | qwen3.8-27b-fp8 | tg32 @ d4096 | 81.56 ± 0.09 | 84.19 ± 0.09 | | | | | qwen3.8-27b-fp8 | pp2048 @ d8192 | 5642.19 ± 3.35 | | 1862.27 ± 1.10 | 1814.96 ± 1.10 | 1862.27 ± 1.10 | | qwen3.8-27b-fp8 | tg32 @ d8192 | 81.00 ± 0.01 | 83.61 ± 0.01 | | | | | qwen3.8-27b-fp8 | pp2048 @ d16384 | 5431.23 ± 5.45 | | 3441.21 ± 3.25 | 3393.89 ± 3.25 | 3442.26 ± 3.26 | | qwen3.8-27b-fp8 | tg32 @ d16384 | 80.63 ± 0.14 | 83.23 ± 0.14 | | | | | qwen3.8-27b-fp8 | pp2048 @ d32768 | 5059.44 ± 0.49 | | 6928.84 ± 0.58 | 6881.52 ± 0.58 | 6929.81 ± 1.32 | | qwen3.8-27b-fp8 | tg32 @ d32768 | 79.43 ± 0.26 | 81.99 ± 0.27 | | | | Quad RTX 5060 Ti 16GB | model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) | |:----------------|----------------:|---------------:|-------------:|----------------:|----------------:|----------------:| | qwen3.8-27b-fp8 | pp2048 | 3218.53 ± 1.05 | | 689.86 ± 0.35 | 636.52 ± 0.35 | 689.86 ± 0.35 | | qwen3.8-27b-fp8 | tg32 | 43.35 ± 0.02 | 44.74 ± 0.02 | | | | | qwen3.8-27b-fp8 | pp2048 @ d4096 | 3005.17 ± 0.55 | | 2097.92 ± 0.37 | 2044.59 ± 0.37 | 2097.92 ± 0.37 | | qwen3.8-27b-fp8 | tg32 @ d4096 | 42.85 ± 0.01 | 44.23 ± 0.01 | | | | | qwen3.8-27b-fp8 | pp2048 @ d8192 | 2927.62 ± 1.87 | | 3551.28 ± 2.59 | 3497.95 ± 2.59 | 3551.28 ± 2.59 | | qwen3.8-27b-fp8 | tg32 @ d8192 | 42.66 ± 0.06 | 44.03 ± 0.06 | | | | | qwen3.8-27b-fp8 | pp2048 @ d16384 | 2821.32 ± 1.44 | | 6586.68 ± 3.60 | 6533.34 ± 3.60 | 6587.62 ± 3.66 | | qwen3.8-27b-fp8 | tg32 @ d16384 | 42.25 ± 0.05 | 43.61 ± 0.05 | | | | | qwen3.8-27b-fp8 | pp2048 @ d32768 | 2654.29 ± 0.46 | | 13170.73 ± 2.42 | 13117.40 ± 2.42 | 13171.73 ± 2.56 | | qwen3.8-27b-fp8 | tg32 @ d32768 | 41.30 ± 0.07 | 42.63 ± 0.07 | | | | \--- My plan more or less from a while back, with hardware notes pretty much up to date My justification to target 128GB/ All Blackwell: - 128GB = DGX Spark, RTX Spark And Strix 128GB AIOs so will be relevant for a couple of years - All Blackwell = FP8 fast now NVFP4 = faster once mature/production ready? (late 2026?) - Hopefully significantly faster than say DGX Spark - Need 192GB VRAM?: -- Build another Node etc (assuming RTX GPUs still relevant to AI inference, one day used will be < $$) - if not, easier to sell 1 x Node (or worse case 4 x GPUs individually per Node) than 1 x nonolithic DGX Spark or whatever once future wonder AI execution chips exist My justification to target PEX 88096 Backplanes - GPU P2P within each node with patched Nvidia drivers on Linux - Maximise capabilities of (relative to Node1) constrained RTX 5060 Ti 16GB PCIe x8 - Backends e.g. vLLM with say TP=4, PP=2 hopefully maximize architecture - PCIe4 = less bandwidth BUT: -- SO VERY much more forgiving re interference -- MUCH less $$ than anything PCIe5 -- NVidia p2pbandwidthlatencytest shows < 1us latency GPU P2P within each switch Each Node - Modular/ self contained, just chuck a spare SFF-8654 PCI host card into any PC and go AI LLM Inference Tiers - <= 64GB -- Performance / production tier: Node1 (GPU 0-3): vLLM TP=4 = FAST -- Experimentation / second model tier: Node2 (GPU 0-3): vLLM TP=4: Fast enough?? - > 64GB and <= 128GB = Node1 + Node2: Capacity tier: -- Node1 (GPU 0-3) + Node2 (GPU 0-3): vLLM TP=4 PP=2: Constrained to at best Node2 speed, good enough? - Other: -- Host RTX 5080: Embedding eg Qwen3-VL-Embedding-8B watever -- Node2 GPU4 RTX 3080: STT/TTS whatever Host GPU RTX 5080 - Use standalone for utility eg Embedding / Vision / Spec Decoding etc Host (Host 128GB DDR5-6000) - Ryzen 5 9600X - MSI MPG B850 Edge TI WiFi - Jonsbo D41 Mesh Black - XPG Core Reactor II VE 850W - iGPU only to Monitor - SATA SSD for each of WIN / Linux OS - Gen5 M2 SSD on PCIe5 x4 (CPU): -- 2TB = Docker -- 4TB = AI Assets HOT - SATA 2 x 28TB Barracuda HDD Mirrored (Linux) -- AI Assets COLD - RTX 5080 16GB in PCI_E3 (PCIe4 x4 Chipset) Node1 Performance node - 64GB VRAM (TP=4 Parallel) - Chassis: ThermalTake Core V21 - PSU: MSI MEG Ai1600T - PLX/PEX 88096 PCI 4 slot switch (with downstream SFF-8654 ports) -- Slot 1/4: RTX 5070 Ti -- Slot 2/4: RTX 5070 Ti -- Slot 3/4: RTX 5070 Ti -- Slot 4/4: RTX 5070 Ti -- All GPU PCIe4 x16 within PEX 88096 Node2 Capacity / Secondary node - 64GB VRAM Secondary node (TP=4) - 16GB VRAM Utility (RTX 3080) - Chassis: ThermalTake Core V21 - PSU: DeepCool PN1200M - PLX/PEX 88096 PCI 5 slot switch = Tensor Parallel 64GB VRAM (Secondary/capacity node) -- Slot 1/5: RTX 5060 Ti 16GB -- Slot 2/5: RTX 5060 Ti 16GB -- Slot 3/5: RTX 5060 Ti 16GB -- Slot 4/5: RTX 5060 Ti 16GB -- Slot 5/5: Alienware RTX 3080 OEM 10GB (I have a 5th slot and a spare 3080 so..) -- 4 x RTX 5060 Ti 16GB GPU PCIe4 x8 (Due to 5060 x8 electrically) within PEX 88096 HOST <-- SlimSAS PCIe4 x16 --> Node1 <-- SlimSAS PCIe4 x8 --> Node2 Hardware porn before transitioing to the PCI switch approach I started this build Nov 2025 slowly sourcing the parts for a new standalone PC meant as my triple 4k gaming + AI experimentation rig, then I found I wasnt gaming and I kept adding GPUs (well, VRAM really) V1: RTX 5080 + RTX 5070 Ti 16GB \(early build photo, missing a few other parts\) V2: RTX 5090 + RTX 5070 Ti V2: RTX 5090 + RTX 5080 + RTX 5070 Ti 16GB Then I picked up a 5060 Ti 16GB and was thinking how the hell do I squeeze this in, looked at M2 to what ever adaptors etc etc, yes my MB has bifurication etc etc, ordered a couple then thought nah thats all getting pretty manky, hence the pivot to the current approach, more or less homogenous nodes re generation / vram, plug any node into any PC with spare PCI slot, easy to move and so on Random POC / mid build photos 5 way 88096 PCB will it Post? = YES Host to 4 slot switch daisychain to 5 way switch, will they post/can I see GPUs? = YES \- 4 GPU in let downstream green (5 way switch) 2 in upstream black 4 way switch at which point I ran out of room / power cables / risers / bits of wood, but hey, I could see all bits in linux in a massive pci tree Host to 4 slot switch daisychain to 5 way switch, will they post\/can I see GPUs? = YES How to mount GPU array in these cases? Was looking for a blade asthetic RTX 5060 pretty easy 5070 a little tighter thats a lot of transistors PCB adapter plates in progress this one got moved about 4 times due to cable constraints etc Work with the fragile risers, dont fight the bends they came with [Now you get the apprao

show remaining 1,211 characters

ch](https://preview.redd.it/nyl5mifowirh1.png?width=1024&format=png&auto=…) 4 x RTX 5060 Ti + 1 x RTX 3080 10GB Part populated just 2 gpus in each, 88096 80mm cooling fans installed, running ok for some early work patching for P2P etc etc 88096 80mm cooling fans installed \(Middle\) Quad 5060 Ti populated and running, left still waiting parts Now.. should I sell the 5090 that is now in my older rescurected gaming PC? ( 5600x PCIe4 32Gb DDR4). Probably yes... Note: For a forum dedicated to AI there sure seems a 'wierd 'I hate AI managed posts/slop' herein, so you guys relax I personally fingered every word above (except for some of the vLLM command env vars/switches) Laters

▲
10
+1
5👁
r/LocalLLaMA · u/Jromagnoli · 15d ago
I'm new and it's kinda overwhelming to get into

Hi, sorry if this doesn't belong here. Getting to the point basically, I've been using online-only AI like GPT/Gemini since 2022, and have been interested in local models but am clueless overall. Yes I'm extremely late. I only use laptop (I'm a student), and I currently own: - Acer swift 5 SF514-55TA (main, budget laptop) | .| . | | --- | --- | | Installed Physical Memory (RAM) | 16.0 GB | | Total Physical Memory | 15.8 GB | | Available Physical Memory | 4.49 GB | | Total Virtual Memory | 25.3 GB | | Available Virtual Memory | 6.33 GB | - Acer Nitro 5 AN515-53 (not used currently) | . | . | | --- | --- | | Installed Physical Memory (RAM) | 8 GB | | Total Physical Memory | 7.85 GB | | Available Physical Memory | 4.96 GB | | Total Virtual Memory | 9.72 GB | | Available Virtual Memory | 5.89 GB | &nbsp; I don't know if it's even possible to set up anything on these. I vaguely know the basics of local models (but I'm probably too dumb for stuff like finetuning, prompts, personal setups, all that fancy shit I see online), and I have a ton of information, models, etc bookmarked on my browser which makes it hard to choose something I know (and am interested in image generation, Chatting (models)? "agents" (task programs?). I guess that needs a lot of separate programs/installations? Or is it possible to have a single client to run varying programs? (im not sure if that's even referred to correctly?) (I'll likely start small) I assume a "program" like LM studio is a good starting point?

▲
0
 
5👁
r/LocalLLaMA · u/LearnNTeachNLove · 15d ago
Any local open source AI model you would recommend?

Are there people in the community trying/testing the open source local AI models? If yes, any recommendation? Did you find a model that fits your regular expectations or even competes with the closed source private AIs? I suppose it depends on the hardware performance and on the expectations/work of the user, but just curious to know your respective experience. And it is highly probable that your recommendation changes every month… PS: from the various answers I am reading i understand that recommendations depend on the material i am using and on what i want to do with it. I have a very basic equipment with 8GB vram 4060 with 32GB of RAM. I do not really know to which extent i can ask a local AI to do something. I mostly saw demonstrations of people asking the AI to make html games, a snake, flappy bird game, to automate some simple tasks, to generate web pages, in a sense i can understand the hype on the other hand on my side i just have a feeling of not knowing what to do with it. It will just hallucinate, or reply in loop… i do not see myself extra spending for a few GB more of VRAM, on the other hand i do not want to use the models of these big AI companies… privacy, freedom, self-autonomy… Maybe the question i should ask is for what kind of activities/hobbies/work do you use local AI (which one?)? How performant is it from your point of view?

▲
40
+1
7👁
r/LocalLLaMA · u/jacek2023 · 15d ago
FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face

from FreedomIntelligence: HuatuoGPT-3-27B is a medical LLM built on Qwen3.8-27B with One-stage Policy Optimization (OnePO). OnePO adapts language models to medicine in a single reinforcement-learning stage, without preceding domain-specific supervised fine-tuning. Teacher responses provide temporary guidance and are retired as the model improves. We release the training code, medical RL dataset, and 8B rubric grader. (last week they released https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B)

▲
1
 
2👁
r/LocalLLaMA · u/mrgreatheart · 15d ago
Epyc for inference

Hi. My current system is an Intel Ultra 7 with 64Gb DDR5 at 6000. It has 4 GPUs totalling 72Gb: A 3090 and 5070 Ti on x8 CPU connected PCIe slots plus two 5060 Ti - one on an x4 CPU connected m.2 socket, and the other on an x4 chipset connected PCIe slot. I am considering upgrading to an Epyc 7443 system with 256Gb of DDR4 8 channel. Given prices I’ll probably end up with 2400 speed sticks giving a theoretical memory bandwidth of 153Gb/s. The obvious benefits are getting all the GPUs on proper CPU connected x16 and x8 PCIe with room for another at some point plus enough RAM to overflow bigger models than I can fit in VRAM. I’m struggling to find clear information on whether this would actually be worth it for the cost. I am currently able to run Qwen3.8-flash-next in two rather painful configurations (not using the chipset connected 5060 because it hurts too much): \- IQ4\_XS at 300 pp / 40 gen in llama.cpp \- EXL3.05bpw at 1,200 pp / 25 gen in exllamav3 The low tg in exllamav3 appears to be because I have to offload 10 layers to CPU. Obviously the extra PCIe slots would bring the other 5060 into play but I’d like to know what possibilities the extra RAM and bandwidth would open up. Is anyone here offloading larger quants (or other large models) to CPU on Epyc, and if so how usable is it?

▲
0
 
3👁
r/LocalLLaMA · u/lots_of_puppies · 15d ago
Would this be super fast for Qwen 27b & next? 32GB isn't much overhead for Q8 + 256k context

I have a 128gb m5 max now. It runs Qwen 3.8 27b well but as context grows the PP and output is so slow. Qwen Next is faster but still too slow to do lots of agentic work (coding) with. I am not a video gamer and I feel really bad if I contribute to hurting them if I buy this computer just for LLMs, but I would love to have a very speedy local Qwen 3.8 27b :p https://www.corsair.com/us/en/p/gaming-computers/cs-9060022-na/vengeance-a8200-gaming-pc-amd-ryzen-9-9950x3d-geforce-rtx-5090-64gb-ddr5-6tb-m-2-ssd-win11-pro-cs-9060022-na From what I've read with those two qwen models, I think my PP would be over 3,000 and my tks out would be 80 to 150? That sounds amazing to me. The only issue is the 5090 only has 32GB ram though so I'm scared I would have to use a very low quant like Q4 and that I couldn't fit a full 262k context. I was wondering if anyone knows if this build above would be worth it? Thank you!

▲
9
+1
4👁
r/LocalLLaMA · u/daphatty · 15d ago
Misrepresentation of this community?

Ever since I joined this subreddit, I’ve noticed something odd. Every r/localllama post that appears in my Latest feed is a propaganda post either for or against open/closed LLMs. Every single one. However, visiting the subreddit directly tells a different story. There are many helpful and enlightening discussions, the kind that made me subscribe in the first place. So what gives? Why is this subreddit being misrepresented in my Latest feed? It’s easy to blame “the algorithm” but what does that even mean? I’m certainly not a tinfoil hat type and I have zero interest in the pro/con discussion. I just want to read and learn more about self hosting LLMs.

▲
0
 
3👁
r/LocalLLaMA · u/skeole · 15d ago
Dynamic Abliteration: Non-Destructive Refusal Suppression via Multi-Layer Engram Steering

Cool use case for engrams! Realtime abliteration, model weights stay intact.

▲
0
 
2👁
r/LocalLLaMA · u/llo7d · 15d ago
Local AI you can give to your Mom

I made this thing called Hey Taby Its and easy and cool way to use local AI, it has a cute face that lives in the top of your screen. You can give it to your mom (i dont mean it like that) on https://heytaby.com

▲
8
-1
2👁
r/LocalLLaMA · u/bakatristan · 15d ago
I built an open-weight alternative to Jev / TypeSafe - introducing OpenJudgement-4B (early preview)

I’m releasing OpenJudgement-4B-Preview, an experimental Qwen-based model fine-tuned on custom datasets for classification, scoring and true/false judgments. It scores answer options directly, and Python formats the results into JSON with probabilities. It still uses an LLM backbone, but doesn’t generate the response token by token. It’s unfinished and isn’t at Jev’s level yet. I’d love feedback, especially examples where it gets things wrong. Use it via api at: https://kitani.ai/models/kitani/OpenJudgement-4B-Preview (paid) Model and inference code: https://huggingface.co/kitaniai/OpenJudgement-4B-Preview

▲
2
+2
2👁
r/LocalLLaMA · u/bruns20 · 15d ago
Advice for single 9700 on windows

Hey guys, I've recently gotten a r9700, which I'm very excited about. I've been trying to research best setups and Llama flags, but I'm finding that a lot of the advice I'm seeing is based around 2x9700's and/or Linux. I know windows is the devil, but I use this computer for other things as well, so I'm not looking to switch. Anybody have some up to date advice on running a single 9700 on windows? Edit: For anybody in the future, I ended up using WSL to load this beautiful man's vLLM build : https://www.reddit.com/r/LocalLLaMA/comments/1wiws8e/153\_toks\_on\_1x\_amd\_radeon\_r9700\_running\_qwen38/ PP increased almost 5x and considerable bump in decode speed too

▲
52
+2
2👁
r/LocalLLaMA · u/Public_Umpire_1099 · 15d ago
R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5]

Here's my \*first\* implementation of KVA projectors on QFN (just the uncensored model for now) the highlights are basically as follows for using the projectors at each different layer: Starting at layer 12, prompt processing speeds up 1.85x \[1700 t/s -> 3150 t/s\] at the tradeoff of increasing perplexity a total of +8% At layer 16, prompt processing speeds up 1.7x \[1700 t/s -> 2900 t/s\] at the tradeoff of increasing perplexity a total of +5% At layer 24, prompt processing speeds up 1.45x \[1700 t/s -> 2500 t/s\] at a tradeoff of increasing perplexity a total of +2.6% This method is different than the other KVA projectors I have seen for the following reasons: My method uses one full map per layer that uses Tikhonov regularization/ridge regression vs a per layer + training correction heads thats applied to 4 streams, then averaged out. This translates to higher accuracy and less perplexity, at the cost of more VRAM. The other methods use apx 400mb while mine uses 1.5ish GB. Other methods predict later layers keys, values, and inputs directly, while my method predicts strictly the inputs to the later layers, and depends upon the models actual weights to compute keys/values. Finally, the other methods I've seen are not variable by which layer implementation starts at (usually locked to 24 i believe), while my method is variable and allows you to determine your own risk tolerance for increasing PP speeds at the cost of increased perplexity. I have a lot of faith that this idea can be expanded and become hugely useful based off my initial indications. In less technical terms, its sort of like MTP for pp instead of tg, \*except it's not lossless\*. The error does get ingested by the model. Models like DSV4.1 and likely MiMo v3 are likely trained alongside this type of implementation, so they may be more tolerant to the ppl increase already. Models that havent been trained against this, like QFN, will continue to see that ppl increase where error occurs. Here's a summary of the BetterBench results. |Metric|Result|Detail| |:-|:-|:-| |Prefill|3,600 t/s|@ 64k tok| |Decode|74.7 t/s|Weighted combined| |Concurrent|70.3 t/s|@ 8 streams (48/48 ok)| |TTFT (P50)|338 ms|Single stream| |Update (P99)|51.5 ms|Stream stutter| BUT WAIT, THERE'S MORE! Here's my 2nd implementation. Based off of the HySparse2 paper, it appears that they are using a similar method but multi-layered instead of single layer. Based off of this, I've built an initial early version of this. Here's what the preliminary results show: Multi-layer - 1.55x speedup at only a +2% of perplexity Using this, I strongly believe that this can be implemented for a total of 1.5x speedup while <1% ppl increase. If anyone wants to adapt this to other engines and models, just note that I found more training to be virtually worthless, it's purely architectural levers that need to move IMO. Currently this is in very early testing- R9V is updating with this capability and this projector is getting uploaded to HF, but it is NOT CONFIRMED STABLE. The V1 iteration of KVA Projectors is however stable. V2 projectors are behind a config flag ( --ced quality) that you can choose if you wish. Here is the HF repo for the full IQ4\_XS QFN Model + MTP + KVA Projector (V1) - this one is directly usable in R9V now. https://huggingface.co/Dyluhn/Qwen3.8-Flash-Next-Uncensored-R9V-IQ4\_XS Here is the repo with just the projectors, V1+V2, with a short explanation on how to get started on implementing this method in other VLLM projects https://huggingface.co/Dyluhn/Qwen3.8-Flash-Next-Uncensored-CED-Projector/tree/main NOTE- R9V is specifically built for 2x R9700 setups with significant host RAM. RAM usage floats around 50ish GB during use for expert storage. You'll likely need an SSD to handle PLE/en-grams at usable speeds, or just keep them in RAM. Enjoy! Join https://discord.gg/launch80 if you are interested in working on and with some of the latest and greatest implementations for RDNA4 (there's even better projects than this one in there) Also, need to acknowledge that https://huggingface.co/kishida was the first one, as far as I can tell, to determine that this feature can be borrowed from DSV4.1 separately from the model architecture. Bravo.

▲
52
+8
2👁
r/LocalLLaMA · u/Mxmtm · 15d ago
Mac Studio M5 Ultra 96GB vs M5 Max 128GB for local LLMs?

I'm about to buy a Mac Studio mainly for running LLMs locally and I'm stuck between two configs: M5 Ultra (30/64) with 96GB: 1.2 TB/s bandwidth, roughly 1.7x faster generation and much faster prefill M5 Max (40-core GPU) with 128GB: 614 GB/s, but 32GB more memory and a bit cheaper The models I care about most right now are Qwen 3.8 27B and Qwen3.8-Flash-Next. The 27B fits easily on both, so the real question is Flash-Next. With the n-gram table offloaded to SSD, it seems to fit on 96GB, but only with the leanest 4-bit builds and very little headroom left for macOS. On 128GB you get more room for better quants, longer context or a second model loaded at the same time. A few questions for those who already made the call: 1. Which one did you go for, and do you regret it? 2. If you're running Flash-Next on a 96GB machine, how's it working in practice? Any issues with memory pressure or long contexts? 3. Is the speed of the Ultra worth giving up the extra memory, or will 96GB feel tight? Thanks!

▲
48
 
1👁
r/LocalLLaMA · u/nickm_27 · 15d ago
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (… · ggml-org/llama.cpp@70c4e15

This has made a massive improvement in performance on my 7900XTX before: `` | model | size | params | backend | ngl | n_ubatch | fa | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | --------------: | -------------------: | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 | 3410.53 ± 22.72 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 | 135.64 ± 0.80 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 @ d8192 | 2477.27 ± 76.78 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 @ d8192 | 120.42 ± 0.46 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 @ d16384 | 2130.36 ± 35.16 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 @ d16384 | 118.42 ± 0.14 | ` after: ` | model | size | params | backend | ngl | n_ubatch | fa | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | --------------: | -------------------: | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 | 4331.21 ± 132.67 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 | 142.76 ± 0.92 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 @ d8192 | 2885.67 ± 80.24 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 @ d8192 | 125.29 ± 0.14 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 @ d16384 | 2402.13 ± 44.16 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 @ d16384 | 120.82 ± 0.18 | ``

▲
37
 
1👁
r/LocalLLaMA · u/crusaderky · 15d ago
MiniMax M3.1 (Space Bunny Alpha) thinks in caveman mode

The CoT of MinimMax M3.1, currently available in openrouter and opencode under the guise of "Space Bunny Alpha", has the familiar look of caveman mode in order to save tokens. This has no impact on the final output. (note: in the first screenshot, pi-caveman is set to off; in the second one I uninstalled it altogether to make sure).

▲
51
-4
9👁
r/LocalLLaMA · u/Balance- · 18d ago
convaiinnovations/laya (multilingual, non-autoregressive System 1 decision model)

Laya is an open-weight (Apache 2.0) "System 1" decision model from Convai Innovations, built by Nandakishor M as an open alternative to TypeSafe's closed Jev API. Instead of generating text, it takes a state (text, an email, a ticket, or JSON) plus typed questions (choice to pick a label, score to place something on an ordinal rubric, and noul for a yes/no probability) and answers all of them in one forward pass in about 33–40 ms on a GPU, so there's no output to parse and nothing to hallucinate. It comes in three checkpoints: a 421M-parameter English model on ModernBERT-large, a faster 322M multilingual model on mmBERT-base covering 100+ languages, and a variant fine-tuned for typed-decisions workflows. A built-in Router detects the input's script and sends it to the right checkpoint. It's trained with RLCD, a reinforcement learning method whose reward uses strictly proper scoring rules, so the model maximizes reward only by reporting honest probabilities; it also has an act-vs-escalate head for deciding when to hand off to a human. The author reports strong results, including beating Jev on AG News, emotion classification, and the typed-decisions benchmark while running roughly 6–8× faster, though the Jev figures are third-party numbers rather than head-to-head runs. The model card is also candid about its limits: the base checkpoints are near chance on typed-decisions without fine-tuning, accuracy drops sharply with 50+ options (Banking77: 0.425 vs. Jev's 0.870), ordinal scoring is its weakest question type, the English checkpoint fails on non-Latin scripts, and the models ship overconfident, so you need to fit a temperature on your own data before trusting the probabilities.

▲
54
-1
5👁
r/LocalLLaMA · u/WonderRico · 15d ago
Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark. post image

I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all runs and here are some noticeable changes: Flash Next is still king, but the benefit of xhigh vs medium reasoning effort is now properly showing. Same for 3.8 27B (however in everyday tasks I personally still prefer using medium)

▲
50
 
6👁
▲
52
-4
4👁
r/LocalLLaMA · u/youcloudsofdoom · 19d ago
One more 'you should try ExllamaV3/exl3 for flash next' appreciation post

After seeing a few posts on here about it, I finally tried exl3 3bpw and exllamav3 for running flash next - with amazing results. On 3x3090s, 128GB DDR4: 1500 prefill, 80 tps decode On 1x5090, 128G. DDR4: 1500 prefill, 29 tps decode Both at 262k context, both with vision/spec decoding. Really impressed, definitely replacing vllm/llama.cpp for me on this model. Quant capacity seems good so far, going to gest the 4bpw later for comparison. Check it out if you were sleeping on it like I was!

▲
866
+5
40👁
r/LocalLLaMA · u/tiensss · 16d ago
Jev isn't new tech. Its marketing targets people who think AI started with LLMs.

I keep seeing Jev presented as some new class of decision model, but most of what’s being advertised is just normal classifier behavior with modern zero-shot capabilities.

It outputs probabilities over constrained choices, doesn’t generate autoregressively, can’t output an invalid class, and can use labels defined at inference time. None of that is new. Zero-shot/NLI classifiers, embedding models, cross-encoders and rerankers have been doing variations of this for years.

The weird part is that most of the impressive Jev comparisons are against LLMs. Of course a specialized classifier is faster and cheaper than making an autoregressive LLM generate an answer. That doesn’t establish a new paradigm. The meaningful comparison is against strong existing classifiers. The purpose of this is to mislead.

There are already benchmarks like BTZSC evaluating dozens of zero-shot classifiers across 22 datasets, including NLI models, embedding models and rerankers. I haven’t seen Jev properly benchmarked across that landscape yet.
(https://proceedings.iclr.cc/paper\_files/paper/2026/hash/417e1c15b3d49852fceded8aa104107d-Abstract-Conference.html)

Where people have compared Jev with conventional classifiers, the story is much less magical. One Banking77 experiment got 93.3% from BGE-small + logistic regression versus 83.2% for Jev, at about 9ms locally.
(https://github.com/ickma2311/jev-baselines-eval)

Some of the marketing also goes into the misleading territory. The “can’t hallucinate” framing is very sus, for example. Their own explanation admits the 0% hallucination figure is not empirical, and what they actually guarantee is that Jev returns an answer matching the allowed schema. That prevents invalid outputs, it does not prevent confidently choosing the wrong valid answer. (https://typesafe.ai/blog/introducing-system-one-models-and-jev)

So color me a skeptic. Look, Jev might even be a good product. Maybe their unpublished architecture or RLCD training method is genuinely novel. But nothing we've seen so far establishes that "System One Models" are a new class of AI. What the public evidence mostly establishes is that using a specialized classifier for classification can be much cheaper and faster than using an autoregressive LLM, which we already knew. It only sounds novel if your idea of AI begins and ends with LLMs.

💬 333 (-1) open on reddit ↗
▲
466
+5
37👁
r/LocalLLaMA · u/returnity · 21d ago
Is HF starting to move against abliterated models?
Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.
The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The scale of the problem is massive: Hugging Face, which hosts open source AI models, currently lists over 6,000 abliterated models.

I can't really tell what exactly the implications are of this "partnership" or what it exactly would impact on HF's model-hosting side. However, I do find it concerning that HF is announcing a collaboration on 'infrastructure safety' with publicity that specifically calls out "dangerous" uncensored models. Thoughts?

💬 213 (-1) open on reddit ↗
▲
302
+3
34👁
r/LocalLLaMA · u/More-Curious816 · 15d ago
Folks, have you purchased the Mac M5 Ultra with 256GB yet? We need serious benchmarks, because we only get YouTube clowns influencers results

Ok, so the Mac M5 Ultra (256GB) hit the market, but the only publicly available benchmarks material are flashy YouTube "clown influencers" videos. We need serious numbers to evaluate whether Apple’s silicon can actually compete with Nvidia’s current GPU‑centric workflows or not. Every damn video I watched, it was from somebody who only know the bare basics and use 8b models, like wtf.

I know it's seriously fucked up price, but following this sub I know some of you already owned it.

💬 236 (-1) open on reddit ↗
▲
1294
+2
35👁
r/LocalLLaMA · u/Acrobatic_Stress1388 · 16d ago
Mods: can we do something about half the forum getting filled with these advertising posts for Jev?

Jev is a paid product that dumped a lot of venture capitol money into shill their product here and in other subreddits. Obvious shill posts are obvious.

💬 242 (-2) open on reddit ↗
▲
713
+9
36👁
r/LocalLLaMA · u/ECrispy · 18d ago
16GB (and in many cases 12GB) is the max vram most people will ever reasonably have

This sub is, needless to say very niche and skewed towards the high end. There are tons of extremely high end setups here with multiple gpu's etc.

Even 24GB is out of reach of most people financially, forget about the 3x3090 or 5090 or even higher setups. Macs/Strix Halo/dgspark etc are all similarly expensive. 16GB is pretty much the high end for most. And this completely changes in most of the rest of the world where even 12GB would be a luxury.

Things have changed recently (I think even last 6 months have been huge) and even agentic coding is now feasible on 16GB cards (eg with Qwen 27B quants).

I think/hope things will continue to improve. Of course there's going to be a hard limit on how much world knowledge these smaller models will have.

The holy grail is new architecture that supercedes the Transformer and new techniques that don't depend on vram/bandwidth.

💬 521 (-2) open on reddit ↗
▲
991
+7
34👁
r/LocalLLaMA · u/Atagor · 16d ago
Pirate Face - pirate bay for LLMs

The title says for itself

In case someone desides to censor huggingface, we'll have an alternative

Edit:

A lot of responses so I'll leave it here:

  1. I'm not the author.
  2. If I were the author I wouldn't use the word "piracy".
  3. If you're the author, please, rename the domain! What is free in the first place must be named as such, we're not pirating anything.
💬 123 (-4) open on reddit ↗