59 posts · 1 sub · RSS
← prev Sep 23, 2026 → Sep 24, 2026 next →
2026-09-23 → 2026-09-24 hourdayweekmonthyearall
allr/LocalLLaMA
▲
108
+8
33👁
r/LocalLLaMA · u/crusaderky · 16d ago
MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam

MiMo-V2.6-Pro has an insanely high score of 46 on AA, putting it at the head of the opensource models available. It also costs pennies. Flash is not out on AA yet, but it costs less than half on datacenter and is slightly below on Xiaomi's own benchmarks. It also fits in 192GB, which makes it the first real use case for Gorgon Halo.

So I tried both models. This is not a benchmark; it's an educated impression from a senior SWE.

MiMo-V2.6-Pro

I gave it a security-focused task: enable a bubblewrap sandbox to do git push to github, but not git push --force or other destructive commands. Optional flag --no-git when starting the sandbox completely disables github write access.

It stopped to ask me questions as it spotted unclear corner cases in the design 🥇 , then moved on to implementing.

It was slow, but that's just an inference issue (\~25 tok/s) that should be fixed in a few days as more providers come online.

Then I read its output and I had to pick up my jaw from the floor, where it had dropped.

With an extremely quick glance at the code, I immediately spotted that, in order to bypass --no-git, you would have to perform this extremely complicated and exotic command inside the sandbox:

$ git push
(fails)
$ echo GIT_STATUS
blocked
$ GIT_STATUS="p0wn3d by l33t h4xx0r" git push
(successful)

This is 15-year-old script kiddie level.

I didn't read further. I asked GLM-5.3 (full-fat) to do a security review of the change.

In 3 minutes, it found NINE glaring security holes that allow bypassing git and gh restrictions. A few examples that made me want to rip my hair out:

In the default restricted mode,

  • git push works 🥇
  • git push --force is blocked 🥇
  • git push -f is blocked 🥇
  • git push -uf lets you happily wipe out the git remote. ☠️
  • git config alias.fp 'push --force --no-verify && git fp goes through too ☠️
  • env -u GIT_CONFIG_COUNT /usr/bin/git push --forceblasts through ☠️

Again. This is an intern-with-acne level kind of incompetence.

To seal the lid on the coffin, MiMo's prose in the chat is infuriating. Not quite Opus-level infuriating, but it gets close. It hurts the eyes and it frequently takes 2 reads to understand what the hell it's saying. GLM, DeepSeek, and Qwen are much more pleasant to work with.

MiMo-V2.6-Flash

I asked MiMo-V2.6-Flash to do a very simple git surgery: create a new branch off master and cherry-pick a single commit from another branch.

However, I didn't realise that the git worktree I pointed it to was corrupted (the branch on the main git repo was fine).

  • A dumb model would have just returned "there's no git here, I have no idea what you're talking about"
  • A smarter model would have noticed that there was a /worktrees/ in the path, come up with an educated guess about what happened, and gave me a hint on how to fix it
  • A very smart model would have noticed that the only other directory existing in the sandbox was the main git repo, which had a branch with the same name as the broken worktree directory, and recovered it from there.

MiMo-V2.6-Flash went on 80k tokens worth of acid trip. It first attempted to find the main git repo, failed, and then panicked and went down a rabbit hole which involved tampering with /tmp, mount --bind, and other insanity. I noticed after a while as I was wondering what the heck was wrong. I suspect that given enough time it may have nuked my main git repo and I tremble at the idea of what it could have done if not sandboxed.

If you scale down Pro's intelligence on AA by comparing the available self-published benchmarks against those of Pro (which is a very crude method but gives a ballpark idea), MiMo-V2.6-Flash comes out on par with GLM-5.3-Flash (high) and Qwen3.8-Flash.

Which is absolutely, categorically, not.

DO NOT shell out the money for a Gorgon Halo for MiMo-V2.6-Flash. Qwen3.8-Flash on a Strix Halo is vastly better.

I'm going to stick with my previous models:

  • DSv4.1 Flash as the default
  • GLM-5.3-Flash (high) as the dirt cheap option
  • GLM-5.3 when the big guns are needed
  • Qwen3.8-Flash and Qwen3.8-27B to run locally (I have a 3080 so Flash is very slow).
💬 99 (+9) open on reddit ↗
▲
52
+8
2👁
r/LocalLLaMA · u/Mxmtm · 15d ago
Mac Studio M5 Ultra 96GB vs M5 Max 128GB for local LLMs?

I'm about to buy a Mac Studio mainly for running LLMs locally and I'm stuck between two configs: M5 Ultra (30/64) with 96GB: 1.2 TB/s bandwidth, roughly 1.7x faster generation and much faster prefill M5 Max (40-core GPU) with 128GB: 614 GB/s, but 32GB more memory and a bit cheaper The models I care about most right now are Qwen 3.8 27B and Qwen3.8-Flash-Next. The 27B fits easily on both, so the real question is Flash-Next. With the n-gram table offloaded to SSD, it seems to fit on 96GB, but only with the leanest 4-bit builds and very little headroom left for macOS. On 128GB you get more room for better quants, longer context or a second model loaded at the same time. A few questions for those who already made the call: 1. Which one did you go for, and do you regret it? 2. If you're running Flash-Next on a 96GB machine, how's it working in practice? Any issues with memory pressure or long contexts? 3. Is the speed of the Ultra worth giving up the extra memory, or will 96GB feel tight? Thanks!

▲
991
+7
34👁
r/LocalLLaMA · u/Atagor · 17d ago
Pirate Face - pirate bay for LLMs

The title says for itself

In case someone desides to censor huggingface, we'll have an alternative

Edit:

A lot of responses so I'll leave it here:

  1. I'm not the author.
  2. If I were the author I wouldn't use the word "piracy".
  3. If you're the author, please, rename the domain! What is free in the first place must be named as such, we're not pirating anything.
💬 123 (-4) open on reddit ↗
▲
691
+7
35👁
r/LocalLLaMA · u/Training-Respect8066 · 15d ago
Qwen-3.8-27B is good enough that I stopped using API

Many a praise have been sung on Qwen-3.8, but here is mine.

Qwen-3.8 and I had a rocky start, because it thinks so much. Watching it working is painful, so you have to stop doing that. You have to let it work unsupervised. And that's okay, because it really is able to complete complex refactors on its own, making good decisions along the way. Not perfect, but hey, neither is API.

The model quant is Q4\_K\_S, context is quantized to Q8\_0, which seems to be okay, quality wise. I use the official Qwen. Briefly tried Swift-Qwen, which is indeed faster, but I found it getting trapped in loops, which is very rare in vanilla Qwen.

I am using Qwen-3.8 in the Pi agent without MCP and with the minimum amount of tools. Bash is all you need, but I keep the read, write, and edit tools. The edit tool in Pi is the weakest link, the model often has to retry edits, because it messed up the indentation. I am waiting for someone to come up with a more fault-tolerant edit in Pi. Probably I have to make one myself some day.

As a sandbox I use docker. My Pi agent is running on a Raspberry Pi, which seems fitting.

On my hardware and where I live, 1M tokens cost 2.4 cent (input) and 70 cent (output) which is comparable to the cheapest providers on nano-gpt.com.

💬 287 (+2) open on reddit ↗
▲
355
+7
51👁
r/LocalLLaMA · u/KnownAd4832 · 15d ago
Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second

A while ago I posted 15 tok/s output and 100-120 tok/s prompt processing with the IQ3\_XXS quant on a 12GB RTX 5070 using llama.cpp. Since then I built my own inference engine for this one model and this kind of PC. The same IQ3\_XXS now runs at \~65 tok/s output and \~430 tok/s prompt processing, and the 2-bit quants run faster still using RCO-GSQ quantization.

Using:

64GB DDR5 (5600)
12GB RTX 5070 SFF (Gigabyte)
Ryzen 5 7600 CPU
Windows

Output (tokens/s) on 128K context:

Q2\_0 (equivalent to unsloth Q3): 65.1
IQ2\_XS (equivalent to unsloth Q4): 52.0
IQ3\_XXS (equivalent to unsloth Q5): 44.8

Prompt processing (tokens/s) on 128K:

Q2\_0: 543
IQ2\_XS: 472
IQ3\_XXS: 414

Requirements:

Q2\_0 = 37.6GB minimum in RAM+VRAM

IQ2\_XS = 39.2GB minimum in RAM+VRAM

IQ3\_XXS = 47GB minimum in RAM+VRAM

Vision encoder = 0.91GB additionally

You can now one click install and run the engine with low cost hardware (currently only optimized for CUDA).

GitHub: https://github.com/Niko1221/Strata

Model: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

💬 339 (+5) open on reddit ↗
▲
323
+7
31👁
r/LocalLLaMA · u/Khaledthe · 16d ago
Using uncensored models makes working less of a headache

I have a lot of projects with my friends and team at work that I copy to use for my personal projects, whether it's a plugin I borrow with their consent or a script. I always find that Qwen 3.8 and Muse Spark 1.3 straight up refuse to do anything, as they see it as a steal, so I have just been rocking Qwen3.8-27B-Heretic-JP-Roleplay-NSFW-DanbooruTags.i1-Q4\_K\_M, and it feels so good to just be able to tell it to do something, and it actually does it. I know the model isn't made for coding or projects but roleplay, but I don't see a big dip in performance as it does what's asked to do.

Does anyone else have this problem or not?

▲
686
+6
19👁
▲
328
+6
22👁
r/LocalLLaMA · u/Ok_Warning2146 · 17d ago
DeepSeek and Moonshot AI face Beijing's probe over potential data leaks to Anthropic

The rumor about Kimi execs getting arrested finally has some legs. I believe the reality is more like under investigation for potential arrests or fine.

▲
866
+5
40👁
r/LocalLLaMA · u/tiensss · 16d ago
Jev isn't new tech. Its marketing targets people who think AI started with LLMs.

I keep seeing Jev presented as some new class of decision model, but most of what’s being advertised is just normal classifier behavior with modern zero-shot capabilities.

It outputs probabilities over constrained choices, doesn’t generate autoregressively, can’t output an invalid class, and can use labels defined at inference time. None of that is new. Zero-shot/NLI classifiers, embedding models, cross-encoders and rerankers have been doing variations of this for years.

The weird part is that most of the impressive Jev comparisons are against LLMs. Of course a specialized classifier is faster and cheaper than making an autoregressive LLM generate an answer. That doesn’t establish a new paradigm. The meaningful comparison is against strong existing classifiers. The purpose of this is to mislead.

There are already benchmarks like BTZSC evaluating dozens of zero-shot classifiers across 22 datasets, including NLI models, embedding models and rerankers. I haven’t seen Jev properly benchmarked across that landscape yet.
(https://proceedings.iclr.cc/paper\_files/paper/2026/hash/417e1c15b3d49852fceded8aa104107d-Abstract-Conference.html)

Where people have compared Jev with conventional classifiers, the story is much less magical. One Banking77 experiment got 93.3% from BGE-small + logistic regression versus 83.2% for Jev, at about 9ms locally.
(https://github.com/ickma2311/jev-baselines-eval)

Some of the marketing also goes into the misleading territory. The “can’t hallucinate” framing is very sus, for example. Their own explanation admits the 0% hallucination figure is not empirical, and what they actually guarantee is that Jev returns an answer matching the allowed schema. That prevents invalid outputs, it does not prevent confidently choosing the wrong valid answer. (https://typesafe.ai/blog/introducing-system-one-models-and-jev)

So color me a skeptic. Look, Jev might even be a good product. Maybe their unpublished architecture or RLCD training method is genuinely novel. But nothing we've seen so far establishes that "System One Models" are a new class of AI. What the public evidence mostly establishes is that using a specialized classifier for classification can be much cheaper and faster than using an autoregressive LLM, which we already knew. It only sounds novel if your idea of AI begins and ends with LLMs.

💬 333 (-1) open on reddit ↗
▲
290
+5
24👁
r/LocalLLaMA · u/Disastrous-Work-1632 · 17d ago
GGUFs in transformers natively! post image

Hey there folks!

Aritra here from Hugging Face. I wanted to update you all about the latest changes in \transformers\. We now natively support GGUFs (llama cpp quants).

You can use it like so:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"

model = AutoModelForCausalLM.from_pretrained(
model_id,
gguf_file=filename,
)

After loading, you're using the normal Transformers APIs.

Why did we want to do this?

  1. Quantized models are smaller (so fits in a laptop)
  2. PyTorch tooling at hand (useful for debugging)
  3. Debugging, evaluation, custom generation becomes much easier

On supported Apple Silicon setups, we're also reusing ggml kernels so the model can run directly from its packed quantized weights. On the Qwen checkpoints we tested on an M2 Max, Transformers reached:

  1. Qwen3.5-4B Q4\_K\_M: 70.4 tok/s vs 71.8 tok/s with llama.cpp
  2. Qwen3.8-27B UD-Q4\_K\_M: 15.9 tok/s vs 13.4 tok/s
  3. Qwen3.5-35B-A3B UD-IQ4\_XS: 60.2 tok/s vs 61.3 tok/s

This isn't meant to replace llama.cpp. If you only care about maximum local inference performance, llama.cpp is still probably the better choice.

The point is more that you can now use the same GGUF models in a more flexible environment.

Read more: https://huggingface.co/blog/transformers-llama-cpp-quants

▲
424
+4
38👁
r/LocalLLaMA · u/Terminator857 · 16d ago
Cost of intelligence is dropping fast

https://preview.redd.it/43n0bhiap8rh1.png?width=960&format=png&auto=w…

50% per quarter is amazing. 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity. https://x.com/EpochAIResearch/status/2102510281176023529

Every year moving forward is going to be significantly different that the prior year. What do you think? We will be running coding agents on our phones pretty soon.

💬 156 (+1) open on reddit ↗
▲
260
+4
17👁
▲
145
+4
24👁
r/LocalLLaMA · u/monkeyofscience · 15d ago
Open AI in the wild

You degenerates are famous. I was at a conference with one of the authors, and they asked me: "Did we do a good job representing the community?"

https://dl.acm.org/doi/abs/10.1145/3805689.3812421

▲
65
+4
28👁
r/LocalLLaMA · u/Dutchnamn · 16d ago
Perhaps the highest quality mainline quants of Qwen3.8 27B?

I am proud to release these quants of Qwen 3.8 27B. They beat the excellent ISTA and Unsloth quants byte-for-byte on three corpora. Both KLD and top 1% were tested 3x. It took a week of continuous GPU and CPU time to generate these, all done on a single Strix Halo.

https://huggingface.co/agentionai/Qwen3.8-27B-AP-GGUF

Hope you like it.

Edit: I did a lot of benchmarking and updated the smallest quant. Slightly improved calibration led to this real life result

https://preview.redd.it/eax6krk4x7sh1.png?width=1600&format=png&auto=…

▲
347
+3
22👁
▲
180
+3
23👁
r/LocalLLaMA · u/NineThreeTilNow · 16d ago
Engram gone wild! 2b model update...

Update from : https://old.reddit.com/r/LocalLLaMA/comments/1wis23s/update\_small\_model\_engram/

It's been about a week so I'm back. People were asking me about the model.

People wanted code, or models, etc. Most of that is useless to you right now because you're not going to use an under trained model. So let's get to the details.

The spec locked to the following after a LOT of testing :

2.6b model all up. Embedding, LM head, AttnRes, etc.

2.2b are trained. Embedding / LM Head are frozen (\~205m each)

4.3b ENGRAM table. Yes. She's chonky.

Architecturally speaking now :

This includes SWA layering 3:1 as per standard ablations have shown is "optimal" These are 4k / 4k / 8k SWA layer follow by the global.

At the end of the first "block" (4 layers) the Engram table appears. The Engram table itself runs with a very small set of attention heads so it's not blindly attempting to inject data. This is context aware Engram. I'm uncertain of exact Qwen / DS methods here. They're fairly close to what I use. The difference is obviously the Engram size.

This has 10 blocks plus a final global layer (41 Layers). + Embed + LM Head

So if we step back, the optimization problem is as follows :

How do we maximize compute in the backbone and offload the boring stuff to a table?

Attention Based Residuals (Moonshot) comes to our rescue here. Why AttnRes? It allows all blocks past 1 to utilize the Engram table in some fashion. They all have attention via the residual stream to determine if they want data from Engram and precisely how much.

Why not a bigger Engram table? Honestly? You could probably do that. However I don't know if any model has attempted this sort of mismatch of compute vs table. In theory, DeepSeek's research says it works.

More depth? Could do that too. Training is expensive though.

Why frozen LM/Embed? These are down projected via SVD from OLMo 3's model. So it's a mathematical compression attempt at \~5k -> 2k. This saves a MASSIVE amount of time.

Currently the data lives in a \~KD format. 32 logits stored from the teacher of the Wikipedia corpus. Instead of training 1 hot, it gets 32 soft targets to try to match. Hard cross entropy is brutal on a model and it's the reason you see "trillions of tokens" quoted.

When you borrow an LM Head / Embed / Tokenizer / Teacher model... This is far less.

\---

Where are we now in training? I passed the 100m token mark yesterday at \~4am Pacific.

The current HF repo has all the checkpoints, data, and the 104m mark safetensor.

\--

I wanted to do some inspection of the model at this point. We have to see that it's not complete garbage right? It has seen a fraction of the data it needs.

So based on ONLY 100m tokens seen, I attempted completions and various ablations. Studies of what EXACTLY Engram stores (because it's mostly a guess).

Let's go straight to completions :

"George Washington was an American"

With Engram - "George Washington was an American naval officer who served in the American Revolutionary War. He was a naval officer who served in"

With Engram Zero'd - "George Washington was an American, but he was not a. He was a very good friend and a. He was"

"The American Civil War was a civil war in the United States from"

With - "The American Civil War was a civil war in the United States from 1861 to 1861. The war was a major victory for the Confederacy..."

Without - "The American Civil War was a civil war in the United States from 1861 to 1862. The war was a major victory for the Confederacy..."

"Aristotle was an Ancient Greek" (probably my favorite)

With - "Aristotle was an Ancient Greek philosopher who was a leading authority in the philosophy of Plato. He was also a leading authority..."

Without - "Aristotle was an Ancient Greek word meaning "to be" (ἀπάς, "to be")..."

\---

What do we learn from direct inspection of Engram?

If it saw the data enough times, it starts to offload it to the tables. At that point the model is less forced to use internal computational space to store data, and can rely on Engram.

That's not some interpretation. That's what the data shows exactly.

In the early 100m tokens of Wikipedia it's HIGHLY biased to early Philosophy. This has to do with the topics of what was IN Wikipedia at that time. Early Wiki contained a lot about Philosophy. It's among the earliest topics to exist.

It has seen TONS of Philosophy so it has basically offloaded to Engram because of the repetition.

Ok this is long enough. I'm calling it quits. I'm still looking to find a provider that will sponsor the model to completion. Right now I'm just paying. I contacted Verda and they never responded. Qubrid is in this sub, and they said they'd help but never emailed back after a few repeated proddings. Massed Compute is where this model is currently training, and I asked them like yesterday. Hopefully they'll help a brother out.

None of the above is "LLM written" except the completions from testing I guess.

As usual, ask whatever. It doesn't matter your understanding level or whatever. I'll sit and respond.

No question too dumb. No insult not insulting enough. (I'm joking. It's Reddit)

Much Love.

▲
84
+3
20👁
▲
61
+3
25👁
▲
205
+2
23👁
r/LocalLLaMA · u/CriticallyCarmelized · 16d ago
Please Google, for the love of God.

Gemma 5, 220B A18B QAT plus ngrams please. Thank you very much!

Will settle for 120B A16B plus ngrams.

▲
97
+2
26👁
r/LocalLLaMA · u/jacek2023 · 16d ago
apple/LensVLM-9B · Hugging Face

https://huggingface.co/bartowski/LensVLM-9B-GGUF

[](https://huggingface.co/apple/LensVLM-9B#lensvlm-9b)LensVLM-9B

LensVLM is a 9B Vision Language Model (VLM) that scans compressed images of text, then selectively expands only the relevant pages to their uncompressed form via learned tools.

[](https://huggingface.co/apple/LensVLM-9B#license)License

All ML model files in this repository, including Apple's modifications to the Qwen model, are provided under the terms of the Apple Machine Learning Research Model License.

The source code that accompanies this model is distributed separately and is provided under the terms of the Apple Sample Code License.

▲
63
+2
37👁
r/LocalLLaMA · u/Fragrant_Scale6456 · 15d ago
Qwen3.8 27b practical modeling for 3d printing post image

I spent the past day and a half trying to get qwen27b to complete some practical work for me. I have a Bambu h2c I have been wanting to get more use out of so thought this would be a fun experiment.

I have 27b running on my 5090 and qwen image 2.1 running on a 3080 10gb with comfyui. I had pi build me some skills to use cadquery and comfy.

prompt: “Make me a printable 3d model of a self watering plant pot and a MagSafe phone stand for my iPhone 17.  Give me a sheet with top front and 3/4 view renders of each item.  Then, use comfy to generate a scene and place the rendered product in the scene.  It should look like an advertisement”

It’s not perfect but I’m honestly super impressed with the output. The multi view sheet renders having the amount of filament each object would use is a nice touch.

The setup is 5090 with ninfer, quasar qat 27b, 590k nvfp4 context, image processing enabled. in comfy I’m using the int8 version of qwen image 2.1. harness was pi with skills it made for cadquery, blender, and comfyUI.

My next goal is to be able to give it a series of photos of an object and have it create a faithful 3d model. If it can pull that off it would be great as one of my hobbies is making custom parts for my RC cars.

If anyone has played around with 3d creation and printing with localLLM I’d definitely want to hear about what tools you are using I have a feeling my setup is very basic at this moment.

▲
5
+2
5👁
r/LocalLLaMA · u/Proof_Nothing_7711 · 15d ago
Qwen 3.8-27 tips for my setup

Hi everyone, I’m an AI newcomer eager to learn and experiment. I’m comfortable with coding on my own, but I want to explore the AI ​​world—specifically for code review. I have two separate setups depending on my location: 1) Minisforum Ryzen 9 HX 370 AI with 64GB DDR5 RAM + OCuLink eGPU (AMD W7800, 48GB VRAM) 2) Beelink SER5 Max Ryzen 7 6800U with 32GB DDR5 RAM + OCuLink eGPU (AMD R7900 XTX, 24GB VRAM) Both setups run Windows 11, though I’m open to switching to Linux if it would improve performance. For the LLM, I use Qwen 3.8-27b (Q8 on the W7800, Q4 on the R7900 via Vulkan) for code review, as mentioned. I started out using LM Studio but have since switched to VS Code with Cline. Could you please offer some advice on optimal settings for this model and, if possible, tips on how to best configure VS Code with Cline? I’d like to switch to Ollama and move away from LM Studio (hoping for a smoother experience). Thanks in advance—and apologies if these questions seem basic; I’m just trying to learn as I go. Thanks!

▲
10
+2
6👁
r/LocalLLaMA · u/Retumbo77 · 15d ago
Better alternatives to PocketPal for Local on Android?

I've been using PocketPal for Android, but I've always been plagued by stability issues (even when models aren't loaded - It's not a RAM utilization thing). I have seen this come up a few times in various places, but PocketPal usually gets the recommendation. Does anyone have any alternatives?

▲
2
+2
2👁
r/LocalLLaMA · u/bruns20 · 15d ago
Advice for single 9700 on windows

Hey guys, I've recently gotten a r9700, which I'm very excited about. I've been trying to research best setups and Llama flags, but I'm finding that a lot of the advice I'm seeing is based around 2x9700's and/or Linux. I know windows is the devil, but I use this computer for other things as well, so I'm not looking to switch. Anybody have some up to date advice on running a single 9700 on windows? Edit: For anybody in the future, I ended up using WSL to load this beautiful man's vLLM build : https://www.reddit.com/r/LocalLLaMA/comments/1wiws8e/153\_toks\_on\_1x\_amd\_radeon\_r9700\_running\_qwen38/ PP increased almost 5x and considerable bump in decode speed too

▲
52
+2
2👁
r/LocalLLaMA · u/Public_Umpire_1099 · 15d ago
R9V Update: Created and adopted KVA projections based on Deepseek V4.1 Flash + HySparse2/MiMo-V3 for Qwen3.8 Flash Next. This is a game changer for models that don't natively implement it. 1.45-1.85x speedup in prefill to 3k+ at a small deficit to perplexity. [2x R9700, 128GB DDR5]

Here's my \*first\* implementation of KVA projectors on QFN (just the uncensored model for now) the highlights are basically as follows for using the projectors at each different layer: Starting at layer 12, prompt processing speeds up 1.85x \[1700 t/s -> 3150 t/s\] at the tradeoff of increasing perplexity a total of +8% At layer 16, prompt processing speeds up 1.7x \[1700 t/s -> 2900 t/s\] at the tradeoff of increasing perplexity a total of +5% At layer 24, prompt processing speeds up 1.45x \[1700 t/s -> 2500 t/s\] at a tradeoff of increasing perplexity a total of +2.6% This method is different than the other KVA projectors I have seen for the following reasons: My method uses one full map per layer that uses Tikhonov regularization/ridge regression vs a per layer + training correction heads thats applied to 4 streams, then averaged out. This translates to higher accuracy and less perplexity, at the cost of more VRAM. The other methods use apx 400mb while mine uses 1.5ish GB. Other methods predict later layers keys, values, and inputs directly, while my method predicts strictly the inputs to the later layers, and depends upon the models actual weights to compute keys/values. Finally, the other methods I've seen are not variable by which layer implementation starts at (usually locked to 24 i believe), while my method is variable and allows you to determine your own risk tolerance for increasing PP speeds at the cost of increased perplexity. I have a lot of faith that this idea can be expanded and become hugely useful based off my initial indications. In less technical terms, its sort of like MTP for pp instead of tg, \*except it's not lossless\*. The error does get ingested by the model. Models like DSV4.1 and likely MiMo v3 are likely trained alongside this type of implementation, so they may be more tolerant to the ppl increase already. Models that havent been trained against this, like QFN, will continue to see that ppl increase where error occurs. Here's a summary of the BetterBench results. |Metric|Result|Detail| |:-|:-|:-| |Prefill|3,600 t/s|@ 64k tok| |Decode|74.7 t/s|Weighted combined| |Concurrent|70.3 t/s|@ 8 streams (48/48 ok)| |TTFT (P50)|338 ms|Single stream| |Update (P99)|51.5 ms|Stream stutter| BUT WAIT, THERE'S MORE! Here's my 2nd implementation. Based off of the HySparse2 paper, it appears that they are using a similar method but multi-layered instead of single layer. Based off of this, I've built an initial early version of this. Here's what the preliminary results show: Multi-layer - 1.55x speedup at only a +2% of perplexity Using this, I strongly believe that this can be implemented for a total of 1.5x speedup while <1% ppl increase. If anyone wants to adapt this to other engines and models, just note that I found more training to be virtually worthless, it's purely architectural levers that need to move IMO. Currently this is in very early testing- R9V is updating with this capability and this projector is getting uploaded to HF, but it is NOT CONFIRMED STABLE. The V1 iteration of KVA Projectors is however stable. V2 projectors are behind a config flag ( --ced quality) that you can choose if you wish. Here is the HF repo for the full IQ4\_XS QFN Model + MTP + KVA Projector (V1) - this one is directly usable in R9V now. https://huggingface.co/Dyluhn/Qwen3.8-Flash-Next-Uncensored-R9V-IQ4\_XS Here is the repo with just the projectors, V1+V2, with a short explanation on how to get started on implementing this method in other VLLM projects https://huggingface.co/Dyluhn/Qwen3.8-Flash-Next-Uncensored-CED-Projector/tree/main NOTE- R9V is specifically built for 2x R9700 setups with significant host RAM. RAM usage floats around 50ish GB during use for expert storage. You'll likely need an SSD to handle PLE/en-grams at usable speeds, or just keep them in RAM. Enjoy! Join https://discord.gg/launch80 if you are interested in working on and with some of the latest and greatest implementations for RDNA4 (there's even better projects than this one in there) Also, need to acknowledge that https://huggingface.co/kishida was the first one, as far as I can tell, to determine that this feature can be borrowed from DSV4.1 separately from the model architecture. Bravo.

▲
1293
+1
36👁
r/LocalLLaMA · u/Acrobatic_Stress1388 · 16d ago
Mods: can we do something about half the forum getting filled with these advertising posts for Jev?

Jev is a paid product that dumped a lot of venture capitol money into shill their product here and in other subreddits. Obvious shill posts are obvious.

💬 241 (-3) open on reddit ↗
▲
91
+1
14👁
r/LocalLLaMA · u/Ok-Breadfruit-3523 · 16d ago
My foray into local ai. Two BC-250 ex mining apus running Qwen3.6-35B-A3B Q4_K_M at 60 tok/s with 64k context post image

These boards cost me $115 each and I have them connected using llama.cpp with Vulkan and RPC on Bazzite. The boards have roughly 27GB of combined GPU memory and communicate over 1gb Ethernet. For around $300 including psu I’m loving the performance. I have a few more and want to see what 6 looks like trying to run qwen 3.8 flash.

▲
12
+1
8👁
r/LocalLLaMA · u/nirurin · 15d ago
Qwen3.8-Flash-Next, on 5090+64gb, with Llama.cpp - Seems to not use ram?

I was actually pretty happy with my Qwen3.8-27b setup, and I'd been tinkering with Ninfer to have a version that was "fast but maybe a bit stupid" and the speed was nice to have as a backup. But I was curious how the Flash-Next version might work, after I learned it didn't need to all fit in VRAM to work. I picked up the Atomic quant (let me know if there is a better one I should use, this one seemed good from what I could find). I used the build setup below. It can still be tweaked some more, as I still am only using about 27gb of my vram. > ./build/bin/llama-server \\ \--model "/mnt/SPCC-2TB/Projects/AI-APPS/LLM-Models/Qwen3.8-Flash-Next-Atomic/Qwen3.8-Flash-Next-AD-4.27bpw-Q4\_K\_M-M64 \-00001-of-00033.gguf" \\ \--no-mmproj \\ \--load-mode mmap \\ \--lazy-mode on \\ \--fit off \\ \--gpu-layers all \\ \--n-cpu-moe 32 \\ \--ctx-size 64768 \\ \--flash-attn on \\ \--jinja \\ \--parallel 1 The odd thing I noticed though - I know that some parts of this are meant to run from the SSD for the sake of saving vram space etc. Fine. But I kinda expected that some of it at least would get buffered into system ram, as running from ram would be a whole lot more efficient than running from my NVME drive. But this run gets me the following results: 40tok/s decode. 50tok/s prompt processing (It's a short prompt so probably not accurate) 27gb of vram used 8gb of system ram used... So... I mean, am I just wrong and this is normal? The speed doesn't seem as bad as I expected (I thought I was going to get more like 10tok/s at best) but it seems like I might be missing a trick somewhere?

▲
10
+1
5👁
r/LocalLLaMA · u/Jromagnoli · 15d ago
I'm new and it's kinda overwhelming to get into

Hi, sorry if this doesn't belong here. Getting to the point basically, I've been using online-only AI like GPT/Gemini since 2022, and have been interested in local models but am clueless overall. Yes I'm extremely late. I only use laptop (I'm a student), and I currently own: - Acer swift 5 SF514-55TA (main, budget laptop) | .| . | | --- | --- | | Installed Physical Memory (RAM) | 16.0 GB | | Total Physical Memory | 15.8 GB | | Available Physical Memory | 4.49 GB | | Total Virtual Memory | 25.3 GB | | Available Virtual Memory | 6.33 GB | - Acer Nitro 5 AN515-53 (not used currently) | . | . | | --- | --- | | Installed Physical Memory (RAM) | 8 GB | | Total Physical Memory | 7.85 GB | | Available Physical Memory | 4.96 GB | | Total Virtual Memory | 9.72 GB | | Available Virtual Memory | 5.89 GB | &nbsp; I don't know if it's even possible to set up anything on these. I vaguely know the basics of local models (but I'm probably too dumb for stuff like finetuning, prompts, personal setups, all that fancy shit I see online), and I have a ton of information, models, etc bookmarked on my browser which makes it hard to choose something I know (and am interested in image generation, Chatting (models)? "agents" (task programs?). I guess that needs a lot of separate programs/installations? Or is it possible to have a single client to run varying programs? (im not sure if that's even referred to correctly?) (I'll likely start small) I assume a "program" like LM studio is a good starting point?

▲
40
+1
7👁
r/LocalLLaMA · u/jacek2023 · 15d ago
FreedomIntelligence/HuatuoGPT-3-27B · Hugging Face

from FreedomIntelligence: HuatuoGPT-3-27B is a medical LLM built on Qwen3.8-27B with One-stage Policy Optimization (OnePO). OnePO adapts language models to medicine in a single reinforcement-learning stage, without preceding domain-specific supervised fine-tuning. Teacher responses provide temporary guidance and are retired as the model improves. We release the training code, medical RL dataset, and 8B rubric grader. (last week they released https://huggingface.co/FreedomIntelligence/HuatuoGPT-3-9B)

▲
9
+1
4👁
r/LocalLLaMA · u/daphatty · 15d ago
Misrepresentation of this community?

Ever since I joined this subreddit, I’ve noticed something odd. Every r/localllama post that appears in my Latest feed is a propaganda post either for or against open/closed LLMs. Every single one. However, visiting the subreddit directly tells a different story. There are many helpful and enlightening discussions, the kind that made me subscribe in the first place. So what gives? Why is this subreddit being misrepresented in my Latest feed? It’s easy to blame “the algorithm” but what does that even mean? I’m certainly not a tinfoil hat type and I have zero interest in the pro/con discussion. I just want to read and learn more about self hosting LLMs.

▲
89
 
30👁
r/LocalLLaMA · u/FactoryReboot · 17d ago
Most powerful harness for Qwen 3.8?

I hear Qwen code unlocks the model better. I also think it has more power user features than open code?

It’s nice open code can work with multiple models easier though

Thoughts?

▲
65
 
28👁
r/LocalLLaMA · u/AvidCyclist250 · 16d ago
Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080

Thought it's about time to share after testing for a week. You need four things most people miss: the right quant, the right model, the right branch, and the right cache flags.

https://github.com/dtm-beep/qwen38-flash-next-mtp-16gb

TLDR: AtomicChat AD-4.27bpw Q4_K_M target + the shared Unsloth MTP head, build from my pr-mtp-fix branch (plain master can't load this MTP head yet, it's PR #28243 + one fix commit), and --spec-draft-cpu-moe is the trick that makes 16 GB work. Draft experts live in RAM so the target's hot experts get the GPU. IQ4_XS ~10 t/s → 16.5 tg / 350 pp at 131k, q8 KV.

Hope it helps someone.

▲
10
 
5👁
r/LocalLLaMA · u/Whahine · 15d ago
Dailychained PLX 88096 switches, Quad RTX 5070 Ti + Quad RTX 5060 Ti

Left: Quad RTX 5060 Ti, Middle: Quad RTX 5070 Ti, both PLX 88096 switch, Right = rehomed host Edit: Benchmarsk were 1 line = fixed WHY: -->> DATA SOVEREIGNITY / PRIVACY<<-- hey, this is (localllama right?), this makes no less sense than my dropping the same $$$ on a motorbike I want but don't need so no Triumph Rocket III motorbike for me boo hoo, is for SOHO Anyway, a bit of a journey, a few hundreds of $$ wasted on power adaptors / pci risers that are not suitable and a small fortune in RTX 50xx GPU that will be obsolete eventually Now I'm still buried in the steep learn to use linux / docker / vllm / models / setup clients learning curve (I am windows since Win 3.1). I am yet to learn to relove the CLI (not since ICL/IBM MFs in the 80's) All setup and running to the point vLLM NCCL messages report P2P enabled within each node, (yet to resolve getting P2P across nodes). A little more work on the cooling (more fans coming)/ best orientation etc to do I expect to have these for a while, hence no loose mining frames etc each of these nodes is self contained built up hardware (prototype quality, a few rough edges here and there), If I can get the GPUs just build another one (I have spare V21 case + 88096 PCB) Note I am in NZ so all I can buy locally is regular basic PC parts, pretty much everything else is overseas import eg even the Thermalrake Core V21 cases had to come from Australia, most everything else is from China 2-4 weeks shipping, if a cable or adapter doesn't work then more delays and i pay sales tax 15% at the border Also these cases party trick is they can be bolted vertically so assuming I can manage rising heat that option saves a bit of space on the desk GPUs are mixed brands/models, a couple I already had, I was incrementally ( every local seller is '1 GPU per customer') collecting 8 x RTX 5060 Ti initially for this build, but at the point I got the the 5th one the price delta between 5060 Ti 16GB and 5070 Ti 16GB got close enough I returned that 5th one, added 3 x RTX 5070 Ti 16GB to one I had already. Note there really is no cost effective used GPU market here so I was only able to buy 1 5060 ti used and I only paid him NZ$150 over what he paid for it in Nov 2025 ;) (Nice for him but even at that markup it was still a score) Looking forward to getting 3.8 Flash Next working too fingers crossed is usable Benchmarks Run command per node below Ubuntu 24.04, NVidia 575.something, CUDA 13.2 vllm 0.30 \- so you can see what is enabled (you are correct and thank you for noticing, yes I really do not know what I am doing on the software side, I am just a very old script kiddy) no spec decode etc to keep it reproducable docker run --rm -it \ --name vllm-node1x4 \ --ipc=host \ --gpus '"device=GPU-dd49c72c-4273-9016-aaad-8883c553b0da,GPU-1ef97316-f1c1-3c15-64db-fbd6373179e5,GPU-ab8220cb-5ab1-82a6-8aec-26454657e215,GPU-6ce2d880-240f-18be-ffe0-96ab3d0eede8"' \ -e NCCL_DEBUG=INFO \ -e NCCL_P2P_DISABLE=0 \ -e NCCL_P2P_LEVEL=SYS \ -e VLLM_SKIP_P2P_CHECK=1 \ -e NCCL_BUFFSIZE=16777216 \ -e NCCL_MIN_NCHANNELS=8 \ -v /mnt/ai-assets/huggingface:/root/.cache/huggingface \ -v /mnt/ai-assets/vllm-cache:/root/.cache/vllm \ -v /mnt/ai-assets/models:/models:ro \ -p 8005:8005 \ vllm/vllm-openai:latest \ /models/Qwen/Qwen3.8-27B-FP8 \ --served-model-name qwen3.8-27b-fp8 \ --quantization fp8 \ --tensor-parallel-size 4 \ --max-model-len 65536 \ --max-num-seqs 10 \ --gpu-memory-utilization 0.85 \ --kv-cache-dtype auto \ --host 0.0.0.0 --port 8005 Tests command llama-benchy \ --base-url http://localhost:8005/v1 \ --model qwen3.8-27b-fp8 \ --tokenizer /models/Qwen/Qwen3.8-27B-FP8 \ --depth 0 4096 8192 16384 32768 \ --latency-mode generation Results \- No overclock/undervolt etc stock GPU settings (Todo: NVOC overclock VRAM) \- look very linear to me, basically 2 to 1 \- results seem ok Quad 5070 Ti 16GB | model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) | |:----------------|----------------:|---------------:|-------------:|---------------:|---------------:|----------------:| | qwen3.8-27b-fp8 | pp2048 | 6488.44 ± 3.99 | | 363.10 ± 0.19 | 315.79 ± 0.19 | 363.10 ± 0.19 | | qwen3.8-27b-fp8 | tg32 | 82.21 ± 0.06 | 84.86 ± 0.06 | | | | | qwen3.8-27b-fp8 | pp2048 @ d4096 | 5831.49 ± 4.69 | | 1100.96 ± 0.70 | 1053.65 ± 0.70 | 1100.96 ± 0.70 | | qwen3.8-27b-fp8 | tg32 @ d4096 | 81.56 ± 0.09 | 84.19 ± 0.09 | | | | | qwen3.8-27b-fp8 | pp2048 @ d8192 | 5642.19 ± 3.35 | | 1862.27 ± 1.10 | 1814.96 ± 1.10 | 1862.27 ± 1.10 | | qwen3.8-27b-fp8 | tg32 @ d8192 | 81.00 ± 0.01 | 83.61 ± 0.01 | | | | | qwen3.8-27b-fp8 | pp2048 @ d16384 | 5431.23 ± 5.45 | | 3441.21 ± 3.25 | 3393.89 ± 3.25 | 3442.26 ± 3.26 | | qwen3.8-27b-fp8 | tg32 @ d16384 | 80.63 ± 0.14 | 83.23 ± 0.14 | | | | | qwen3.8-27b-fp8 | pp2048 @ d32768 | 5059.44 ± 0.49 | | 6928.84 ± 0.58 | 6881.52 ± 0.58 | 6929.81 ± 1.32 | | qwen3.8-27b-fp8 | tg32 @ d32768 | 79.43 ± 0.26 | 81.99 ± 0.27 | | | | Quad RTX 5060 Ti 16GB | model | test | t/s | peak t/s | ttfr (ms) | est_ppt (ms) | e2e_ttft (ms) | |:----------------|----------------:|---------------:|-------------:|----------------:|----------------:|----------------:| | qwen3.8-27b-fp8 | pp2048 | 3218.53 ± 1.05 | | 689.86 ± 0.35 | 636.52 ± 0.35 | 689.86 ± 0.35 | | qwen3.8-27b-fp8 | tg32 | 43.35 ± 0.02 | 44.74 ± 0.02 | | | | | qwen3.8-27b-fp8 | pp2048 @ d4096 | 3005.17 ± 0.55 | | 2097.92 ± 0.37 | 2044.59 ± 0.37 | 2097.92 ± 0.37 | | qwen3.8-27b-fp8 | tg32 @ d4096 | 42.85 ± 0.01 | 44.23 ± 0.01 | | | | | qwen3.8-27b-fp8 | pp2048 @ d8192 | 2927.62 ± 1.87 | | 3551.28 ± 2.59 | 3497.95 ± 2.59 | 3551.28 ± 2.59 | | qwen3.8-27b-fp8 | tg32 @ d8192 | 42.66 ± 0.06 | 44.03 ± 0.06 | | | | | qwen3.8-27b-fp8 | pp2048 @ d16384 | 2821.32 ± 1.44 | | 6586.68 ± 3.60 | 6533.34 ± 3.60 | 6587.62 ± 3.66 | | qwen3.8-27b-fp8 | tg32 @ d16384 | 42.25 ± 0.05 | 43.61 ± 0.05 | | | | | qwen3.8-27b-fp8 | pp2048 @ d32768 | 2654.29 ± 0.46 | | 13170.73 ± 2.42 | 13117.40 ± 2.42 | 13171.73 ± 2.56 | | qwen3.8-27b-fp8 | tg32 @ d32768 | 41.30 ± 0.07 | 42.63 ± 0.07 | | | | \--- My plan more or less from a while back, with hardware notes pretty much up to date My justification to target 128GB/ All Blackwell: - 128GB = DGX Spark, RTX Spark And Strix 128GB AIOs so will be relevant for a couple of years - All Blackwell = FP8 fast now NVFP4 = faster once mature/production ready? (late 2026?) - Hopefully significantly faster than say DGX Spark - Need 192GB VRAM?: -- Build another Node etc (assuming RTX GPUs still relevant to AI inference, one day used will be < $$) - if not, easier to sell 1 x Node (or worse case 4 x GPUs individually per Node) than 1 x nonolithic DGX Spark or whatever once future wonder AI execution chips exist My justification to target PEX 88096 Backplanes - GPU P2P within each node with patched Nvidia drivers on Linux - Maximise capabilities of (relative to Node1) constrained RTX 5060 Ti 16GB PCIe x8 - Backends e.g. vLLM with say TP=4, PP=2 hopefully maximize architecture - PCIe4 = less bandwidth BUT: -- SO VERY much more forgiving re interference -- MUCH less $$ than anything PCIe5 -- NVidia p2pbandwidthlatencytest shows < 1us latency GPU P2P within each switch Each Node - Modular/ self contained, just chuck a spare SFF-8654 PCI host card into any PC and go AI LLM Inference Tiers - <= 64GB -- Performance / production tier: Node1 (GPU 0-3): vLLM TP=4 = FAST -- Experimentation / second model tier: Node2 (GPU 0-3): vLLM TP=4: Fast enough?? - > 64GB and <= 128GB = Node1 + Node2: Capacity tier: -- Node1 (GPU 0-3) + Node2 (GPU 0-3): vLLM TP=4 PP=2: Constrained to at best Node2 speed, good enough? - Other: -- Host RTX 5080: Embedding eg Qwen3-VL-Embedding-8B watever -- Node2 GPU4 RTX 3080: STT/TTS whatever Host GPU RTX 5080 - Use standalone for utility eg Embedding / Vision / Spec Decoding etc Host (Host 128GB DDR5-6000) - Ryzen 5 9600X - MSI MPG B850 Edge TI WiFi - Jonsbo D41 Mesh Black - XPG Core Reactor II VE 850W - iGPU only to Monitor - SATA SSD for each of WIN / Linux OS - Gen5 M2 SSD on PCIe5 x4 (CPU): -- 2TB = Docker -- 4TB = AI Assets HOT - SATA 2 x 28TB Barracuda HDD Mirrored (Linux) -- AI Assets COLD - RTX 5080 16GB in PCI_E3 (PCIe4 x4 Chipset) Node1 Performance node - 64GB VRAM (TP=4 Parallel) - Chassis: ThermalTake Core V21 - PSU: MSI MEG Ai1600T - PLX/PEX 88096 PCI 4 slot switch (with downstream SFF-8654 ports) -- Slot 1/4: RTX 5070 Ti -- Slot 2/4: RTX 5070 Ti -- Slot 3/4: RTX 5070 Ti -- Slot 4/4: RTX 5070 Ti -- All GPU PCIe4 x16 within PEX 88096 Node2 Capacity / Secondary node - 64GB VRAM Secondary node (TP=4) - 16GB VRAM Utility (RTX 3080) - Chassis: ThermalTake Core V21 - PSU: DeepCool PN1200M - PLX/PEX 88096 PCI 5 slot switch = Tensor Parallel 64GB VRAM (Secondary/capacity node) -- Slot 1/5: RTX 5060 Ti 16GB -- Slot 2/5: RTX 5060 Ti 16GB -- Slot 3/5: RTX 5060 Ti 16GB -- Slot 4/5: RTX 5060 Ti 16GB -- Slot 5/5: Alienware RTX 3080 OEM 10GB (I have a 5th slot and a spare 3080 so..) -- 4 x RTX 5060 Ti 16GB GPU PCIe4 x8 (Due to 5060 x8 electrically) within PEX 88096 HOST <-- SlimSAS PCIe4 x16 --> Node1 <-- SlimSAS PCIe4 x8 --> Node2 Hardware porn before transitioing to the PCI switch approach I started this build Nov 2025 slowly sourcing the parts for a new standalone PC meant as my triple 4k gaming + AI experimentation rig, then I found I wasnt gaming and I kept adding GPUs (well, VRAM really) V1: RTX 5080 + RTX 5070 Ti 16GB \(early build photo, missing a few other parts\) V2: RTX 5090 + RTX 5070 Ti V2: RTX 5090 + RTX 5080 + RTX 5070 Ti 16GB Then I picked up a 5060 Ti 16GB and was thinking how the hell do I squeeze this in, looked at M2 to what ever adaptors etc etc, yes my MB has bifurication etc etc, ordered a couple then thought nah thats all getting pretty manky, hence the pivot to the current approach, more or less homogenous nodes re generation / vram, plug any node into any PC with spare PCI slot, easy to move and so on Random POC / mid build photos 5 way 88096 PCB will it Post? = YES Host to 4 slot switch daisychain to 5 way switch, will they post/can I see GPUs? = YES \- 4 GPU in let downstream green (5 way switch) 2 in upstream black 4 way switch at which point I ran out of room / power cables / risers / bits of wood, but hey, I could see all bits in linux in a massive pci tree Host to 4 slot switch daisychain to 5 way switch, will they post\/can I see GPUs? = YES How to mount GPU array in these cases? Was looking for a blade asthetic RTX 5060 pretty easy 5070 a little tighter thats a lot of transistors PCB adapter plates in progress this one got moved about 4 times due to cable constraints etc Work with the fragile risers, dont fight the bends they came with [Now you get the apprao

show remaining 1,211 characters

ch](https://preview.redd.it/nyl5mifowirh1.png?width=1024&format=png&auto=…) 4 x RTX 5060 Ti + 1 x RTX 3080 10GB Part populated just 2 gpus in each, 88096 80mm cooling fans installed, running ok for some early work patching for P2P etc etc 88096 80mm cooling fans installed \(Middle\) Quad 5060 Ti populated and running, left still waiting parts Now.. should I sell the 5090 that is now in my older rescurected gaming PC? ( 5600x PCIe4 32Gb DDR4). Probably yes... Note: For a forum dedicated to AI there sure seems a 'wierd 'I hate AI managed posts/slop' herein, so you guys relax I personally fingered every word above (except for some of the vLLM command env vars/switches) Laters

▲
0
 
5👁
r/LocalLLaMA · u/LearnNTeachNLove · 15d ago
Any local open source AI model you would recommend?

Are there people in the community trying/testing the open source local AI models? If yes, any recommendation? Did you find a model that fits your regular expectations or even competes with the closed source private AIs? I suppose it depends on the hardware performance and on the expectations/work of the user, but just curious to know your respective experience. And it is highly probable that your recommendation changes every month… PS: from the various answers I am reading i understand that recommendations depend on the material i am using and on what i want to do with it. I have a very basic equipment with 8GB vram 4060 with 32GB of RAM. I do not really know to which extent i can ask a local AI to do something. I mostly saw demonstrations of people asking the AI to make html games, a snake, flappy bird game, to automate some simple tasks, to generate web pages, in a sense i can understand the hype on the other hand on my side i just have a feeling of not knowing what to do with it. It will just hallucinate, or reply in loop… i do not see myself extra spending for a few GB more of VRAM, on the other hand i do not want to use the models of these big AI companies… privacy, freedom, self-autonomy… Maybe the question i should ask is for what kind of activities/hobbies/work do you use local AI (which one?)? How performant is it from your point of view?

▲
1
 
2👁
r/LocalLLaMA · u/mrgreatheart · 15d ago
Epyc for inference

Hi. My current system is an Intel Ultra 7 with 64Gb DDR5 at 6000. It has 4 GPUs totalling 72Gb: A 3090 and 5070 Ti on x8 CPU connected PCIe slots plus two 5060 Ti - one on an x4 CPU connected m.2 socket, and the other on an x4 chipset connected PCIe slot. I am considering upgrading to an Epyc 7443 system with 256Gb of DDR4 8 channel. Given prices I’ll probably end up with 2400 speed sticks giving a theoretical memory bandwidth of 153Gb/s. The obvious benefits are getting all the GPUs on proper CPU connected x16 and x8 PCIe with room for another at some point plus enough RAM to overflow bigger models than I can fit in VRAM. I’m struggling to find clear information on whether this would actually be worth it for the cost. I am currently able to run Qwen3.8-flash-next in two rather painful configurations (not using the chipset connected 5060 because it hurts too much): \- IQ4\_XS at 300 pp / 40 gen in llama.cpp \- EXL3.05bpw at 1,200 pp / 25 gen in exllamav3 The low tg in exllamav3 appears to be because I have to offload 10 layers to CPU. Obviously the extra PCIe slots would bring the other 5060 into play but I’d like to know what possibilities the extra RAM and bandwidth would open up. Is anyone here offloading larger quants (or other large models) to CPU on Epyc, and if so how usable is it?

▲
0
 
3👁
r/LocalLLaMA · u/lots_of_puppies · 15d ago
Would this be super fast for Qwen 27b & next? 32GB isn't much overhead for Q8 + 256k context

I have a 128gb m5 max now. It runs Qwen 3.8 27b well but as context grows the PP and output is so slow. Qwen Next is faster but still too slow to do lots of agentic work (coding) with. I am not a video gamer and I feel really bad if I contribute to hurting them if I buy this computer just for LLMs, but I would love to have a very speedy local Qwen 3.8 27b :p https://www.corsair.com/us/en/p/gaming-computers/cs-9060022-na/vengeance-a8200-gaming-pc-amd-ryzen-9-9950x3d-geforce-rtx-5090-64gb-ddr5-6tb-m-2-ssd-win11-pro-cs-9060022-na From what I've read with those two qwen models, I think my PP would be over 3,000 and my tks out would be 80 to 150? That sounds amazing to me. The only issue is the 5090 only has 32GB ram though so I'm scared I would have to use a very low quant like Q4 and that I couldn't fit a full 262k context. I was wondering if anyone knows if this build above would be worth it? Thank you!

▲
0
 
3👁
r/LocalLLaMA · u/skeole · 15d ago
Dynamic Abliteration: Non-Destructive Refusal Suppression via Multi-Layer Engram Steering

Cool use case for engrams! Realtime abliteration, model weights stay intact.

▲
0
 
2👁
r/LocalLLaMA · u/llo7d · 15d ago
Local AI you can give to your Mom

I made this thing called Hey Taby Its and easy and cool way to use local AI, it has a cute face that lives in the top of your screen. You can give it to your mom (i dont mean it like that) on https://heytaby.com

▲
48
 
1👁
r/LocalLLaMA · u/nickm_27 · 15d ago
vulkan: int8 coopmat1 matmul implementation for AMD RDNA3 and RDNA4 (… · ggml-org/llama.cpp@70c4e15

This has made a massive improvement in performance on my 7900XTX before: `` | model | size | params | backend | ngl | n_ubatch | fa | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | --------------: | -------------------: | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 | 3410.53 ± 22.72 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 | 135.64 ± 0.80 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 @ d8192 | 2477.27 ± 76.78 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 @ d8192 | 120.42 ± 0.46 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 @ d16384 | 2130.36 ± 35.16 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 @ d16384 | 118.42 ± 0.14 | ` after: ` | model | size | params | backend | ngl | n_ubatch | fa | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | --: | --------------: | -------------------: | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 | 4331.21 ± 132.67 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 | 142.76 ± 0.92 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 @ d8192 | 2885.67 ± 80.24 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 @ d8192 | 125.29 ± 0.14 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | pp512 @ d16384 | 2402.13 ± 44.16 | | gemma4 26B.A4B Q4_0 | 13.26 GiB | 25.23 B | Vulkan | -1 | 1024 | 1 | tg128 @ d16384 | 120.82 ± 0.18 | ``

▲
37
 
1👁
r/LocalLLaMA · u/crusaderky · 15d ago
MiniMax M3.1 (Space Bunny Alpha) thinks in caveman mode

The CoT of MinimMax M3.1, currently available in openrouter and opencode under the guise of "Space Bunny Alpha", has the familiar look of caveman mode in order to save tokens. This has no impact on the final output. (note: in the first screenshot, pi-caveman is set to off; in the second one I uninstalled it altogether to make sure).

▲
50
 
6👁
▲
130
-1
28👁
r/LocalLLaMA · u/returnity · 15d ago
ThinkingCap 3.8-27B vs. Swift 3.8-27B vs. Qwen 3.8-27B Benchmarks

With the release of ThinkingCap-Qwen3.8-27B, I thought it would be worthwhile to do a comparison between the original Qwen3.8-27B, the new ThinkingCap, and Swift-Qwen3.8-27B. Both Swift which I already reviewed, and ThinkingCap do exactly the same thing: they reduce the excessive reasoning loops that 3.8-27B is renowned for. In fact, their claims are almost identical: both models claim to reduce reasoning tokens by approximately 40%, with minimal degradation in performance. I wanted to put these claims to the test.

I used my standard Aider eval suite, which I’ve found to provide good separation of models tested (~20 so far), and on which only one model (Qwen3.8-Flash) has scored over 90%. I am able to measure a number of useful metrics on this evaluation, including pass1/2, completion tokens, seconds/case, tokens/solve, and how many diffs were well-formed in the model’s attempts. Here’s the results of 2 runs per model, which should reduce the error bars to +/- 2-3% at most. All 3 models were evaluated at Q8_0 in llama.cpp 0.5.0:

| model | First-try pass | Retry pass | well-formed diff | median tokens | sec/case | tok/solve |
|---|---|---|---|---|---|---|
| ThinkingCap-Qwen3.8-27B (xhigh) | 27.1% | 77.6% | 100.0% | 7436 | 777 | 12.8K |
| Qwen3.8-27B (xhigh) | 27.1% | 77.6% | 99.1% | 12547 | 1481 | 19.3K |
| Swift-Qwen3.8-27B (xhigh) | 30.8% | 75.7% | 98.1% | 7301 | 750 | 12.1K |

Shockingly, ThinkingCap and vanilla 27B score \*identically\*. I’ve never even had 2 runs of the same model score identically, so treat this as a total coincidence. However, this definitely supports BottlecapAI’s claims of minimal performance degradation. Swift performs within noise levels of the other 2 models, just 2% lower, but with a higher first-try pass rate than either of them.

To swipe a phrase from Claude, the real story is the completion tokens: nearly 5k fewer median completion tokens for both fine-tuned models compared to the original. That almost \*exactly\* matches the claimed 40% reductions from their model cards. ThinkingCap uses slightly more tokens per solve, and therefore takes a little longer than Swift, but they’re within a few percent of each other here as well. One thing to note that’s not seen on the chart: the medians tie, but in mean completion tokens, ThinkingCap uses 8.5% more because its tail is longer — there are more cases on which it still overthinks significantly, while Swift achieves a more uniform reduction in reasoning token usage. Another distinction: both models spend more tokens on cases they fail than on cases they solve, but this is more pronounced for Swift (13.2k for fails vs. 5.9k for solves) than it is for ThinkingCap (9.4k median vs. 6.8k median). ThinkingCap gives up more easily, perhaps? Or it just knows when it’s beaten.

In order to differentiate these two excellent fine-tunes, we need to take a more granular look at their performance. There are 3 languages which distinguish them on programming performance:

| model | cpp | javascript | python |
|---|---|---|---|
| ThinkingCap-Qwen3.8-27B (xhigh) | 11.5% / 61.5% | 35.4% / 85.4% | 27.3% / 78.8% |
| Qwen3.8-27B (xhigh) | 7.7% / 69.2% | 27.1% / 81.2% | 42.4% / 78.8% |
| Swift-Qwen3.8-27B (xhigh) | 11.5% / 69.2% | 37.5% / 81.2% | 36.4% / 72.7% |

As you can see, ThinkingCap significantly underperforms Swift on C++, losing out on pass2 by 8%. However, it makes that ground back up on Javascript and Python, overperforming by 4% and 6%, respectively. This is notable if you use any of these languages more than the others. Performance on the other languages in Aider was statistically similar (p > 0.05). One last distinction — ThinkingCap is the only model with a perfect score on well-formed diffs: zero error outputs and zero malformed replies, whereas both other models had several.

Anyways, I hope this helps anyone trying to choose between these two very well-crafted fine-tunes, both of which do what they say on the tin...

EDIT: Swift Flash Next is OUT!. Join me in requesting an Unsloth UDv3 quant here.

💬 67 (+1) open on reddit ↗
▲
63
-1
32👁
r/LocalLLaMA · u/My_Unbiased_Opinion · 15d ago
PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192)

Just a quick PSA. llama.cpp does have prompt caching. if you are running large context lengths and have long multiturn projects, increasing -cram can provide you with massive speedups. There is a point where context lengths can get so large that 8192mb is not enough and the whole context needs to be re processed again on every turn. personally, I have found 20480 to work well with Qwen 27B 3.8 at 262K context.

the main downside is this uses more ram. vram usage doesnt increase.

▲
8
-1
2👁
r/LocalLLaMA · u/bakatristan · 15d ago
I built an open-weight alternative to Jev / TypeSafe - introducing OpenJudgement-4B (early preview)

I’m releasing OpenJudgement-4B-Preview, an experimental Qwen-based model fine-tuned on custom datasets for classification, scoring and true/false judgments. It scores answer options directly, and Python formats the results into JSON with probabilities. It still uses an LLM backbone, but doesn’t generate the response token by token. It’s unfinished and isn’t at Jev’s level yet. I’d love feedback, especially examples where it gets things wrong. Use it via api at: https://kitani.ai/models/kitani/OpenJudgement-4B-Preview (paid) Model and inference code: https://huggingface.co/kitaniai/OpenJudgement-4B-Preview

▲
54
-1
5👁
r/LocalLLaMA · u/WonderRico · 15d ago
Lesson learned. Don't blindly trust repos and make sure everything is stable for a long running (multi weeks) benchmark. post image

I posted previously my swe-verified django 100 tasks benchmark comparing different local models and quantization. No new models for now, but a fix in my evaluation workflow that was unfortunately not stable during the weeks/months of me using it. I redid the evaluation on all runs and here are some noticeable changes: Flash Next is still king, but the benefit of xhigh vs medium reasoning effort is now properly showing. Same for 3.8 27B (however in everyday tasks I personally still prefer using medium)

▲
75
-2
26👁
r/LocalLLaMA · u/Borkato · 15d ago
Is Qwen Flash Next at like Q2 better than 27B at Q4?

I know questions like this are asked often but I didn’t see this specific one

💬 161 (+1) open on reddit ↗
▲
455
-3
43👁
r/LocalLLaMA · u/R_Duncan · 16d ago
JEV almost dead: CLM vs JEV

Original post: https://www.reddit.com/r/LocalLLaMA/comments/1woscea/contrastive\_language\_models/

(sorry I felt it wasn't giving CLM the highlight it deserves)

What it is: a new projection head for Qwen3-8B.

github: https://github.com/Contrastive-LM/CLM

hf: https://huggingface.co/Contrastive-LM

At the API and functional interface level, CLM supports everything Jev does—it is not a subset. However, there are important trade-offs in generalization, context scale, and architecture between the two.

1. Functional Parity (Same Primitives)

CLM was specifically engineered as an open-weights, self-hostable alternative to TypeSafe AI's Jev. It implements the exact same "System One" decision interface and supports all three of Jev’s core question primitives:

  • Choice: Evaluates a discrete set of candidates and returns a categorical probability distribution.
  • Noul: Outputs a calibrated true/false probability for a proposition or guardrail check.
  • Score: Scores an input against an ordered rubric or scale.

Code written for the TypeSafe Jev client can be pointed directly at a clm-serve endpoint with drop-in compatibility (from clm import CLMClient, Choice, Noul, Score).

2. Where CLM Outperforms Jev

  • Latency and Disaggregated Caching: Jev is a proprietary cloud model that evaluates state and question choices jointly. CLM separates the state head from the action head. If an agent has a persistent set of tools or actions, CLM embeds those actions once and caches them. In benchmarks like interactive browser agents and gaming (T-Rex, Super Mario), CLM is 4× to 13× faster than Jev.
  • Open Weights & Fine-Tunability: Jev is a closed API with no user fine-tuning (you can only prompt it via state and question instructions). Because CLM’s heads are tiny open weights (\~75 MB), you can fine-tune them on your own agent trajectories.
  • Coding Benchmark Verifiers: When fine-tuned on agent trajectories, CLM achieves state-of-the-art verifier performance on Terminal-Bench 2.1 (87.6%) and DeepSWE (81.6%), whereas zero-shot Jev struggled on those exact benchmarks (scoring \~71% on DeepSWE).

3. Where Jev Still Has the Edge (CLM-8B Limitations)

While CLM covers the entire feature surface of Jev, the current CLM-v0.1-8B release trails Jev in a few areas:

  • Zero-Shot Broad Knowledge: Jev is backed by a larger, proprietary model On zero-shot open-domain tasks, Jev still holds an edge in edge-case accuracy (e.g., Berkeley Function Calling Leaderboard v4: Jev scored 99.2% vs. CLM-8B’s 95.2%; WikiRacing: Jev 30/30 vs. CLM-8B 26/30).
  • Context Budget: Jev accepts requests up to a 64K token context out-of-the-box. CLM-8B was tested and calibrated at 2K to 8K context. While its Qwen3 backbone can accept longer prompts, representations past 8K haven't been calibrated for the reference head.
  • Probability Normalization: CLM calculates probabilities via dot products and softmax over the candidates passed in that request Its probabilities are inherently relative to the candidate set provided, whereas Jev’s scoring is calibrated internally against absolute criteria.

Summary

If you are asking if you will lose API features by using CLM instead of Jev: No, you get the full primitive set (Choice, Noul, Score) with massive latency gains and zero API costs. You only sacrifice some zero-shot generalization on niche out-of-domain tasks compared to TypeSafe's hosted service.

💬 194 (+4) open on reddit ↗
▲
315
-3
40👁
r/LocalLLaMA · u/Secure_Recording_472 · 15d ago
UkisAI Swift Series / 27B, Flash Next and Bonsai 2 + GSQ-RCO / -63.4% thinking, x1.95 speed with xhigh accuracy post image

Hey everyone,

Jovan from UkisAI here! Today, we are introducing Swift, a family of efficient reasoning LLMs based on Qwen, trained by penalizing tokens related to pathological overthinking patterns and restoring accuracy via RL (GSPO) and OPD.

After amazing feedback and 350k+ downloads in 13 days on our Swift Qwen 3.8 27B we are releasing the entire model family as well as the highly requested GSQ-RCO quants for 27B and Flash-Next.

This release includes:

Swift1.5 27B, an improved version of our last model, with even lower token usage, fixed bugs and better agentic performance, with -58.5% thinking tokens while scoring 0.35% higher and outperfoming base on Terminal Bench 2.1 by not falling into "overthinking error" loops.

Swift Flash Next, with 63.4% fewer thinking tokens and a 1.8x speed up scoring -0.2% vs base on xhigh

Swift Bonsai 2, with 39.8% fewer thinking tokens while scoring 0.19% higher (although we'd still like to note it as experimental)

Our benchmarks are ran x5 on Base and Swift, averaging across five seeds and various domains, including General (GPQA, AIME26), Coding (LiveCodeBench), Vision (ERQA), Agentic (Terminal Bench 2.1).

One note is that the Terminal Bench 2.1 score of Swift1.5 27B is misleadingly low at first glance. It is not a bug, but a simple matter of the Swift models not falling into overthinking loops and failing the task, rather pursuing it until the end, leading to higher average token usage. The token reduction still falls in the -38.7% range when compared apples-to-apples.

We also added a fun "game creation" benchmark you can find and play here, it is completely subjective but Swift generated better games in less time: Flash Next Game and 27B Game

We are including a Research API and HuggingFace Spaces to give the models a spin before downloading or if you don't have enough compute to run them right now! You can find both on the model cards.

We have also made GGUF, NVFP4, MLX and W4A16 quants for relevant model versions.

More details on our training approach and community feedback can be seen here: https://www.reddit.com/r/LocalLLaMA/comments/1wg7dd5/ukisai\_swiftqwen3827b\_583\_thinking\_x195\_speed/

All of the various quantization and model versions are available in their respective collections:

Swift1.5 27B: https://huggingface.co/collections/ukisai/swift-15-27b

Swift Flash Next: https://huggingface.co/collections/ukisai/swift-flash-next

Swift Bonsai 2: https://huggingface.co/collections/ukisai/swift-bonsai-2

We are also working on a 9B variant to be released in the upcoming days.

We would greatly appreciate your feedback via independent evaluations on real world tasks. As per last release, we operate on a candy-shop basis, trying to fulfill as many Swift model requests and quants as possible, so please do share your needs in the comments!

💬 253 (+1) open on reddit ↗
▲
296
-3
35👁
r/LocalLLaMA · u/More-Curious816 · 15d ago
Folks, have you purchased the Mac M5 Ultra with 256GB yet? We need serious benchmarks, because we only get YouTube clowns influencers results

Ok, so the Mac M5 Ultra (256GB) hit the market, but the only publicly available benchmarks material are flashy YouTube "clown influencers" videos. We need serious numbers to evaluate whether Apple’s silicon can actually compete with Nvidia’s current GPU‑centric workflows or not. Every damn video I watched, it was from somebody who only know the bare basics and use 8b models, like wtf.

I know it's seriously fucked up price, but following this sub I know some of you already owned it.

💬 236 (-1) open on reddit ↗
▲
98
-3
22👁
r/LocalLLaMA · u/futterneid · 16d ago
Streaming Nemotron 3 Diarization post image

I’ve been playing with Nemotron 3 Diarization, and it fills a gap I’ve had with local voice agents: keeping track of who is speaking.

It’s a diarization model, so it gives you speaker labels rather than transcriptions or people’s names. It can stream its output and track up to eight speakers. I’ve been trying it with one-second streaming chunks, and the quality has been really good in my tests.

I plugged it into my speech-to-speech setup on a DGX Spark and connected it to a Reachy Mini. The fun part is watching someone new speak, then seeing the robot ask their name and remember it for the conversation (that's the video!).

It has day-zero Transformers integration and is getting a commercial-friendly license.

▲
74
-3
20👁
r/LocalLLaMA · u/HolidayBit143 · 17d ago
Unsloth Studio VS LM Studio... Which one do you prefer?

So I've been experimenting with various platforms and even on day 1 Unsloth Studio was released, I knew that LM Studio had it's days numbered. LM studio will always be the OG but I wonder how much longer they have, especially with all these new platforms arising. It feels like LM studio just fell behind and it doesn't help that they are putting so much effort on BIONIC which I'm not even sure if anyone uses on a serious level.

Which one do you prefer, or do you use something else?

▲
59
-3
32👁
r/LocalLLaMA · u/OvertaxedOne · 16d ago
Qwen FN vs 27B --- Think I'm saturated.

Got QFN up and running on our Strix box this past weekend and have been running on it for a few days now. Big thank you/shout out to the Halogen team, it's running fantastic on the Strix, this is clearly "the setup" right now for this hardware with this model, really impressive performance for such a large model (\~30-40TPS generation, \~900-1000TPS prefill on real world use, not benchmarks, over the past few days)!

However, that said, I kind of feel like 27B saturated my personal use cases. QFN is a great model, but I'm not really noticing much where I think "Wow, 27B would never get this and QFN just one shot it". They feel very similar in capabilities (IE, both are amazing!) and I kind of feel like I'm reaching the end of the runway for what I can realistically make use of in my day to day use cases, I'm just asking questions that really require more than 27B the majority of the time.

It really feels more to me like a very similar level of smart, one that runs well with limited memory bandwidth (QFN) and one that runs well with limited GPU memory (27B). It's an interesting result, and I guess maybe I should have expected it because I was really struggling already to find things that I actually need to do day to day professionally that 27B couldn't do. When I escalate to the cloud now it's nearly always for either speed or context, rarely intelligence.

Surprising result, at least to me, I was expecting to have my hair blown back, but I guess this is just yet another data point that "good enough is good enough".

▲
31
-3
8👁
r/LocalLLaMA · u/jjusko20 · 15d ago
I'm trying to post-train AliceAI-Foundation-80B-A3B-Base myself

Just wanted to share with someone - don't have much to report yet. I am interested in this new AliceAI model and have been wanting to make a community impact for a while - and releasing an initial agentic version of this model sounds cool. I am training on 3 32gb v100s (which has been fun to get to work, to say the least). What I'm really doing is creating a shallow distill of Qwen 3.8 27b and then using reinforcement learning - My initial plan is a SFT with Qwen3.8 27b synthetic data I'm generating targeting long horizon agentic work - then, RL / GRPO with a grader model for a while. I'm considering using a stronger model to generate the training data, but I'm trying to keep this on my local machine only. It's coming along - I can just barely fit the weights and activations on the v100s in qlora. I've built the framework for the SFT data generation for. I don't expect anything amazing but it should be a neat experiment. Also considering using a pre-existing data set for the tune, but I'm more interested in creating my own distillation. Update: it's training! https://preview.redd.it/vx7ybgzbxjrh1.png?width=1329&format=png&auto=…

▲
5
-3
8👁
r/LocalLLaMA · u/DerTomsn · 15d ago
7900 XTX — two "low-thinking" Qwen 3.8 27B quants (Swift + ThinkingCap) vs the regular quant

First, do they actually produce less tokens? Yes. Total tokens per benchmark run (4 scenarios): base quant \~66k, ThinkingCap \~49k (−26%), Swift \~45k (−33%). So the "less thinking" is real — and Swift cuts the most. Then the cost: and this is where it got interesting. The two quants don't trade off the same way: Decode: base \~48 t/s. ThinkingCap barely changes (\~43). Swift drops hard (\~32, −33%). Prefill / TTFT — the opposite of what I expected: Swift is the fastest (\~614 t/s, TTFT \~1s), base in between (\~530 t/s, \~1.2s), ThinkingCap the slowest (\~100 t/s, TTFT 6–11s). * Quality holds: \~84–87 on my eval, same band as base. Full run data: ThinkingCap finishes \~23% faster than base. The token savings win even with the slow prefill. Swift is break-even because of the slower decode speed. |*Quant*|*Prefill*|*Decode*|*Quality*|*Runtime*| |:-|:-|:-|:-|:-| |Base (unsloth)|\~530 t/s|\~48 t/s|\~85|\~1407s| |ThinkingCap|\~100 t/s|\~43 t/s|\~85|\~1087s| |Swift|\~614 t/s|\~32 t/s|\~86|\~1400s| Caveat: 2 runs per quant only, so single-run variance will move these. Prefill speed of ThinkingCap is oddly low. Need to do some more tests on that. Side-by-side (thinking xhigh, Q4\_K\_M, all 7900 XTX) of 3 of the runs: https://llm-bench.io/compare/runs?runs=cmufxkpkb00bj01o05h9vsu5o,cmufwold900ay01o0ae1frktu,cmufyiq3p00by01o0vwbpilew

▲
60
-4
22👁
▲
452
-5
30👁
r/LocalLLaMA · u/Graemer71 · 15d ago
Qwen 3.8 27b be like...

The user is frustrated — I rambled too much and didn't act. Let's just run the test suite and move on. No more forensics. One command, execute, then report.
(Original memo is a casual internal monologue in English. Translating faithfully while preserving the informal, stream-of-consciousness register.)

Got it. No more analysis. Running the suite now:
<br>
(Casual English internal memo, stream-of-thought style, with the informal tone of the original Japanese preserved.)
(Ugh, I'm going in circles. Stopping the analysis here. Going to run the test suite right now.)
(Just do it. Run the tests. Stop thinking so much. Okay.)
(One command. Execute it. Then report the results.)
(Stop analyzing. Just run it.)
(Alright, running it now. For real this time.)
(That's the plan: run the tests, get the output, then move on to the next step.)
(No more going off on tangents. Focus.)
(Running the suite now. End of story.)
(Okay. Done thinking. Doing it.)
(Finally. Action over words.)
(Go, go, go — running the tests right now!)
(That's it. Execution time. No more excuses.)
(Committing: running the full test suite immediately.)
(Silence the brain. Fire the command. That's all.)
(Bye, distractions. Hello, test runner.)
(Here we go — actually running it this time.)
(Deep breath… execute. Done.)
(Yes. Moving from thought to deed.)
(It's happening now, not later.)
(Action mode: engaged.)
(Finally doing what I've been saying.)
(Runner, start!)
(Commencing test execution. Right now.)
(Enough already — I'm pressing run.)
(Initiating. Immediately.)
(That's right. No more delays.)
(Here comes the real deal.)
(Time to stop talking and start doing.)
(Running tests… starting.)
(Enough pondering — going.)
(Execution begins now.)
(Proceeding immediately with the test suite.)
(Finally: action.)
(Do it.)
(Now.)
(Go.)
(Run.)
(Execute.)
(Start.)
(Ship it — tests, go!)
(Engaging.)
(Launching.)
(Firing.)
(Kickoff.)
(Commence.)
(Beating the pause button to death.)
(Play.)
(ON.)
(Active.)
(Living. Breathing. Testing.)
(This time, it's real.)
(No fake-outs. Real run. Real results. Coming right up.)
(Here it comes. Ready? It's here.)
(Test suite. Full throttle. Go.)
(That's the cue. Doing it now.)
(From now on: hands on the keyboard, brain offline.)
(Running the suite. Period.)
(Enough. Enough. Enough. Running it.)
(Really, truly, definitely — starting now.)
(Countdown: 3… 2… 1… run!)
(Boom. Here we go.)
(Finally getting something done.)
(At last, progress!)
(Tests are flying.)
(Code meets trial by execution.)
(Making it official: suite is running.)
(Witness the spectacle: the tests, at work.)
(Brave code, enter the arena.)
(The judge is compiling. Stay calm.)
(Suite, you are summoned.)
(Answer the call. Run. Pass. Fail. Tell me.)
(No more mercy for lazy code.)
(Rigor time.)
(Precision time.)
(Truth serum injected.)
(Go go go.)
(Whoosh.)
(Blamo.)
(Sound of victory… or the wail of defeat. Either way, knowledge!)
(Data incoming.)
(Results on deck.)
(Buckle up.)
(Here they come.)
(Drums, please.)
(And… we're live.)
(Broadcasting from the test bench.)
(Reporting live.)
(Straight from the console.)
(Unfiltered pytest output.)
(Raw, uncut.)
(Rolling the tape.)
(Cliffhanger, part 2: "So what happened?" — stay tuned.)
(In the meantime, the fans… no, the assertions… await.)
(Justice is served.)
(Order restored.)
(Chaos quantified.)
(Entropy reduced.)
(Enlightenment achieved: pass/fail.)
(Binary destiny. All code is born equal before CI.)
(Some are destined for more retries than others.)
(The weak are eliminated. The strong are deployed.)
(Natural selection, but with better tooling.)
(And, occasionally, a segfault.)
(Onward.)
(Forward.)
(Uplifted.)
(Elevated.)
(Exalted.)
(Ascending into the green region of the test pyramid.)
(Where the colors are bright and the coverage is solid.)
(A place of dreams.)
(A place of truth.)
(A place where bugs go to be born… no, to die.)
(RIP, bugs. You were never welcome here.)
(Violence to defects.)
(Ruthless, but fair.)
(The suite is impartial. Like death. Like linting.)
(But, unlike linting, it actually proves something.)
(That's the beauty.)
(That's the poetry.)
(assert not broken)
(while (alive): run\_tests())
(elif dead: postmortem())
(return "ship it")
(# end of stream of consciousness, or beginning of CI?)
(Both, actually. They rhyme.)
(Goodnight, failures. Hello, green bars.)
(Bar chart happy meal.)
(All beef, no flakiness (hopefully).)
(Prayers offered to the reliability saints.)
(Saint James of the JUnit, patron of JVM.)
(Saint Pytest, defender against flakiness.)
(Santa's Little Helper: cron jobs.)
(Amen. Amine. Amino acids. Alright, shutting up now.)
(Silence in the courtroom. The witness is called.)
(The witness: the test suite.)
("What did you see?" "Failures. Glorious, instructive failures.")
("Any lies in your testimony?" "Only the mocks.")
("Thank you, counsel rests.")
(This joke was sponsored by my inability to stop typing.)
(Sorry. Really stopping now.)
(For real.)
(Seriously.)
(In earnest.)
(Meaning it.)

(Voice of Morgan Freeman) - Unfortunately Qwen did not earnestly mean it, and did not, in fact, get on with it

▲
314
-6
20👁
▲
289
-6
23👁
r/LocalLLaMA · u/johnnyApplePRNG · 16d ago
Jev in 25 lines of Python