38 posts · 1 sub · RSS
← prev Sep 22, 2026 → Sep 23, 2026 next →
2026-09-22 → 2026-09-23 hourdayweekmonthyearall
allr/LocalLLaMA
▲
109
+9
36👁
r/LocalLLaMA · u/crusaderky · 17d ago
MiMo-V2.6 (both Pro and Flash) is a benchmaxxed scam

MiMo-V2.6-Pro has an insanely high score of 46 on AA, putting it at the head of the opensource models available. It also costs pennies. Flash is not out on AA yet, but it costs less than half on datacenter and is slightly below on Xiaomi's own benchmarks. It also fits in 192GB, which makes it the first real use case for Gorgon Halo.

So I tried both models. This is not a benchmark; it's an educated impression from a senior SWE.

MiMo-V2.6-Pro

I gave it a security-focused task: enable a bubblewrap sandbox to do git push to github, but not git push --force or other destructive commands. Optional flag --no-git when starting the sandbox completely disables github write access.

It stopped to ask me questions as it spotted unclear corner cases in the design 🥇 , then moved on to implementing.

It was slow, but that's just an inference issue (\~25 tok/s) that should be fixed in a few days as more providers come online.

Then I read its output and I had to pick up my jaw from the floor, where it had dropped.

With an extremely quick glance at the code, I immediately spotted that, in order to bypass --no-git, you would have to perform this extremely complicated and exotic command inside the sandbox:

$ git push
(fails)
$ echo GIT_STATUS
blocked
$ GIT_STATUS="p0wn3d by l33t h4xx0r" git push
(successful)

This is 15-year-old script kiddie level.

I didn't read further. I asked GLM-5.3 (full-fat) to do a security review of the change.

In 3 minutes, it found NINE glaring security holes that allow bypassing git and gh restrictions. A few examples that made me want to rip my hair out:

In the default restricted mode,

  • git push works 🥇
  • git push --force is blocked 🥇
  • git push -f is blocked 🥇
  • git push -uf lets you happily wipe out the git remote. ☠️
  • git config alias.fp 'push --force --no-verify && git fp goes through too ☠️
  • env -u GIT_CONFIG_COUNT /usr/bin/git push --forceblasts through ☠️

Again. This is an intern-with-acne level kind of incompetence.

To seal the lid on the coffin, MiMo's prose in the chat is infuriating. Not quite Opus-level infuriating, but it gets close. It hurts the eyes and it frequently takes 2 reads to understand what the hell it's saying. GLM, DeepSeek, and Qwen are much more pleasant to work with.

MiMo-V2.6-Flash

I asked MiMo-V2.6-Flash to do a very simple git surgery: create a new branch off master and cherry-pick a single commit from another branch.

However, I didn't realise that the git worktree I pointed it to was corrupted (the branch on the main git repo was fine).

  • A dumb model would have just returned "there's no git here, I have no idea what you're talking about"
  • A smarter model would have noticed that there was a /worktrees/ in the path, come up with an educated guess about what happened, and gave me a hint on how to fix it
  • A very smart model would have noticed that the only other directory existing in the sandbox was the main git repo, which had a branch with the same name as the broken worktree directory, and recovered it from there.

MiMo-V2.6-Flash went on 80k tokens worth of acid trip. It first attempted to find the main git repo, failed, and then panicked and went down a rabbit hole which involved tampering with /tmp, mount --bind, and other insanity. I noticed after a while as I was wondering what the heck was wrong. I suspect that given enough time it may have nuked my main git repo and I tremble at the idea of what it could have done if not sandboxed.

If you scale down Pro's intelligence on AA by comparing the available self-published benchmarks against those of Pro (which is a very crude method but gives a ballpark idea), MiMo-V2.6-Flash comes out on par with GLM-5.3-Flash (high) and Qwen3.8-Flash.

Which is absolutely, categorically, not.

DO NOT shell out the money for a Gorgon Halo for MiMo-V2.6-Flash. Qwen3.8-Flash on a Strix Halo is vastly better.

I'm going to stick with my previous models:

  • DSv4.1 Flash as the default
  • GLM-5.3-Flash (high) as the dirt cheap option
  • GLM-5.3 when the big guns are needed
  • Qwen3.8-Flash and Qwen3.8-27B to run locally (I have a 3080 so Flash is very slow).
💬 99 (+9) open on reddit ↗
▲
2109
+15
54👁
r/LocalLLaMA · u/Salah_H_Hasan · 18d ago
Qwen 4 Announced at Apsara Conference

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

💬 568 (+3) open on reddit ↗
▲
424
+4
38👁
r/LocalLLaMA · u/Terminator857 · 17d ago
Cost of intelligence is dropping fast

https://preview.redd.it/43n0bhiap8rh1.png?width=960&format=png&auto=w…

50% per quarter is amazing. 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity. https://x.com/EpochAIResearch/status/2102510281176023529

Every year moving forward is going to be significantly different that the prior year. What do you think? We will be running coding agents on our phones pretty soon.

💬 156 (+1) open on reddit ↗
▲
133
-1
25👁
r/LocalLLaMA · u/No_Algae1753 · 18d ago
Am I going insane for thinking that these are all Bot comments?

I know that there are some bots active on this sub but wow these comments really look AI generated. Is it just me or are those really bot comments?

https://preview.redd.it/20st4vivq3rh1.png?width=955&format=png&auto=w…

source:

https://www.reddit.com/r/LocalLLaMA/comments/1wn9a2w/what\_underrated\_ai\_tools\_have\_actually\_made\_you/

💬 148 (+1) open on reddit ↗
▲
120
+3
26👁
r/LocalLLaMA · u/Melted_gun · 18d ago
What underrated AI tools have actually made you more productive in 2026?

I asked this back in 2025, but the AI landscape has changed a lot since then.

Not looking for the usual ChatGPT, Claude, Gemini, Midjourney, etc. I'm curious about the lesser-known tools that you actually kept using.

Could be for research, coding, design, video, writing, automation, planning, journaling, local AI, or even something oddly specific.

Free or paid doesn't matter.

What tool genuinely saved you time or improved your workflow this year? And what do you actually use it for?

💬 106 (+1) open on reddit ↗
▲
686
+6
19👁
▲
512
+5
21👁
▲
431
-2
18👁
r/LocalLLaMA · u/Sitkin_Marrel · 18d ago
New 6B image model coming, AntLing just open sourced the Ming-Image-0.1-Design family post image

• Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard. https://huggingface.co/inclusionAI/Ming-Image-0.1-Design https://huggingface.co/inclusionAI/Ming-Image-0.1-Design-Layer

▲
347
+3
22👁
▲
343
+3
28👁
r/LocalLLaMA · u/ironicstatistic · 18d ago
Ngram and world knowledge - why are we just building a coding model?

This post is written by a human and I'd appreciate it if you treated it as such. Thanks.

So, I've been noticing a pretty clear interest in developing as good a coding and agentic tool-calling model as possible, especially at smaller sizes, sub-50 gigs. However, I'm finding that at least for my use of AI, if I really want to move away from big providers, I am going to require a model that has better world knowledge than the current offerings.

Qwen 3.8 27B is a truly fantastic model for tons and tons of stuff. It's highly intelligent, super good at designing applications and coding and working on my system. However, its world knowledge sucks​ compared to the frontier, especially at the Q4 quant that I have to run it at.

So, that leaves me with a question. With the new N-gram technology that we're seeing being baked into Qwen 3.8 Next and that presumably will run on future models, why can't a model be made that has a smaller set of intellectual capabilities but a greater amount of world knowledge? I understand that right now everyone is optimizing towards making as smart a model as possible fit into as small a space as possible. But why don't we leverage the SSD to give the model a lot of world knowledge and make models that are better at dealing with screenshots, multilingual capabilities, doing things like pixel art or answering physics questions?

ngram seems like the answer to the "can't fit in vram" question... Qwen 3.8 Next really opens my mind to the possibility that there could be a totally different and better paradigm for how these models are developed, at least for many use cases. Having a relatively smart model with a large amount of world knowledge might be better than having as smart a model as possible...

Not to mention that this would mean that a model's training cut off would become less relevant, because it could just be fashioned a new ngram.

Obviously, the main interest is in creating a model that can code as well as possible because that's what'll capture market share. But I am curious if there are any efforts into this kind of thing or if anybody has an idea on why these things aren't done more often.

Please tell me why I'm wrong, how I'm wrong, and in how many ways I'm wrong because I'm sure that that's all you really want to tell me, but at least I'll learn something, because as is obvious from this post I have no idea what I'm talking about.

Thanks have a good day :)

edit: I found this post and I guess it provides a lot of what I was asking:

https://www.reddit.com/r/LocalLLaMA/comments/1vzgtqf/ngram\_vs\_experts\_explained/

edit2:

this one is even better, recconend reading. ty reddit suggestions:

https://www.reddit.com/r/LocalLLaMA/comments/1w0198r/no\_engrams\_wont\_let\_you\_run\_1t\_models\_locally\_it/

▲
328
+6
22👁
r/LocalLLaMA · u/Ok_Warning2146 · 17d ago
DeepSeek and Moonshot AI face Beijing's probe over potential data leaks to Anthropic

The rumor about Kimi execs getting arrested finally has some legs. I believe the reality is more like under investigation for potential arrests or fine.

▲
314
-6
20👁
▲
289
-6
23👁
r/LocalLLaMA · u/johnnyApplePRNG · 17d ago
Jev in 25 lines of Python
▲
283
-2
25👁
r/LocalLLaMA · u/Disastrous-Work-1632 · 17d ago
GGUFs in transformers natively! post image

Hey there folks!

Aritra here from Hugging Face. I wanted to update you all about the latest changes in \transformers\. We now natively support GGUFs (llama cpp quants).

You can use it like so:

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "unsloth/Qwen3.5-4B-GGUF"
filename = "Qwen3.5-4B-Q4_K_M.gguf"

model = AutoModelForCausalLM.from_pretrained(
model_id,
gguf_file=filename,
)

After loading, you're using the normal Transformers APIs.

Why did we want to do this?

  1. Quantized models are smaller (so fits in a laptop)
  2. PyTorch tooling at hand (useful for debugging)
  3. Debugging, evaluation, custom generation becomes much easier

On supported Apple Silicon setups, we're also reusing ggml kernels so the model can run directly from its packed quantized weights. On the Qwen checkpoints we tested on an M2 Max, Transformers reached:

  1. Qwen3.5-4B Q4\_K\_M: 70.4 tok/s vs 71.8 tok/s with llama.cpp
  2. Qwen3.8-27B UD-Q4\_K\_M: 15.9 tok/s vs 13.4 tok/s
  3. Qwen3.5-35B-A3B UD-IQ4\_XS: 60.2 tok/s vs 61.3 tok/s

This isn't meant to replace llama.cpp. If you only care about maximum local inference performance, llama.cpp is still probably the better choice.

The point is more that you can now use the same GGUF models in a more flexible environment.

Read more: https://huggingface.co/blog/transformers-llama-cpp-quants

▲
249
+4
17👁
r/LocalLLaMA · u/108er · 18d ago
Qwen image 2.1 (Fast FP8) generates premium quality images post image

Don't know how they did it, but for under 10GB model, the results are astonishing. I am running it on Unsloth Studio. They just released the update, so if you are not seeing the option, I recommend updating your Unsloth Studio. Cheers!

▲
205
+2
23👁
r/LocalLLaMA · u/CriticallyCarmelized · 17d ago
Please Google, for the love of God.

Gemma 5, 220B A18B QAT plus ngrams please. Thank you very much!

Will settle for 120B A16B plus ngrams.

▲
194
-3
22👁
r/LocalLLaMA · u/simpleuserhere · 18d ago
Laya model playing Flappy Bird on a CPU using OpenVINO INT8 inference post image

A 421M-parameter model just played Flappy Bird on my desktop CPU (OpenVINO int8)

Running on my Intel Core i7 12th gen CPU

Converted laya system one model to OpenVINO and quantized to int8

Repo : https://github.com/rupeshs/flappy-laya-openvino-cpu

▲
180
+3
23👁
r/LocalLLaMA · u/NineThreeTilNow · 17d ago
Engram gone wild! 2b model update...

Update from : https://old.reddit.com/r/LocalLLaMA/comments/1wis23s/update\_small\_model\_engram/

It's been about a week so I'm back. People were asking me about the model.

People wanted code, or models, etc. Most of that is useless to you right now because you're not going to use an under trained model. So let's get to the details.

The spec locked to the following after a LOT of testing :

2.6b model all up. Embedding, LM head, AttnRes, etc.

2.2b are trained. Embedding / LM Head are frozen (\~205m each)

4.3b ENGRAM table. Yes. She's chonky.

Architecturally speaking now :

This includes SWA layering 3:1 as per standard ablations have shown is "optimal" These are 4k / 4k / 8k SWA layer follow by the global.

At the end of the first "block" (4 layers) the Engram table appears. The Engram table itself runs with a very small set of attention heads so it's not blindly attempting to inject data. This is context aware Engram. I'm uncertain of exact Qwen / DS methods here. They're fairly close to what I use. The difference is obviously the Engram size.

This has 10 blocks plus a final global layer (41 Layers). + Embed + LM Head

So if we step back, the optimization problem is as follows :

How do we maximize compute in the backbone and offload the boring stuff to a table?

Attention Based Residuals (Moonshot) comes to our rescue here. Why AttnRes? It allows all blocks past 1 to utilize the Engram table in some fashion. They all have attention via the residual stream to determine if they want data from Engram and precisely how much.

Why not a bigger Engram table? Honestly? You could probably do that. However I don't know if any model has attempted this sort of mismatch of compute vs table. In theory, DeepSeek's research says it works.

More depth? Could do that too. Training is expensive though.

Why frozen LM/Embed? These are down projected via SVD from OLMo 3's model. So it's a mathematical compression attempt at \~5k -> 2k. This saves a MASSIVE amount of time.

Currently the data lives in a \~KD format. 32 logits stored from the teacher of the Wikipedia corpus. Instead of training 1 hot, it gets 32 soft targets to try to match. Hard cross entropy is brutal on a model and it's the reason you see "trillions of tokens" quoted.

When you borrow an LM Head / Embed / Tokenizer / Teacher model... This is far less.

\---

Where are we now in training? I passed the 100m token mark yesterday at \~4am Pacific.

The current HF repo has all the checkpoints, data, and the 104m mark safetensor.

\--

I wanted to do some inspection of the model at this point. We have to see that it's not complete garbage right? It has seen a fraction of the data it needs.

So based on ONLY 100m tokens seen, I attempted completions and various ablations. Studies of what EXACTLY Engram stores (because it's mostly a guess).

Let's go straight to completions :

"George Washington was an American"

With Engram - "George Washington was an American naval officer who served in the American Revolutionary War. He was a naval officer who served in"

With Engram Zero'd - "George Washington was an American, but he was not a. He was a very good friend and a. He was"

"The American Civil War was a civil war in the United States from"

With - "The American Civil War was a civil war in the United States from 1861 to 1861. The war was a major victory for the Confederacy..."

Without - "The American Civil War was a civil war in the United States from 1861 to 1862. The war was a major victory for the Confederacy..."

"Aristotle was an Ancient Greek" (probably my favorite)

With - "Aristotle was an Ancient Greek philosopher who was a leading authority in the philosophy of Plato. He was also a leading authority..."

Without - "Aristotle was an Ancient Greek word meaning "to be" (ἀπάς, "to be")..."

\---

What do we learn from direct inspection of Engram?

If it saw the data enough times, it starts to offload it to the tables. At that point the model is less forced to use internal computational space to store data, and can rely on Engram.

That's not some interpretation. That's what the data shows exactly.

In the early 100m tokens of Wikipedia it's HIGHLY biased to early Philosophy. This has to do with the topics of what was IN Wikipedia at that time. Early Wiki contained a lot about Philosophy. It's among the earliest topics to exist.

It has seen TONS of Philosophy so it has basically offloaded to Engram because of the repetition.

Ok this is long enough. I'm calling it quits. I'm still looking to find a provider that will sponsor the model to completion. Right now I'm just paying. I contacted Verda and they never responded. Qubrid is in this sub, and they said they'd help but never emailed back after a few repeated proddings. Massed Compute is where this model is currently training, and I asked them like yesterday. Hopefully they'll help a brother out.

None of the above is "LLM written" except the completions from testing I guess.

As usual, ask whatever. It doesn't matter your understanding level or whatever. I'll sit and respond.

No question too dumb. No insult not insulting enough. (I'm joking. It's Reddit)

Much Love.

▲
161
-4
16👁
r/LocalLLaMA · u/Akainu_Fan · 18d ago
Did Alibaba abandon 35B A3B?

Basically the title.We did not get a new moe model with qwen 3.8 and Alibaba did not announce any small moe models on apsara.I know we might get an announcement later but ngl I kinda lost hope

▲
161
+4
24👁
r/LocalLLaMA · u/power97992 · 18d ago
Now Opus 5.5 is 58 on Artificial Analysis , how long do you have wait until an open model hits 58?

It is 12 points higher than the best open model mimo 2.6 pro and a big jump from fable 5.1. Crazy, glm 5.5 and qwen 4 will be on par with gpt 6 sol or better since it has a score of 48

If it took 2 months for the best open model to go from 44 to 46, then at this rate, in 6 months , they will reach 58? It is quite possible they will reach it in 4-5 months, since they have will more leaps in intelligence as they deploy more gpus and scale up the parameters, data and compute and improve the architecture .Wow sol 6 is worse than 5,6 at deepswe?

▲
126
+3
24👁
r/LocalLLaMA · u/niacolhealth · 18d ago
AntLing open sourced the Ming-Image-0.1-Design family

AntLing open sourced the Ming-Image-0.1-Design family:
• Ming-Image-0.1-Design, 6B
• Ming-Image-0.1-Design-Layer, 6B
• Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill

Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard

▲
119
+4
14👁
▲
97
-4
23👁
r/LocalLLaMA · u/futterneid · 17d ago
Streaming Nemotron 3 Diarization post image

I’ve been playing with Nemotron 3 Diarization, and it fills a gap I’ve had with local voice agents: keeping track of who is speaking.

It’s a diarization model, so it gives you speaker labels rather than transcriptions or people’s names. It can stream its output and track up to eight speakers. I’ve been trying it with one-second streaming chunks, and the quality has been really good in my tests.

I plugged it into my speech-to-speech setup on a DGX Spark and connected it to a Reachy Mini. The fun part is watching someone new speak, then seeing the robot ask their name and remember it for the conversation (that's the video!).

It has day-zero Transformers integration and is getting a commercial-friendly license.

▲
97
+2
26👁
r/LocalLLaMA · u/jacek2023 · 17d ago
apple/LensVLM-9B · Hugging Face

https://huggingface.co/bartowski/LensVLM-9B-GGUF

[](https://huggingface.co/apple/LensVLM-9B#lensvlm-9b)LensVLM-9B

LensVLM is a 9B Vision Language Model (VLM) that scans compressed images of text, then selectively expands only the relevant pages to their uncompressed form via learned tools.

[](https://huggingface.co/apple/LensVLM-9B#license)License

All ML model files in this repository, including Apple's modifications to the Qwen model, are provided under the terms of the Apple Machine Learning Research Model License.

The source code that accompanies this model is distributed separately and is provided under the terms of the Apple Sample Code License.

▲
92
+2
31👁
r/LocalLLaMA · u/zyxciss · 18d ago
Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW post image

Someone recommended that I try the ByteShape Qwen 3.8 27B IQ3-XXS GGUF after seeing my previous testing of the GSQ quant.

So I did.

And the result was… surprisingly bad.

For context, I'm running:

  • RTX 3060 12GB
  • 16GB DDR4 RAM, single channel
  • CachyOS / Arch Linux
  • llama.cpp
  • Qwen 3.8 27B
  • MTP/speculative decoding where applicable

The two low-bit quants I compared were:

ISTA-DASLab / GSQ-RCO-IQ3-XXS

  • \~10.4GB
  • roughly 2.5 BPW territory
  • MTP enabled
  • \~29 tok/s around full context
  • \~34–40 tok/s at lower context
  • This was the quant I had already been using in my previous web-development test.

ByteShape IQ3-XXS

  • roughly 500MB smaller
  • also around the same ultra-low-bit range
  • advertised as having extremely high similarity to the BF16 model based on KL-divergence measurements

On paper, the ByteShape quant looked very interesting.

It was smaller, while apparently retaining extremely high similarity to the original BF16 model. It was also being compared in size to significantly higher-BPW quants.

So naturally I expected it to at least be competitive with the GSQ version.

It wasn't.

The actual result is shown Above

ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF was able to generate a 3D voxel diorama in one shot under 55k tokens, and Byteshape's 3.8 27B model, took roughly three shots and still hasn't completed with over 98K tokens spent already.

Same with web development not impressive as advertised in Here

Any New Model Suggestions for RTX 3060?

▲
89
 
30👁
r/LocalLLaMA · u/FactoryReboot · 17d ago
Most powerful harness for Qwen 3.8?

I hear Qwen code unlocks the model better. I also think it has more power user features than open code?

It’s nice open code can work with multiple models easier though

Thoughts?

▲
74
-3
20👁
r/LocalLLaMA · u/HolidayBit143 · 17d ago
Unsloth Studio VS LM Studio... Which one do you prefer?

So I've been experimenting with various platforms and even on day 1 Unsloth Studio was released, I knew that LM Studio had it's days numbered. LM studio will always be the OG but I wonder how much longer they have, especially with all these new platforms arising. It feels like LM studio just fell behind and it doesn't help that they are putting so much effort on BIONIC which I'm not even sure if anyone uses on a serious level.

Which one do you prefer, or do you use something else?

▲
81
+5
29👁
r/LocalLLaMA · u/returnity · 18d ago
mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)

Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:

  • Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
  • Self-supervised learning: trains itself on new material constantly
  • Weights stored on SSD and paged in on-demand
  • Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
  • Dynamic size: builds new experts and increases parameter counds as-needed
  • Evolutionary growth: trials newly generated experts, unused ones are pruned back
  • No tokenizer: it reads raw bytes directly
  • Catastrophic forgetting prevented by slow trunk/fast experts learning rate split

Weights will be released in "a couple weeks" once training progress reaches \~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.

What do you guys think?

▲
72
-3
28👁
r/LocalLLaMA · u/BagComprehensive79 · 18d ago
About Mimo 2.6 Architecture

I was checking nee Mimo 2.6 architecture on huggingface page and it looks very simple. I dont mean in a bad way but when we compare recent open models, their architecture is very simple. They dont use any Gated DeltaNet, no mHC or similar architecture, no engram. Just ordinary simple architecture and very good RL i guess.

What are you guys thinking about this?

▲
64
-1
29👁
r/LocalLLaMA · u/AvidCyclist250 · 17d ago
Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080

Thought it's about time to share after testing for a week. You need four things most people miss: the right quant, the right model, the right branch, and the right cache flags.

https://github.com/dtm-beep/qwen38-flash-next-mtp-16gb

TLDR: AtomicChat AD-4.27bpw Q4_K_M target + the shared Unsloth MTP head, build from my pr-mtp-fix branch (plain master can't load this MTP head yet, it's PR #28243 + one fix commit), and --spec-draft-cpu-moe is the trick that makes 16 GB work. Draft experts live in RAM so the target's hot experts get the GPU. IQ4_XS ~10 t/s → 16.5 tg / 350 pp at 131k, q8 KV.

Hope it helps someone.

▲
63
+1
34👁
r/LocalLLaMA · u/OvertaxedOne · 17d ago
Qwen FN vs 27B --- Think I'm saturated.

Got QFN up and running on our Strix box this past weekend and have been running on it for a few days now. Big thank you/shout out to the Halogen team, it's running fantastic on the Strix, this is clearly "the setup" right now for this hardware with this model, really impressive performance for such a large model (\~30-40TPS generation, \~900-1000TPS prefill on real world use, not benchmarks, over the past few days)!

However, that said, I kind of feel like 27B saturated my personal use cases. QFN is a great model, but I'm not really noticing much where I think "Wow, 27B would never get this and QFN just one shot it". They feel very similar in capabilities (IE, both are amazing!) and I kind of feel like I'm reaching the end of the runway for what I can realistically make use of in my day to day use cases, I'm just asking questions that really require more than 27B the majority of the time.

It really feels more to me like a very similar level of smart, one that runs well with limited memory bandwidth (QFN) and one that runs well with limited GPU memory (27B). It's an interesting result, and I guess maybe I should have expected it because I was really struggling already to find things that I actually need to do day to day professionally that 27B couldn't do. When I escalate to the cloud now it's nearly always for either speed or context, rarely intelligence.

Surprising result, at least to me, I was expecting to have my hair blown back, but I guess this is just yet another data point that "good enough is good enough".

▲
65
+4
28👁
r/LocalLLaMA · u/Dutchnamn · 17d ago
Perhaps the highest quality mainline quants of Qwen3.8 27B?

I am proud to release these quants of Qwen 3.8 27B. They beat the excellent ISTA and Unsloth quants byte-for-byte on three corpora. Both KLD and top 1% were tested 3x. It took a week of continuous GPU and CPU time to generate these, all done on a single Strix Halo.

https://huggingface.co/agentionai/Qwen3.8-27B-AP-GGUF

Hope you like it.

Edit: I did a lot of benchmarking and updated the smallest quant. Slightly improved calibration led to this real life result

https://preview.redd.it/eax6krk4x7sh1.png?width=1600&format=png&auto=…

▲
61
+3
25👁
▲
50
 
6👁
▲
872
+11
41👁
r/LocalLLaMA · u/tiensss · 17d ago
Jev isn't new tech. Its marketing targets people who think AI started with LLMs.

I keep seeing Jev presented as some new class of decision model, but most of what’s being advertised is just normal classifier behavior with modern zero-shot capabilities.

It outputs probabilities over constrained choices, doesn’t generate autoregressively, can’t output an invalid class, and can use labels defined at inference time. None of that is new. Zero-shot/NLI classifiers, embedding models, cross-encoders and rerankers have been doing variations of this for years.

The weird part is that most of the impressive Jev comparisons are against LLMs. Of course a specialized classifier is faster and cheaper than making an autoregressive LLM generate an answer. That doesn’t establish a new paradigm. The meaningful comparison is against strong existing classifiers. The purpose of this is to mislead.

There are already benchmarks like BTZSC evaluating dozens of zero-shot classifiers across 22 datasets, including NLI models, embedding models and rerankers. I haven’t seen Jev properly benchmarked across that landscape yet.
(https://proceedings.iclr.cc/paper\_files/paper/2026/hash/417e1c15b3d49852fceded8aa104107d-Abstract-Conference.html)

Where people have compared Jev with conventional classifiers, the story is much less magical. One Banking77 experiment got 93.3% from BGE-small + logistic regression versus 83.2% for Jev, at about 9ms locally.
(https://github.com/ickma2311/jev-baselines-eval)

Some of the marketing also goes into the misleading territory. The “can’t hallucinate” framing is very sus, for example. Their own explanation admits the 0% hallucination figure is not empirical, and what they actually guarantee is that Jev returns an answer matching the allowed schema. That prevents invalid outputs, it does not prevent confidently choosing the wrong valid answer. (https://typesafe.ai/blog/introducing-system-one-models-and-jev)

So color me a skeptic. Look, Jev might even be a good product. Maybe their unpublished architecture or RLCD training method is genuinely novel. But nothing we've seen so far establishes that "System One Models" are a new class of AI. What the public evidence mostly establishes is that using a specialized classifier for classification can be much cheaper and faster than using an autoregressive LLM, which we already knew. It only sounds novel if your idea of AI begins and ends with LLMs.

💬 333 (-1) open on reddit ↗
▲
316
 
32👁
r/LocalLLaMA · u/Khaledthe · 17d ago
Using uncensored models makes working less of a headache

I have a lot of projects with my friends and team at work that I copy to use for my personal projects, whether it's a plugin I borrow with their consent or a script. I always find that Qwen 3.8 and Muse Spark 1.3 straight up refuse to do anything, as they see it as a steal, so I have just been rocking Qwen3.8-27B-Heretic-JP-Roleplay-NSFW-DanbooruTags.i1-Q4\_K\_M, and it feels so good to just be able to tell it to do something, and it actually does it. I know the model isn't made for coding or projects but roleplay, but I don't see a big dip in performance as it does what's asked to do.

Does anyone else have this problem or not?

💬 102 (-1) open on reddit ↗
▲
1296
+4
38👁
r/LocalLLaMA · u/Acrobatic_Stress1388 · 17d ago
Mods: can we do something about half the forum getting filled with these advertising posts for Jev?

Jev is a paid product that dumped a lot of venture capitol money into shill their product here and in other subreddits. Obvious shill posts are obvious.

💬 241 (-3) open on reddit ↗
▲
984
 
35👁
r/LocalLLaMA · u/Atagor · 17d ago
Pirate Face - pirate bay for LLMs

The title says for itself

In case someone desides to censor huggingface, we'll have an alternative

Edit:

A lot of responses so I'll leave it here:

  1. I'm not the author.
  2. If I were the author I wouldn't use the word "piracy".
  3. If you're the author, please, rename the domain! What is free in the first place must be named as such, we're not pirating anything.
💬 123 (-4) open on reddit ↗