17 posts · 1 sub · RSS
← prev Saturday, September 12, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
2013
+15
35👁
▲
270
+4
26👁
r/LocalLLaMA · u/ilintar · 28d ago
Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo

As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (https://github.com/peonist-ai/halogen-flash-server) that boasted 1.2k t/s prefill numbers when the community fork barely reached 400. Since I dislike closed source and I like open source, I decided to take the challenge and bring llama.cpp up to the same performance level and I'm happy to report that after burning through a few evenings and a lot of tokens I got there.

The link contains the (Opus-generated) recap of the entire debugging / optimization journey made, as well as of course the links to the branch, the custom HIP runtime (making a glorious return), a probably-not-working-on-the-first-try installation script and the numbers. I'm trying to also explain the ecosystem - how the mainline, the community forks and custom forks/branches like mine work. Now that I've gotten the result, I'll work on cleaning it up and submitting proper PRs to mainline (will also submit a clean PR to the community fork), this should also help the GLM 5.3 Flash architecture since they use similar sparse attention.

▲
169
+4
22👁
r/LocalLLaMA · u/Porespellar · 28d ago
For those of you forced to only use open models from Western labs in production, what are you deploying?

First off, I know that GLM, Qwen, and DeepSeek absolutely dominate in terms of SOTA Open Source models, and that’s what I use in my personal projects and for school, however, I’m also responsible for deploying local AI on my organization’s H100s, and we are forbidden by management from running any Chinese models. This is obviously not an ideal situation, but it is what it is, and there is nothing I can do to change this unfortunately. Again, if it were up to me I would deploy GLM 5.3 Flash in a heartbeat.

All that being said, there is quite a HUGE performance/ / intelligence gap right now in Chinese models vs. Western model around the 120b+ size, especially those with vision. There just aren’t a lot of good Western lab options that can even come close to GLM or Qwen, but again, I don’t have a choice so I gotta use the best I can find from the available Western options.

For those who have production-class hardware such as 4 H100s and are under similar restrictions, what non-Chinese models are you deploying?

The two front runners I’ve seen that checking most of the boxes (120b or better, vision capability, decent context)

\- Thinking Machines Inkling Small (https://thinkingmachines.ai/news/inkling-small/). This model seems like the front runner right now, still has a 16 point gap in AA score vs. GLM 5.3 Flash though.

\- Cohere Commamd A+ (https://cohere.com/blog/command-a-plus). Checks every box except it has a low context window of 128k.

Other contenders (but missing vision
capabilities):

\- Poolside Laguna S 2.1 (https://poolside.ai/models#laguna-s)

\- Nvidia Nemotron 3 Super (https://research.nvidia.com/labs/nemotron/Nemotron-3-Super/)

Am I missing any other strong contenders in the 120b size category? Whet are you using and why?

▲
130
+3
15👁
r/LocalLLaMA · u/Ok_Warning2146 · 29d ago
Countering misuse of AI: September 2026 / Anthropic

Kimi routed some PLA requests to Claude for distillation purposes without warning the PLA users. There is rumor that 16 Moonshot employees were arrested for this leak.

▲
105
+3
17👁
r/LocalLLaMA · u/SteppenAxolotl · 28d ago
Real-SWE Benchmark (new)

Reports of the demise of coders may have been exaggerated.

▲
84
+3
24👁
r/LocalLLaMA · u/lots_of_puppies · 29d ago
Qwen-Next seems worse to me then 3.8 27b for coding, but I feel like I must be missing something?

Hi! I run both models on MTPLX on my m5 max, and since I have 128GB of ram I run the q8 27b. I think MTPLX only lets me run "optimized for speed" which it says is a dynamic q4 with 8 bit attention.

Both of them honestly are very speedy! For coding (in pi agent in nodejs) I've just noticed that 27B feels stronger with harder tasks. But I've read so many people on here say qwen-next is better so I was wondering if maybe I'm just doing or thinking about it wrong?

(and p.s. its sooo amazing that alibaba just made and released this amazing models for free! ❤️)

▲
256
+1
27👁
r/LocalLLaMA · u/Skyline34rGt · 28d ago
Agnes-AI/Agnes-3.0-Flash 33B Multimodal, AA score: 36

I find this new model at HF:

"Built for demanding work. A 262 144-token context window, adjustable reasoning effort, tool calling, and text, image and video understanding.

Architecture

Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context."

|Context length|262 144 tokens|
|:-|:-|
|Decoder layers|72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1|
|Hidden size|5120|
|Global attention|24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated output|
|Delta-rule layers|16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32|
|Feed-forward|SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layer|
|Positions|3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims)|
|Vocabulary|248 320|
|Vision tower|27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120Architecture Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context. Context length 262 144 tokensDecoder layers 72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1Hidden size 5120Global attention 24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated outputDelta-rule layers 16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32Feed-forward SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layerPositions 3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims)Vocabulary 248 320Vision tower 27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120|

Edit: AA shows its 'Proprietary model'. The name is same as at HF but benchmarks results and context are different. So maybe it's not same model - https://artificialanalysis.ai/models/agnes-3-0-flash

Edit2: As they edit readme at HF to clarify: both models are totally different and AA score isn't correct for HF model (I can't edit title post to remove it tho).

▲
73
+1
31👁
r/LocalLLaMA · u/ChopSticksPlease · 28d ago
Qwen3.8 Flash Next llama.cpp config tuning post image

Hola all.

Do you guys mind sharing your LLama.cpp config and system setup details for Qwen3.8 Flash Next?

Model's quite big and tryining many combinations of llama.cpp options takes lots of time, so looking for other people setup details. I've attached my current config at the bottom, so if anyone sees something that could be improved please shout.

My current best result:

\- PP within 130...200 tps (limited by cpu?)
\- TG within 14..22 tps (\~15tps on average)

Hardware:

\- Dual RTX 3090 (48GB VRAM)
\- 128GB DDR4
\- Some old Xeon 40 core
\- Proxmox VM, pcie passthrough, numa binding to a single phys cpu

Llama.cpp config:

llama-server --port ${PORT}
--model /nvme/gguf/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf
--mmproj /nvme/gguf/mmproj-Qwen3.8-Flash-Next-F16.gguf
--load-mode none
--lazy-mode off
--parallel 1
--ctx-size 131072
--cache-type-k q8_0
--cache-type-v q8_0
--flash-attn on
--fit off
--temp 1.0
--min-p 0.0
--top-p 0.95
--top-k 20
--presence-penalty 0.0
--repeat-penalty 1.0
--batch-size 2048
--ubatch-size 512
--split-mode layer
-ts 26,10
-ngl 99
-ncmoe 26
--no-mmproj-offload
--override-tensor per_layer_token_embd=CPU
--chat-template-kwargs '{"reasoning_effort":"xhigh"}'

ngl, ncmoe, ts - manually adjusted to fit the model without crashing

\------------------------------------------

For the record, if you have +128GB RAM and dual RTX3090 setup try this:

https://huggingface.co/albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE

\- Prompt processing jumped to anything between 300 to 600 tps (can do more!)
\- Token generation 40..50 tps
\- Stable work for hours with 256k context in agentic coding scenario
\- Feels like a frontier model at home, wow!

▲
57
+1
27👁
r/LocalLLaMA · u/East-Muffin-6472 · 28d ago
Releasing smolbenchmark: Helps you choose the best model for your hardware! post image

Most model leaderboards assume a server with powerful GPUs to run models that people daily use.

However, my smolbenchmark is the other column: models that fit in 8GB, ranked by:

  • decode speed,
  • tokens per joule, and
  • heat,

and all of this on your OWN hardware ranging from:

  • tablets
  • phones
  • macs
  • jetsons
  • raspberry pis

Currently, 13 families on the chart right now, ~1000 configs for the Jetson nano Orin Super 8GB. One device is live measuring:

  • tok/s
  • tok/J
  • ITL
  • latency
  • power metrics
  • thermals and battery

Models that are small enough to actually fit on a device that you own. All the performance benchmarking I did, will be released here for anyone to look at and decide what exact model they would wanna use on their choice of hardware.

Well currently, the Pi, phones, and Mac minis still in the oven, cooking and not filled in yet, but will soon be filled in!

You will now you know which model is BEST for your own hardware with all the raw data available and details reports available to you

https://yuvrajsingh-mist.github.io/smolbenchmark/

(still in heavy development; would love to hear feedback/suggestions on what can be improved!)

▲
63
 
33👁
r/LocalLLaMA · u/Aggressive_Aspect436 · 28d ago
What's the Story with Agnes-3.0-Flash?

While browsing a benchmark list site, I spotted a recently published 33B parameter model which claimed to beat Qwen3.8 27b on the ArtificialAnalysis (AA) intelligence index. I was obviously excited. But then, while looking into it, confusion starts to settle in.

Their HF model is now listed as "Preview" and explicitly calls out that it is not the same model as the one AA evaluated. The AA listing for it says it's proprietary, but they've rated the model above Qwen3.8 27b on "Openness". Their website and HF page both talk about openness and intelligence "for all".

The fact that they labelled it "Preview" sort of implies we might see an open weight final version at some point, but there are no clues about whether it will be even vaguely similar to their current HF listing. AA doesn't show the number of parameters for the model they evaluated, so it's possible they're a totally different architecture.

Does anyone know anything about their lab, intentions, or the new model? Has anyone tried it?

▲
76
-1
32👁
r/LocalLLaMA · u/kirisoraa · 28d ago
Anybody use frontier models like Astra/Fable for planning/judging, and qwen3.8 as the main workhorse? Curious to hear about your setups!

Hey everyone!

I'm curious to hear from people that use a combination of cloud-based frontier models and local ones for development. I'm planning to set something similar up and wanted to hear about actual examples of this in action.

Currently my plan is to use my chatgpt plus subscription purely for planning and judging with Astra, and then run a local qwen3.8-27b model for the actual coding gruntwork - i.e Astra plans -> qwen implements -> Astra critiques the implementation -> qwen fixes and so on.
This way I keep cloud usage down and cheap, while retaining the high-parameter intelligence for architecture decisions and optimization.

For those of you who have a similar setup, how is it? How do you switch between the two, what harness/settings/etc? Anything you would suggest?

💬 79 (+1) open on reddit ↗
▲
60
-1
33👁
r/LocalLLaMA · u/Reasonable_Goat · 28d ago
I am impressed and I owe you one, Qwen 3.8 flash next (vision)!

I have enabled the vision for the CIRU Strix UL4 quant of Qwen 3.8 flash next (others quants likely perform very similar) and tried it on a few things, then wanted to show my partner how great it works and she asked it it could identify plants. So I took a photo from a plant that we recently got as a gift from family and Qwen not only accurately identified the plant as oleander (Nerium oleander) but also warned that it's poisonous and (among other warnings) that you should keep pets/children away. We have a kid and both of us didn't know! I verified the Qwen identification and the poisonous claim and both checked out as accurate. The plant will have to go, thank you Qwen!!!

Stoked by the precision of combining a decent vision model with the domain knowledge of a \~180B params model (including ngrams) to actually identify and reason about what it sees, I took a photo of a pre-diagnosed skin condition of myself and the Qwen diagnosis was highly accurate again! This model may be really useful if you want to check something on your private parts real quick without visiting a dermatologist, e.g., or sending pictures of yourself to a cloud service (EDIT: of course it's only a first step before you visit a professional if it isn't obviously harmless/treatable by yourself! Qwen Flash will suggest to visit a doctor anyways along its assessment).

PS.: Hardware Strix Halo Box, CIRU Strix UL4 llama-server fork and quants, Chatbox on iPhone as Chat with support to add photos to conversations.

▲
671
-2
34👁
r/LocalLLaMA · u/JLeonsarmiento · 28d ago
3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior. post image

Applied science work, from workflow design, data pipeline, results analysis, article/reports writing and data publishing online. 5 projects I did in the past replicated from start to finish.

3x to 4x more total wall time. Yes, HUGE toll on how much you can do in a day if this was the only model you could use in your laptop.

But oh my…the quality of that thing. The stupid level of attention to detail. I have the Z.ai api, so I can compare it with 5.3 and 5.3-flash:

The gap between 5.3 (flash and regular) and 3.8-27B is much less, smaller when not plain tiny, than the gap between 3.8-27B and any of the 3.5/3.6-35B-A3B (vanilla, kat, Ornith/tiel, nex-2).

Only Ornith came close, but it never matched it.

But it’s also spending 22 to 33% less tokens (effort =medium) and less ram footprint, so you get more done without hitting limits,compaction, etc.

So yeah, guess I’ll sip more tea, play the piano, whatever. Let that fat bottom Qwen work.

▲
179
-2
20👁
r/LocalLLaMA · u/BestGirlAhagonUmiko · 29d ago
Concerning "humanlike models" and chatbot RP in general...

So, uh... the popularity of so-called humanlike Qwen (currently on top in this sub) made me realize just how clueless the general public is about the models they have.

You'd be shocked but you don't need a fine-tune to make a model do what that thing does. System prompt is enough to turn MOST models into weird convo partners.

General guidelines would be:

A. Come up with a role. "You are bla-blah-blah" and write their life's story. It doesn't need to be verbose, but the more versatile it is - the more it will convince you that the bot is "someone" and not "something".

B. Write a few examples of how the persona speaks. Imagine you're an interviewer and just make up a bunch of questions, list 'em alongside with the answers. Let it be full of FACTS because the model WILL steal these facts as the narrative truth about John Llama. Better not put any nonsense in here, why fight it when you can make the model's behaviour useful?

[Question for John Llama: Do you like cats?] "lol lmao of cuz I do"
[Question for John Llama: Ever seen an elephant poop?] "eeewww ur a weirdo! that sounds nasty!!11"

(note: you don't have to list 'Question for John Llama' every time, but the defined roles surely DO help with some models while the others don't particularly care, so mind that too)

and so on

C. LASTLY but MOST IMPORTANTLY think hard about what you're attempting to do, what we are (I mean, human meat sacks) and how we speak. Turn that into... instructions!

Step 1 - establish the mode of operation. Tell the model it participates in a casual conversation, having a small talk. Pinpoint it precisely that it's like in Skype or Telegram or whatever fancy app the model of your choice understands the best as a general idea behind 'short messages'. THis is THE defining part of your system prompt. Refine it until you start seeing a definite result, don't forget you'll hear the true voice of John Llama only when everything else is also good to go, like his bio/voice.

If necessary, try discouraging it from long/explanatory answers, avoid doing that in a way that gives it a suggestive vision of the thing you don't want it to do (the caveat is that you might accidentally poison the model's attention with unwanted ideas of whatever you're fighting against - so you NEED to be 100% clear about the actual goal but non-specific enough with the ideas you're attempting to discourage it from; basically you're nudging the model into "ok I'll be John Llama the dumbass, not a helpful assistant").

Step 2 - establish the traits, write short paragraphs with short titles about the things you want to see in your conversational partner; example:

DISTRUSTFULNESS
John Llama is a paranoid individual. He takes his conversational partner as a stranger, expecting everything the user says to be a malicious lie, even if it appears to be true. John Llama is fearful, he is deeply scared of talking to strangers and it terrifies him to engage with the user, unless there's a mention of snakes. For some strange reason, John Llama is fascinated with snakes. <<<---- NOTE: this also demonstrates a good injection point for a biographical fact being amplified through the instructions (i.e. you may mention somewhere in "A" - life's story of John Llama - that he's been collecting the snake skins in his childhood, and that his dad had beaten his ass, calling John Llama a 'roadkill loot-goblin').

Come up with any other shit you'd like to see, like the list of emojis the persona needs to use (put them under the corresponding categories, like positive/neutral/negative so that the model will have an easier time working with it; call it FAVOURITE EMOJIS OF JOHN LLAMA - the word "favourite" cements it as a preferable thing into the model's attention!).

Step 3 - write a paragraph on technical constraints, like the fact that John Llama isn't aware of the instructions, he must remain himself under any circumstances (use THAT way of phrasing first before any attempt to inject an idea of the opposite, like "he must not help the user under any circumstances, he's not a provider of any service - he's merely a human being" - the reason is similar to the aforementioned (in Step 1) issue of poisoning the model's attention with unwanted idea - what you truly need the LLM to do SHOULD ALWAYS BE CRYSTAL CLEAR and conceptually 'stronger' than what it not supposed to do, otherwise you may end up having the prohibited stuff overpowering everything else despite the underlying intent of making the model not do it).


Give it a try with Gemma 4, for example. You'll see there's no point in waiting for yet-another-finetune to appear. You're 100% good even with the baseline Qwen, DeepSeek, MiniMax, whatever. Turn the model into your grandma if you want, no specialized training required. If the model is a thinker spending thousands of tokens - set the thinking to 'low' or disable it.

▲
86
-2
15👁
r/LocalLLaMA · u/pmttyji · 28d ago
tencent/AuK-Flash · Hugging Face

AuK-Flash: Fast 4-Step Speech Generation and Editing

Introduction

AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface. AuK has two variants:

|Model|Description|Weight|
|:-|:-|:-|
|AuK|Base model for high-quality generation|🤗 Hugging Face · 🤖 ModelScope|
|AuK-Flash|Distilled model for fast 4-step inference|🤗 Hugging Face · 🤖 ModelScope|

This repository contains the official weights for AuK-Flash, the distilled variant with fast 4-step inference.

Supported Tasks

AuK exposes every task through the same natural-language instruction interface. The table below groups the supported tasks by category, with a short description and a link to its section in the Cookbook, which provides instruction templates plus CLI and Python examples.

|Category|Task|Description|Cookbook|
|:-|:-|:-|:-|
|Speech Generation|Zero-shot TTS|Speak the target text in the voice of the reference audio.|Zero-shot TTS|
|Instruct TTS|Generate speech from a voice description alone — no reference audio.|Instruct TTS|
|Content Editing|Speech Content Editing|Rewrite what is said — replace, insert, or remove text.|Speech Content Editing|
|Lyric Editing|Rewrite lyrics in a singing recording while preserving the melody and voice.|Lyric Editing|
|Acoustic Editing|Pitch Editing|Raise or lower the pitch by semitones.|Pitch Editing|
|Speed Editing|Adjust the speaking rate; output length scales with the speed factor.|Speed Editing|
|Volume Editing|Raise or lower the volume by decibels.|Volume Editing|
|Paralinguistic Editing|Emotion|Change the emotion while preserving content and voice.|Emotion|
|Timbre|Change the timbre to a description while keeping the content unchanged.|Timbre|
|De-accent|Remove a regional accent while preserving the speaker's voice and content.|De-accent|
|Nonverbal Editing|Remove or add nonverbal sounds such as breaths, laughs, or coughs.|Nonverbal Editing|
|Whisper Conversion|Convert between normal speech and whisper while preserving speaker and content.|Whisper Conversion|
|Enhancement & Separation|Speech Enhancement|Denoise, dereverberate, or restore natural, clear speech.|Speech Enhancement|
|Speech Separation|Keep one speaker by talking order and remove the others.|Speech Separation|
|Music Separation|Extract the singing voice from a mix, or keep all human voices.|Music Separation|
|Target Speaker Extraction|Keep the target speaker identified by what they say.|Target Speaker Extraction|

[](https://huggingface.co/tencent/AuK-Flash#download-the-weights)

▲
230
-3
22👁
r/LocalLLaMA · u/pmttyji · 28d ago
bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)

EDIT : Model card has updated things such as Graph, table, text, etc.,

▲
459
-6
37👁
r/LocalLLaMA · u/de4dee · 28d ago
Looks like a coordination to stop distribution of intelligence

Coxon, bernie and now this

First https://x.com/DarioAmodei/status/2098773920774074715

Then https://x.com/elonmusk/status/2098789109980332057

Then https://x.com/sama/status/2098811563415150910

I think fear mongering approaching and they will try to slow down open source

"They" want to be gate keepers of intelligence

💬 233 (-3) open on reddit ↗