15 posts · 1 sub · RSS
← prev Tuesday, September 22, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
2110
+16
52👁
r/LocalLLaMA · u/Salah_H_Hasan · 18d ago
Qwen 4 Announced at Apsara Conference

https://preview.redd.it/bpbc9i6hizqh1.png?width=1270&format=png&auto=…

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

💬 568 (+3) open on reddit ↗
▲
512
+5
21👁
▲
81
+5
29👁
r/LocalLLaMA · u/returnity · 18d ago
mini-AGI: Continual-learning dynamically looped transformer with evolutionary grown (on a laptop)

Saw this today and found it very intriguing. Lots of interesting design choices here, and it's cool to see someone doing something different. Here's a few highlights:

  • Looped transformer: dynamic recurrent depth on a per-token basis, up to 24 cycles
  • Self-supervised learning: trains itself on new material constantly
  • Weights stored on SSD and paged in on-demand
  • Mixture of Experts: 8 active, 32 routed held in VRAM, smart caching of 96 more
  • Dynamic size: builds new experts and increases parameter counds as-needed
  • Evolutionary growth: trials newly generated experts, unused ones are pruned back
  • No tokenizer: it reads raw bytes directly
  • Catastrophic forgetting prevented by slow trunk/fast experts learning rate split

Weights will be released in "a couple weeks" once training progress reaches \~GPT-2 levels. The trend line has held 15-fold so far, but it may bend at some point, so that is definitely a rough estimate of the trajectory.

What do you guys think?

▲
249
+4
17👁
r/LocalLLaMA · u/108er · 17d ago
Qwen image 2.1 (Fast FP8) generates premium quality images post image

Don't know how they did it, but for under 10GB model, the results are astonishing. I am running it on Unsloth Studio. They just released the update, so if you are not seeing the option, I recommend updating your Unsloth Studio. Cheers!

▲
161
+4
24👁
r/LocalLLaMA · u/power97992 · 17d ago
Now Opus 5.5 is 58 on Artificial Analysis , how long do you have wait until an open model hits 58?

It is 12 points higher than the best open model mimo 2.6 pro and a big jump from fable 5.1. Crazy, glm 5.5 and qwen 4 will be on par with gpt 6 sol or better since it has a score of 48

If it took 2 months for the best open model to go from 44 to 46, then at this rate, in 6 months , they will reach 58? It is quite possible they will reach it in 4-5 months, since they have will more leaps in intelligence as they deploy more gpus and scale up the parameters, data and compute and improve the architecture .Wow sol 6 is worse than 5,6 at deepswe?

▲
119
+4
14👁
▲
343
+3
28👁
r/LocalLLaMA · u/ironicstatistic · 18d ago
Ngram and world knowledge - why are we just building a coding model?

This post is written by a human and I'd appreciate it if you treated it as such. Thanks.

So, I've been noticing a pretty clear interest in developing as good a coding and agentic tool-calling model as possible, especially at smaller sizes, sub-50 gigs. However, I'm finding that at least for my use of AI, if I really want to move away from big providers, I am going to require a model that has better world knowledge than the current offerings.

Qwen 3.8 27B is a truly fantastic model for tons and tons of stuff. It's highly intelligent, super good at designing applications and coding and working on my system. However, its world knowledge sucks​ compared to the frontier, especially at the Q4 quant that I have to run it at.

So, that leaves me with a question. With the new N-gram technology that we're seeing being baked into Qwen 3.8 Next and that presumably will run on future models, why can't a model be made that has a smaller set of intellectual capabilities but a greater amount of world knowledge? I understand that right now everyone is optimizing towards making as smart a model as possible fit into as small a space as possible. But why don't we leverage the SSD to give the model a lot of world knowledge and make models that are better at dealing with screenshots, multilingual capabilities, doing things like pixel art or answering physics questions?

ngram seems like the answer to the "can't fit in vram" question... Qwen 3.8 Next really opens my mind to the possibility that there could be a totally different and better paradigm for how these models are developed, at least for many use cases. Having a relatively smart model with a large amount of world knowledge might be better than having as smart a model as possible...

Not to mention that this would mean that a model's training cut off would become less relevant, because it could just be fashioned a new ngram.

Obviously, the main interest is in creating a model that can code as well as possible because that's what'll capture market share. But I am curious if there are any efforts into this kind of thing or if anybody has an idea on why these things aren't done more often.

Please tell me why I'm wrong, how I'm wrong, and in how many ways I'm wrong because I'm sure that that's all you really want to tell me, but at least I'll learn something, because as is obvious from this post I have no idea what I'm talking about.

Thanks have a good day :)

edit: I found this post and I guess it provides a lot of what I was asking:

https://www.reddit.com/r/LocalLLaMA/comments/1vzgtqf/ngram\_vs\_experts\_explained/

edit2:

this one is even better, recconend reading. ty reddit suggestions:

https://www.reddit.com/r/LocalLLaMA/comments/1w0198r/no\_engrams\_wont\_let\_you\_run\_1t\_models\_locally\_it/

▲
126
+3
24👁
r/LocalLLaMA · u/niacolhealth · 18d ago
AntLing open sourced the Ming-Image-0.1-Design family

AntLing open sourced the Ming-Image-0.1-Design family:
• Ming-Image-0.1-Design, 6B
• Ming-Image-0.1-Design-Layer, 6B
• Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill

Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard

▲
120
+3
26👁
r/LocalLLaMA · u/Melted_gun · 18d ago
What underrated AI tools have actually made you more productive in 2026?

I asked this back in 2025, but the AI landscape has changed a lot since then.

Not looking for the usual ChatGPT, Claude, Gemini, Midjourney, etc. I'm curious about the lesser-known tools that you actually kept using.

Could be for research, coding, design, video, writing, automation, planning, journaling, local AI, or even something oddly specific.

Free or paid doesn't matter.

What tool genuinely saved you time or improved your workflow this year? And what do you actually use it for?

💬 106 (+1) open on reddit ↗
▲
92
+2
31👁
r/LocalLLaMA · u/zyxciss · 18d ago
Qwen 3.8 27B at ~3 BPW on an RTX 3060: GSQ vs ByteShape IQ3-XXS 2.88BPW post image

Someone recommended that I try the ByteShape Qwen 3.8 27B IQ3-XXS GGUF after seeing my previous testing of the GSQ quant.

So I did.

And the result was… surprisingly bad.

For context, I'm running:

  • RTX 3060 12GB
  • 16GB DDR4 RAM, single channel
  • CachyOS / Arch Linux
  • llama.cpp
  • Qwen 3.8 27B
  • MTP/speculative decoding where applicable

The two low-bit quants I compared were:

ISTA-DASLab / GSQ-RCO-IQ3-XXS

  • \~10.4GB
  • roughly 2.5 BPW territory
  • MTP enabled
  • \~29 tok/s around full context
  • \~34–40 tok/s at lower context
  • This was the quant I had already been using in my previous web-development test.

ByteShape IQ3-XXS

  • roughly 500MB smaller
  • also around the same ultra-low-bit range
  • advertised as having extremely high similarity to the BF16 model based on KL-divergence measurements

On paper, the ByteShape quant looked very interesting.

It was smaller, while apparently retaining extremely high similarity to the original BF16 model. It was also being compared in size to significantly higher-BPW quants.

So naturally I expected it to at least be competitive with the GSQ version.

It wasn't.

The actual result is shown Above

ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF was able to generate a 3D voxel diorama in one shot under 55k tokens, and Byteshape's 3.8 27B model, took roughly three shots and still hasn't completed with over 98K tokens spent already.

Same with web development not impressive as advertised in Here

Any New Model Suggestions for RTX 3060?

▲
133
-1
25👁
r/LocalLLaMA · u/No_Algae1753 · 18d ago
Am I going insane for thinking that these are all Bot comments?

I know that there are some bots active on this sub but wow these comments really look AI generated. Is it just me or are those really bot comments?

https://preview.redd.it/20st4vivq3rh1.png?width=955&format=png&auto=w…

source:

https://www.reddit.com/r/LocalLLaMA/comments/1wn9a2w/what\_underrated\_ai\_tools\_have\_actually\_made\_you/

💬 148 (+1) open on reddit ↗
▲
431
-2
18👁
r/LocalLLaMA · u/Sitkin_Marrel · 17d ago
New 6B image model coming, AntLing just open sourced the Ming-Image-0.1-Design family post image

• Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard. https://huggingface.co/inclusionAI/Ming-Image-0.1-Design https://huggingface.co/inclusionAI/Ming-Image-0.1-Design-Layer

▲
194
-3
22👁
r/LocalLLaMA · u/simpleuserhere · 18d ago
Laya model playing Flappy Bird on a CPU using OpenVINO INT8 inference post image

A 421M-parameter model just played Flappy Bird on my desktop CPU (OpenVINO int8)

Running on my Intel Core i7 12th gen CPU

Converted laya system one model to OpenVINO and quantized to int8

Repo : https://github.com/rupeshs/flappy-laya-openvino-cpu

▲
72
-3
28👁
r/LocalLLaMA · u/BagComprehensive79 · 18d ago
About Mimo 2.6 Architecture

I was checking nee Mimo 2.6 architecture on huggingface page and it looks very simple. I dont mean in a bad way but when we compare recent open models, their architecture is very simple. They dont use any Gated DeltaNet, no mHC or similar architecture, no engram. Just ordinary simple architecture and very good RL i guess.

What are you guys thinking about this?

▲
161
-4
16👁
r/LocalLLaMA · u/Akainu_Fan · 18d ago
Did Alibaba abandon 35B A3B?

Basically the title.We did not get a new moe model with qwen 3.8 and Alibaba did not announce any small moe models on apsara.I know we might get an announcement later but ngl I kinda lost hope