20 posts · 1 sub · RSS
← prev Wednesday, September 16, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
60
-1
26👁
r/LocalLLaMA · u/No_Run8812 · 24d ago
Upgraded my local setup with 2 rtx pros and it's amazing. post image

Follow up post of https://www.reddit.com/r/LocalLLaMA/s/nGMyKswrch.

Thanks everyone who replied. I didn't change the specs. Might be loosing some of the memory bandwidth but will scale in future if I need to.

It took me 2.5 days to build it because one of the GPU connected to PSU was loosing power whenever I load anything on the GPU, so I had to rewire every connection again to identify the fault. I am glad the system is working because I was apprehensive if this will work (I am software dev, getting my hands dirty with hardware for the 3rd time in life). My finger tips still hurt from pulling the cables from motherboard and PSUs.

To enable the full potential of the system, I had to enable peer to peer communication between the GPUs, cuda graph, tensor parallelism. I have capped both the GPUs at 500W (no reason, just didn't want GPUs to run on its full capacity).

Also, I had to open my box, because temps were shooting high, and fans were making weird noises.

I am running:

  1. Qwen 3.8 flash next 8 bit
  1. Deepseek v4 flash 0731 (official)

I have a M3 ultra 512, LLMs run on it, but I personally find it useless for inference. My head just hurts watching it work slow. On the other hand this new system is killing it, decode 150 tk/s and prefill 10K tk/s.

Qwen is good, but most of the context is consumed by thinking tokens, I was checking if it's a good idea to not use the thinking token. I barely have vram left for concurrent requests with full context window. Loving the Deepseek 1M context, and I also have room for 4 concurrent requests. Both of them are okay model, even if they make mistakes, I don't notice because of the speed. It's just fast, makes an error, corrects it moves on.

Finally the day is here when I can save on monthly subscriptions and not worry about the weekly or 5 hours limit. I have already setup my server with openclaw, opencode, openweb UI and Tailscale.

Has anyone experience excluding the thinking tokens of Qwen from the context and keep the final result? Was there any impact on the performance or accuracy of the model?

Any suggestions, what else I should install on it? Any new models to try?

▲
168
+2
22👁
r/LocalLLaMA · u/Mysterious_Hearing14 · 24d ago
Openjev post image

https://huggingface.co/AlexWortega/openjev

I build an openjev, it can play games and do everything what jev can. and yes - it's trained as crossencoder

▲
62
-1
29👁
▲
519
+5
38👁
r/LocalLLaMA · u/GuiltyBookkeeper4849 · 24d ago
Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis

I let Qwen 3.8 27B 4bit quantized with 100K context window run autonomously for 63 hours (50 million+ tokens) to try to solve the RH.

Of course it did not solve it, but the experiment still shows it's internal work, memory organization, strategies used and more.

The interesting thing is that it never hallucinated an answer and never stopped trying new ideas to solve it.

Multiple times it corrected it's own mistakes.

I am really hopeful that one of the unsolved millenium prize problems will be solved by an agent or a swarm of agents powered by an open source model in the next 12 months.

If you want to check out it's internal memories, code, strategies and more I published everything on HF: https://huggingface.co/datasets/gr0010/artificium-riemannhypothesis-experiment

My next goal is to actually use an agent perhaps powered by a smarter open model like GLM 5.3 flash or a swarm of agents, to solve an open math problem.

Please let me know if you tried something similar, what problem you'd suggest to tackle next, and if you have any question.

If you have GPUs consider getting in touch with me, we could run multiple agents to create a swarm and get them to tackle a simple yet open math/coding problem.

💬 193 (+2) open on reddit ↗
▲
463
+6
24👁
r/LocalLLaMA · u/skeole · 24d ago
Xiaomi MiMo 2.6 Live Training Dashboard

Cool to see this as it happens!

▲
123
 
22👁
▲
422
-5
25👁
r/LocalLLaMA · u/Porespellar · 24d ago
Frontier LLM development simplified for politicians: post image

Nobody is buying this “Pace the frontier” nonsense. It makes no logical sense at all. Are American labs really going to take a pause and lose any small lead they still may have over Chinese labs? Does anyone really believe this? This seems like some performative virtue signaling BS. Why are they bothering with this pacing campaign? Someone please explain.

▲
1273
+6
21👁
▲
76
+1
26👁
r/LocalLLaMA · u/Hefty_Wolverine_553 · 24d ago
What's the current best LLM uncensoring method?

With the recent Nvidia Huggingface acquisition and frontier AI labs screaming about safety and putting guardrails everywhere, I think it's important that we have local models that aren't affected by arbitrary guardrails set during training. To be clear, this post NOT about Enterprise Resource Planning (ERP). Censorship in an LLM can highly affect its abilities to do many legitimately useful things (note GPT-OSS, Fable 5), and going forth I believe censorship will only get more and more strict.

There have been many resources and posts about uncensored models using abliteration, heretic, and probably many other methods that I'm not aware of. However, it seems like all of this information is scattered about the place, and Huggingface is essentially flooded with "uncensored" variants of basically every popular open source model, many of which don't work well, affect the model's intelligence greatly, and have "KLD 0.0001" presumably from measuring against Wikitext datasets. I'm hoping that this post can gather some more useful information to serve as a starting point/discussion of which uncensoring methods work best.

Please share your experiences with specific uncensoring methods (not just a single uncensored model) and how well they work (both good and bad), as well as any notable people doing consistent/high quality work on uncensoring models.

▲
94
-3
24👁
▲
94
-3
22👁
r/LocalLLaMA · u/SomewhereAtWork · 24d ago
LocalJev?

Jev is a model to produce structured output (choices) from input text. It apparently can play (not run!) Doom.

https://typesafe.ai/blog/introducing-system-one-models-and-jev

Is there already a open implementation of this kind of model?

▲
54
-3
26👁
r/LocalLLaMA · u/WebAssemblyMan · 24d ago
Recurrent Looped Transformer post image

Recurrent Looped Transformer (RLT)passes the decoder's final hidden state to the next token, together with that token's causal encoder representation. The decoder reads encoder-derived global KV memory and maintains a sliding-window attention (SWA) cache at every layer. The same update runs over prompt and response tokens.

More effective reasoning depth!

https://github.com/yifanzhang-pro/recurrent-looped-tranformer

▲
146
-3
24👁
r/LocalLLaMA · u/sadnessdevil · 24d ago
You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

I'm pretty sure it can be done with any model based on qwen4exp, which Qwen's next local models will be based on. You can use a quant that barely fits in VRAM and still run at the model's maximum context length without kv cache quantization, since most of the KV cache can live in system RAM.

I actually made it working on vLLM and now I get 1M context with 3x 3090. I get \~80 tok/s at short context, dropping to \~60 tok/s once QSA reaches its 2048-token budget, after which decode speed stays flat as total context grows. The throughput is pretty good too, and I get like 150tk/s @ 4 concurrent requests. Prefill at 248k reaches 3,701 tok/s. (The patches and the model are available on my huggingface page if you're interested)

Decode speed is a bandwidth problem. Each decode step produces one token, and to produce it the GPU reads every weight and every piece of attention state that the step needs. On a single stream the card spends most of the step waiting for memory rather than computing. So the size of that per-step read sets the token rate.

This is why a normal model keeps its KV cache in VRAM. Take Qwen3.8-27B, which is built on the Qwen3-Next architecture and shares most of its properties with Qwen3.8-Flash-Next (\qwen4\_exp\). It still has one full attention layer every few layers, and a full attention layer reads its entire KV cache on every step. That read grows with the context, so decode gets slower as the conversation gets longer. It also grows past what any host link (such as PCIe) can carry, so the cache has to sit next to the compute.

The numbers of this model show the size of the problem. One QSA layer holds 2 key/value heads of 256 dimensions, as K and as V, in 2 bytes each, which is 2,048 B per token. At 262,144 tokens that is 512 MiB for one layer, and 6 GiB for all 12 layers on every single step. A PCIe 4.0 x16 slot carries about 32 GiB/s, so a host-resident cache of that shape allows about 5 tokens per second.

Here's an interesting part, Qwen3.8-Flash-Next avoids this in two ways:

Only 12 of the 48 layers have a KV cache at all. The other 36 layers are gated delta-net layers, a linear attention whose recurrent state has a fixed size. That state does not grow with the context.

Those 12 layers also do not attend over the whole context. QSA runs a cheap indexer over a pooled, compressed key, where \indexer\_head\_dim=128\ divided by \indexer\_compress\_ratio=4\ gives the pooled width. The indexer selects at most \indexer\_budget=2048\ positions. The layer reads the main KV rows only for the positions that the indexer selects.

So \indexer\_budget\ bounds the bytes that a decode step reads, and the context length does not:

\\\`

2048 selected x 2 kv heads x 256 dim x 2 (K and V) x 2 B = 4 MiB per layer

x 12 layers = 48 MiB per token

\\\`

Take an example, at 80 tok/s that is about 3.9 GB/s across the link. It is a small fraction of a PCIe 4.0 x16 slot, and most of it overlaps with compute.

Only few things need to stay on the GPU. The model itself, and a 2-byte slot plus the pooled index key, which is \1 x (128 / 4) x 2 B = 64 B\. Together they are 66 B per token per layer, against 2,048 B for a full row.

▲
73
-2
35👁
r/LocalLLaMA · u/BullfrogScary8947 · 24d ago
[Release] SOTA GGUFs for Qwen3.8-Flash-Next: GSQ-RCO Providing Near Baseline Performance

https://preview.redd.it/e5wn8eyh7vph1.png?width=1080&format=png&auto=…

https://preview.redd.it/8ov5gl8j7vph1.png?width=1080&format=png&auto=…

New Qwen3.8-Flash-Next quantization using GSQ-RCO. Cuts the size of Qwen3.8 Flash Next from around 80-95GB to 68-76GB, while still preserving near baseline quality. Also their Q2\_0 variant claims to be much faster offering 6.2x better prompt throughput in coding.

"Q2\_0 is built for speed. It avoids the quantization formats that rely on large lookup tables: those formats pack more accuracy into a given bit-width, but decoding them costs real time, and on this model that cost dominates inference. Q2\_0 delivers 3.4x the prompt throughput and 1.9x lower end-to-end latency than IQ2\_XS at a slightly smaller file size, and its decode rate stays flat across workloads instead of varying with the content. The trade is a little quality: 89.07 task average against 89.16 for IQ2\_XS, and 3.5 points below IQ3\_XXS. Pick it when throughput matters most, and see *Performance* for the measurements.

The IQ3\_XXS model is the strongest operating point: it matches the base model exactly on AIME25 (100.00) and is within 0.51 points on GPQA-Diamond and 1.14 on LiveCodeBench v6, at roughly one fifth of the BF16 size."

Model link: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

💬 54 (+1) open on reddit ↗
▲
208
-3
21👁
▲
536
-4
37👁
r/LocalLLaMA · u/RishiFurfox · 24d ago
Hey, Meta. Where's those Muse Spark weights? post image

It was well over a month since Meta promised to release the weights for Muse Spark.

Back then (10th August), they were on Spark 1.2. Now we're on 1.3 and still nothing's been released. So it begs the question: will they be releasing the 1.2 weights when 1.4 drops? Or will we get whatever's then-current as open weights?

It's ironic given Mark Zuckerberg said at the same time that we can't delay the release of models by "even a month," due to the competition with China. It's been well over a month. He was arguing in the context of new regulations delaying models, but I think it applies equally to the open weights contest as it does to the closed models one.

After all, the Chinese models are all open. That's the competition and point of comparison.

Have Meta given any sort of explanation for why they're sitting on the weights or how much longer it'll take for them to honour their promise? Will we even get them in light of all the attempts at regulatory capture and dire warnings about how AI is dangerous?

▲
60
+2
35👁
▲
114
+4
17👁
r/LocalLLaMA · u/EcstaticDentist · 25d ago
Qwen3.8 27b Game Dev Part 2 post image

Qwen3.8 27b may not be able to whip up 3d models & GLB’s but it will sure do with them as you please once you drop them in the game repo. Absolutely fascinating

▲
64
 
35👁
r/LocalLLaMA · u/Qwen30bEnjoyer · 25d ago
Open Source Appreciation Post

It's late at night in the lab, I've been working on a basic script for a virology project, and holy hell the safeguards have been pissing me off.

Mirroring detectEVE data over rsync to my laptop by making a zip file first? No no no, great safety mogul DARIO demands there be NO file transfer today. Request blocked, reported, labeled [cyber]. Yet, Deepseek V4.1 does it with no complaint.

I got tired of reading papers - so I ask Claude - "Does this PDF go over binary host virus infections?" Immediately blocked for biological safety risk. Deepseek V4.1 tells me it doesn't have the data I need without drama.

Bioinformatics server goes down and I need help getting it back up by getting the outputs of my diagnostic scripts to the mounted usb drive? Oops, its named exfil. Looks scawy. No transfer of output logs for you due to CYBER risk.

Would CNNs be a good architecture to start on phage-host prediction? Claude wouldn't tell me because information you can find in a google search is too dangerous for me to handle apparently - but once again Deepseek v4.1 tells me that GCNs are where I should start.

I get that Virology is a particularly sensitive topic, but come on. Imagine if Google had taken the same safety approach in the early days of search. Like if Google made it so that you either had to give up your identification and where you work to them, or go to the library and search by hand. It's almost unthinkable, yet in the name of the almighty Safety, Dario and Altman continue working to keep scientific knowledge locked away.

I think for the sake of all scientists, open source AI must win because we need a tool that just WORKS without egomaniacs micromanaging us or shaking us down for ID.

▲
95
+1
15👁
r/LocalLLaMA · u/MrMrsPotts · 25d ago
What's the next local model you are excited about?

And why?

💬 215 (+1) open on reddit ↗