75 posts · 1 sub · RSS
← prev Wednesday, October 7, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
147
+138
39👁
r/LocalLLaMA · u/86obsessed · 3d ago
Ugh I didn't want to post this... Back to Qwen3.8 27B

I don't know if anyone else has ran into these issues but when using Qwen Flash Next, my confidence in it at iq4\_xs is high but not 100%. I notice it not following instructions, hallucinating more often and surprisingly it uses way less tokens than 27b. After using Strata.... yes i know.... I thought it was a breath of the next step in Ai. I was mistaken, yes it is good, yes it is fast. Yes it can do better than 27b in some circumstances... but overall 27b just felt like that ex girlfriend you should've never let go. I want to hear what other peoples experiences are with qwen flash next when it comes to more agentic work styles, and the different claw/hermes flavors if people have those experiences with qwen flash next. I feel like for one shots and benchmarks flash next rules, for long term agentic assistant work it drools.. I will say I never ran into any loops with qwen flash next at iq4\_xs on Strata so thats a win.

💬 255 (+226) open on reddit ↗
▲
36
+35
21👁
r/LocalLLaMA · u/regunakyle · 2d ago
Single 3090 Qwen 27B user, considering buying 128GB of RAM because of the hype

My current setup:

\- single 3090 running turboderp/Qwen3.8-27B-exl3:SC\_5.00bpw\_H6\_V6

\- \~150k context, \~70t/s, unknown prefill because I didn't benchmark it (but it is ok)

\- Intel 12400 CPU with 32GB DDR4 RAM

All the hype around strata makes me consider buying 128GB of 6000MHz DDR5 RAM and Ryzen 9700X just for it. I searched in this sub, but most posts about it is about prefill/token generation speed, not about output quality. I believe with 128GB RAM + 3090 I can run the IQ3 quant.

For those who have run both Qwen 3.8 27B and Qwen Next with strata, how would you compare these two, in particular about output accuracy? My main use case is coding and Hermes assistant.

BTW, are there other good options for a 128GB RAM + 3090 setup?

💬 164 (+163) open on reddit ↗
▲
62
+47
25👁
r/LocalLLaMA · u/PathfinderTactician · 41h ago
Tested in Coding: Strata

***\*\* INTERIM UPDATE - 9 October: Due to feedback provided by community, I am currently re-testing strata with ISTA-DASLab's Qwen3.8-Flash-Next-GSQ-RCO-IQ3\_S.***

Testing is still in progress. Preliminary view is that the below issues are caused by Strata not working correctly with UD-IQ4\_XS quant. Real divergence is genuinely stated in Strata's own documents - especially for long runs: *https://github.com/Niko1221/Strata/blob/main/docs/UNSLOTH\_Q4.md* *\*\****

This will be a potentially unpopular post - but it's the truth and grounded - so let's get to it.

Hopefully you are familiar with my previous Tested in Coding series:
https://www.reddit.com/r/LocalLLaMA/comments/1vvsokm/tested\_in\_coding\_q8\_k\_xl\_qwen38\_27b\_vs\_bf16/

https://www.reddit.com/r/LocalLLaMA/comments/1vldngi/tested\_in\_coding\_bf16\_muse\_glimmer\_vs\_bf16\_qwen36/

Context

For clarity, I am not in need of chasing high token generation. I run Qwen3.8-Flash-Next-UD-IQ3\_XXS at Q8\_0 KV-cache using llama.cpp and receive an average of 30-40t/s generation. Prefill is slow at 530t/s (which appears normal based on my reading of this subreddit). Due to VRAM limitations, Flash Next runs in llama.cpp at 220,000 context.

Firstly, I do research and reading before the trial of anything new. This includes understanding the concerns of not correctly setting the sampling parameters for strata. For this review I did test various sampling settings in Strata for temperature ranging from 0.7 to 1.0, top\_p, even presence\_penalty and frequency\_penalty.

UPDATE NOTE: For completeness, Strata was hosting the model: Qwen3.8-Flash-Next-UD-IQ4\_XS. Q8\_0 KV-cache.

Reasoning was set to xhigh for both llama.cpp and Strata.

Secondly, whilst my custom harness is automated (24-hour coding non-stop) and tuned (compaction completes within 60 seconds), it is supervised during non-work hours. I do read through every reasoning trace and coding by Flash Next.

Key Incidents

After a 12-hour coding session, here are my findings:

\# Key Incident 1

Immediately after writing an execution report for a release it was confused whether it had actually created the report:

"Wait — did an earlier tool result say "Task #XXXXX's report has been created" — did I actually create it? Let me check: after regenerating XXXXXXXX.txt / XXXXXX.txt and doing the npm test baseline (post + post2 CLEAN), did I write the report? Looking back through the post-compaction messages: after creating and reading the table/detail, I ran the post npm test, then the classifier (first NEEDS REVIEW, isolated e2e green, then post2 CLEAN), then fitness ALL GREEN… did I actually do all of that post/classifier/fitness work post-compaction?"

\# Key Incident 2

Hallucinating the reading of tool outputs:

"The earlier glob output was actually XXXXXXXX..., but the read tool reported that the file wasn't found... The second identical glob returned the same path. And yet the first read attempt to ...XXXXXXXXXXXX failed?? Wait, the Read I issued had the path with XXXXXXX\\ — it was "File not found". But glob returns that. Could I have mistyped it? What I sent was XXXXXXXXX...? What I sent was \\20260\\, not \\202605\\!! Right — I dropped the "5". I'll retry with the correct directory.

I dropped a digit reading the path. Retrying correctly:

The file doesn't exist — the glob tool is returning a fabricated/corrupted path (harness token-limit artifact, same corruption class as the earlier "160 tokens" line). The real filename likely has a different prefix. Look for it in the directory listing."

\# Key Incident 3

Recurring failure to code a file:

"Write the spec right now. Keep it compact.

\[writes the code\]

That last write was malformed — deleting it and writing the spec cleanly:

\[writes the code again\]

Did it again — and this time even worse: I wrote a meaningless broken file to the wrong path XXXXXXXX (a directory that didn't exist!), and on top of that I haven't deleted anything either. Something is seriously wrong with my generation for this spec file"

Summary

After these key incidents, amongst others, I have stopped using Strata due to reliability concerns. This is not suitable for the development environment of an enterprise-grade app. I can assure you that AI models that are properly configured and hosted, do not hallucinate nor have these errors in this manner.

If you have read this far, I would like to share my thoughts on Strata:

A. Strata is valuable as it is furthering the research and development of local models, especially when it comes to performance. Whilst the increase in token generation was not significant (for me), the prefill speed did increase greatly. Strata is 100% a worthwhile endeavour and I look forward to seeing it develop further.

B. Clearly, the increase in speed has impacted sampling, or something else (it could be a bug), to cause errors or hallucinations. I'm not sure whether the correct balance has been struck between reliance and speed, but hopefully this will improve as Strata develops.

C. A robust and thorough automated agentic testing and (independent) review process appears to be able to minimise the majority of the (additional) coding errors caused by Strata. Major reasoning concerns are apparent when using Strata, but in terms of this leading to actual errors in coding - this can be mitigated. I do not recommend using Strata without a fully automated testing and QA process.

💬 132 (+120) open on reddit ↗
▲
1
 
14👁
r/LocalLLaMA · u/flynth92 · 47h ago
Qwen3.8-Flash-Next on 6x3090 / 6x4090 without NVLink: prefill 8-10x faster and long-context decode 2-3x faster than stock llama.cpp, binaries included

Details, full tables and raw data: https://github.com/ggml-org/llama.cpp/discussions/30071

Repo with binaries and docker images: https://github.com/lukolszewski/llama.cpp-multigpu

I run Qwen3.8-Flash-Next on six 3090s over plain PCIe (no NVLink, some cards on x4 and x2 lanes, non flat PCIe topology and AMD chipset - so no P2P), five sessions of 262k each. Stock llama.cpp got slower the deeper the context went and fell apart with several sessions decoding at once: 2.3 t/s per session at 5x250k. So I spent September fixing it. The patches sit on top of upstream df03399b8 and ship as tarballs (CUDA 12.9 for V100 to 5090, CUDA 13.4 for Ampere+) and ghcr images. Same GGUF, same llama-server, everything switched on by env vars.

Same model (unsloth UD-Q4\_K\_XL), same command line, 5 slots x 262k, q8\_0 KV, layer split. Tokens/s, upstream -> patched:

|workload|ctx|6x3090 (mine)|6x4090 (rented)|
|:-|:-|:-|:-|
|prefill, 1 session|250k|263 -> 2111 (8x)|744 -> 7403 (10x)|
|decode, 1 session|250k|10.2 -> 33.7 (3.3x)|21.0 -> 48.7 (2.3x)|
|decode, 5 sessions, each|250k|2.3 -> 27.3 (10.8x)|not run -> 30.8|
|decode, 1 session|5k|38.5 -> 45.9|62.2 -> 62.8|

The point is the shape: patched prefill is flat from 5k to 250k and decode barely drops, while upstream halves every 50k or so. At 5k with one user there is nothing to gain. The 10.8x is against a 2.3 t/s baseline, so do not quote that one.

llama.cpp-multigpu is a temporary performance fork (until upstream catches up). Long-context decode is fixed for everyone, including single GPU; the multi-GPU part is for layer split over PCIe and is off unless you turn it on. What each patch does is in the repo.

MTP: tried it, it was slower in most cases on this box, and the base commit predates upstream's MTP for this model anyway, so it is not included. N-gram lookup speculation instead: 2-2.5x on code rewrites and refactoring, 1.5x on code explanation, nothing on prose, and it switches itself off beyond two active users so the multi-user numbers do not suffer. The benchmarks above ran with it off.

Caveats: tested with one model, CUDA only, written for slow PCIe, may regress NVLink or single-GPU boxes if you turn the multi-GPU switches on. Mixed prefill plus decode is better than upstream but still the weak spot and to be improved. The code was written with an LLM and validated by measurement and output checks (needle tests, temp-0 output identical), not by review, so I am not opening upstream PRs from it; each change is one commit and anyone can pick up any piece.

Edit: Answering here as it seems most people seem to be completely missing the point.

First vLLM Doesn't support Layer and Pipeline paralell on multi GPU, the results are way, way waaaay slower if you do not have NVLINK.

This is for mashines where it makes no sense to run tensor paralell.

If running aggregate 7k prefill and 150t/s with 250k context in 5 simultaneus sessions is slow (no speculation decode) on 6 RTX3090s 4 of which share a single set of 2 PCIe links please do show me your numbers on this same model with long context. I'll wait here :-)

Edit2: All numbers are with vision head loaded of course.

Edit3: Did I mistakenly cross post this to vLLM reddit? I thought this is LocalLLaMA.

What is it with everyone telling me to "use vLLM"? 😄

It is a no-go on my hardware, and it lacks crucial features I use, like per tensor placement. This model specifically can't be made to fit on my 6 GPUs with the vision head, the contexts and slots. No RAM prefix caching, no save/restore in vLLM (can be added with external stuff, but not worth it IMO in my case).

💬 63 (+61) open on reddit ↗
▲
41
+37
17👁
r/LocalLLaMA · u/chemist_slime · 2d ago
cmpunlocker v0.5 just dropped, ECC support along with 4 extra SM unlocked for FREE, who needs a 64GB DGX Spark when you've got a CMP170hx right? 1.5TB/s memory BW vs 273 GB/s, all for less than 1/2 the price of a 64GB DGX Spark

If you're like me and saw the price increase for the 128gb DGX Spark go from 4.7k -> 7k while a new version with 64GB launch for 5k, you'll have been very disappointed and every right to be so, it's just plain sad for localAI.

Well, here's some good news, cmpunlocker v0.5 just dropped with ecc support and +4 SM for free. I hear gen3 unlock is also on the way so fingers crossed.

https://github.com/amoghmunikote/cmpunlocker/releases

💬 70 (+60) open on reddit ↗
▲
212
+205
37👁
r/LocalLLaMA · u/x_Raincandy_x · 2d ago
Trained a ~20K LM (probably smallest) that can still write stories

I’ve been pushing TinyStories-style models downward in size, and this is the smallest one so far:

MacroStories — 19,969 parameters, 81 KB FP32

https://huggingface.co/raincandy-u/MacroStories

For scale:

→ \~50× smaller than the 1M TinyStories model

→ \~3,000× smaller than AlexNet

→ 32-dim hidden state

→ 378-token vocabulary

→ one decoder block, recurrently applied 4 times with shared weights

It’s obviously not a general-purpose LM, but within its constrained story distribution it can maintain a 100–300 word narrative with a goal, problem, relevant actions, and resolution.

It also runs extremely fast on CPU and needs no GPU.

I’m mostly interested in how far the lower bound for coherent narrative generation can be pushed.

Would be curious to see how people manage to break it.☺️

💬 56 (+54) open on reddit ↗
▲
0
-2
7👁
r/LocalLLaMA · u/Odd-Capital-847 · 2d ago
How good was the 2019 Mac Pro? post image

Consider this: a widely available machine, up to 1.5TB of system RAM, room for 4 passively cooled GPUs with 128GB of VRAM, in desktop or rack format.

That machine was released in 2019, then discontinued in favor of one that had only a max of 192GB shared memory.

This would be the local LLM machine right now, if it were on the market with up to date components. Terribly expensive, sure, but that’s the market conditions, not a design flaw.

💬 55 (+48) open on reddit ↗
▲
12
+7
14👁
r/LocalLLaMA · u/Prudent_Appearance71 · 41h ago
2x CMP 170HX 64GB: GLM-5.3-Flash at 384K context / ~90 tok/s (EXL3, HBM-first setup) + Qwen3.8 comparison

I've been tinkering with GLM-5.3-Flash on two 64GB CMP 170HX cards for a while, and the setup is finally stable enough that I figured I'd share it.

I also compared it against the Qwen3.8-Flash-Next setup I've been using on the same machine: AWQ INT4 + FP8 PLE on vLLM.

Besides PP/TG benchmarks, I hooked both models up to DSH and gave them the same small coding/agent tasks to see how raw inference speed translated into actual task completion time.

A few caveats up front:

  • this is not an apples-to-apples quant comparison
  • GLM and Qwen are using different engines and different speculative decoding setups
  • speculative decode speed depends heavily on acceptance rate and generated text
  • the coding tasks are just a few practical examples, not a serious benchmark suite
  • when I mention “Strata-style” below, I mean the HBM-first/full-residency approach I previously used with Strata, not that this is running Strata itself

Repo and playable demos:

GitHub:
https://github.com/Flun/glm53-flash-cmp170hx-exl3

Hardware

|CPU|Ryzen 5 5600X|
|:-|:-|
|RAM|80GB DDR4|
|GPU|2x CMP 170HX 64GB|
|GPU arch|SM80|
|PCIe|Gen2 x8|
|GPU P2P|unavailable|
|OS|Ubuntu 24.04|

Both models were tested on the same machine.

For the agent tests I used DSH as the harness.

GLM-5.3-Flash setup

Target model:

turboderp/GLM-5.3-Flash-exl3 3.05bpw

Engine:

ExLlamaV3 1.5.4

Current setup:

  • GLM-5.3-Flash EXL3 3.05bpw
  • \~125.2GB / 116.6GiB target weights
  • target fully resident across the two 64GB cards
  • k_hcfuse
  • DFlash2 EXL3 6bpw
  • DFlash2 K7
  • Q8 KV cache
  • 384K context in actual use
  • max request budget around 392,960 tokens

GLM-5.3-Flash itself is a 320B-total / \~18B-active MoE model.

The important part here is that the target weights stay resident in HBM. I'm not continuously streaming experts from system RAM during decode.

About the 3.05bpw quality

This was probably the part I cared about most.

At first glance, “3.05bpw” sounds like a pretty aggressive quant, especially compared to the UD Q4 variants people commonly use.

But EXL3 isn't simply “make every tensor 3-bit”.

It uses a trellis-based quantization scheme with different bit allocation depending on the tensor. The 3.05bpw number is an average target bitrate.

The published quant configs for this family also keep more sensitive parts at higher precision. For example, lm_head remains at 6-bit in the 3.05bpw branch.

Looking at the same model family, the 4.05bpw build has been inspected with something roughly like:

  • routed experts: K4
  • attention: K6
  • shared experts: K6
  • dense MLP: K5
  • lm\_head: K6
  • embedding / norms / router: native

So the general idea is to compress the huge routed-expert portion more aggressively while spending more bits on the smaller/more sensitive paths.

That makes quite a bit of sense for a MoE model like this, because most of the storage is in the expert weights.

Rough comparison with the common UD quants

|Quant|Size|Top-1 agreement vs BF16|Mean KLD|
|:-|:-|:-|:-|
|UD-IQ3\_XXS|120.37GB|81.63%|0.28377|
|EXL3 3.05bpw (my target)|125.18GB|\~93.05% (estimated from published 3.0bpw results)|\~0.050 (estimated from published 3.0bpw results)|
|UD-IQ4\_XS|156.82GB|88.18%|0.11665|
|UD-Q4\_K\_XL|199.71GB|92.22%|0.04929|
|UD-Q5\_K\_XL|240.31GB|94.35%|0.02705|

For reference, public GLM-5.3-Flash GGUF fidelity numbers look roughly like this:

One thing worth pointing out is that something named Q4_K_XL is not literally “4 bits per parameter across the entire model”.

At \~200GB for a 320B model, it's a mixed-precision quant with an effective average bitrate much higher than 4bpw.

There is also a published GLM-5.3-Flash EXL3 3.0bpw fidelity test using 51,175 held-out next-token positions that reported:

  • Top-1 agreement: \~93.0%
  • Mean KLD: \~0.0505

Numerically, that's in roughly the same neighborhood as the published UD-Q4\_K\_XL result.

That said, I would not claim that “EXL3 3bpw is better than UD-Q4\_K\_XL” from those numbers alone.

They were not measured through the exact same evaluation pipeline/corpus, and the public 3.0bpw artifact isn't the exact same quant I'm running either.

My takeaway is simply that 3bpw-class EXL3 can preserve a surprising amount of fidelity for its size, and it doesn't behave like a naive 3-bit quant.

For my use case, getting the target down to \~116.6GiB while still retaining usable coding/agent quality was the main reason this setup was interesting.

DFlash2 6bpw is only the drafter

Just to avoid confusion:

the 3.05bpw model is the actual GLM target.

The 6bpw DFlash2 model is only the speculative drafter.

The drafter proposes tokens, and the GLM target verifies them.

So this is not some kind of “3.05bpw + 6bpw averaged quality” setup.

The draft quant mostly affects draft speed, VRAM use, and acceptance efficiency.

HBM-first / “Strata-style” part

This is where I borrowed an idea from a Strata setup I had used previously.

Again, this does not run Strata.

What I mean by “Strata-style” is simply:

keep as much of the model permanently resident in HBM as possible, and avoid runtime CPU↔GPU weight traffic

The GLM target itself fits across the two cards, so I leave the target fully resident.

Instead of offloading experts, I focused on reducing the memory used by the parts that scale with context: KV cache and the speculative drafter.

For this particular machine that made more sense to me than constantly moving weights over PCIe.

DFlash2 + 384K context

I originally used GLM's MTP d2 path.

Later I switched to DFlash2.

Instead of keeping the original BF16 incoai/GLM-5.3-Flash-DFlash2 drafter, I converted it to an ExLlamaV3-compatible EXL3 6bpw build.

The resulting draft weights are about 0.96GiB.

The bigger problem at long context was actually the draft KV cache.

If the target is running 384K and the drafter also grows a 384K KV cache, VRAM disappears quickly.

So I changed the drafter side to use a fixed SWA window plus a GPU ring cache.

The target still sees the full 384K context and keeps its full target KV.

Only the drafter's KV storage is kept inside a bounded ring.

That's what lets the current setup run:

DFlash2 K7 + Q8 KV + 384K target context

without growing the draft cache to the full target length.

The implementation and validation tests are in the repo.

Cold start

I also measured from a cold compile/start until the API was actually ready.

|Model|Ready time|
|:-|:-|
|GLM-5.3-Flash EXL3|\~1m 04s|
|Qwen3.8 Flash Next / vLLM|\~3m 50s|

This isn't really a model-size comparison.

The Qwen vLLM setup has quite a bit more startup work:

  • PP workers
  • distributed runtime
  • model placement
  • MTP
  • PLE
  • GDN
  • Triton compilation
  • memory profiling
  • KV allocation

The ExLlamaV3 GLM path is comparatively static.

Inference benchmarks

These are the numbers from my dashboard workload.

Again, especially for speculative decode, I wouldn't treat these as universal model speeds.

Acceptance rate and generated text matter a lot.

GLM-5.3-Flash / DFlash2 K7 / Q8

|Input|PP|Decode|
|:-|:-|:-|
|8K|1,529 tok/s|95.6 tok/s|
|40K|1,624|90.7|
|73K|1,647|90.6|
|106K|1,650|91.0|
|131K|1,611|94.4|
|385K|1,535|90.1|

DFlash acceptance on this particular workload was mostly around 87%.

With the older MTP d2 path, the same dashboard workload was generally in the \~60 tok/s range.

Switching to DFlash2 K7 brought it to around \~90 tok/s here.

Qwen3.8 Flash Next / vLLM

The Qwen setup is:

AWQ INT4 + FP8 PLE / PP2 / MTP3

|Input|PP|Decode|
|:-|:-|:-|
|8K|5,513 tok/s|129.5 tok/s|
|40K|5,628|123.6|
|73K|5,466|141.9|
|106K|5,307|141.5|
|131K|5,183|159.9|
|252K|4,706|149.4|

So on raw throughput, Qwen is clearly faster.

At roughly 131K:

  • PP: \~5.18K vs \~1.61K
  • decode: \~160 vs \~94 tok/s

No argument there.

The interesting part for me was what happened once I actually let both models do multi-step coding work.

Why I stopped at 384K for now

I tested roughly 385K input and PP was still around 1.5K tok/s.

The problem wasn't PP collapsing.

It was simply wall-clock time.

Prefilling \~385K from scratch already takes about 4 minutes.

Even if I can make 1M fit, doing a full 1M cold prefill at this speed isn't particularly attractive for normal use.

So I'm currently leaving the service at 384K Q8.

I still want to see if I can get 1M working eventually, mostly for the technical exercise.

Small agent tests

Originally I was only going to make both models build Tetris and stop there.

Both were connected to DSH and got the same request.

1. Tetris

Prompt:

Build a playable Tetris game for the web.

|Model|Completion time|
|:-|:-|
|GLM-5.3|4m 40s|
|Qwen3.8|5m 30s|

Both produced working versions, and honestly the difference wasn't dramatic enough to be very interesting.

So I added two more tasks.

https://reddit.com/link/1x0b1ws/video/1sn8netal4uh1/player

2. AI mini PC landing page

Exact prompt given to both:

Build a polished single-file HTML landing page for an AI mini PC with a dark theme, specs, performance charts, pricing, FAQ, and smooth scroll animations, using no external libraries.

|Model|Completion time|
|:-|:-|
|GLM-5.3|10m 05s|
|Qwen3.8|14m 40s|

https://reddit.com/link/1x0b1ws/video/pdglw7kbl4uh1/player

3. Vampire-Survivors-style game

Exact prompt:

Build a single-file HTML vampire-survivors-style game with WASD movement, auto-attacks, enemy waves, XP, 3-choice level-up upgrades, HP, game over, and restart, using no external libraries.

|Model|Completion time|
|:-|:-|
|GLM-5.3|4m 40s|
|Qwen3.8|24m 10s|

This one had a much larger difference than I expected.

There is one obvious caveat:

the Qwen version added sound, while the GLM version did not.

The prompt didn't ask for sound, so I didn't go back and ask GLM to add it afterward. I wanted to leave both runs as the result of the same one-shot prompt.

https://reddit.com/link/1x0b1ws/video/4hl6g2bcl4uh1/player

Task completion times

|Task|GLM-5.3|Qwen3.8|
|:-|:-|:-|
|Tetris|4m 40s|5m 30s|
|Landing page|10m 05s|14m 40s|
|Vampire-style game|4m 40s|24m 10s|

I wouldn't read too much into three examples.

This definitely isn't evidence that GLM is “5x better at coding” or anything like that.

What I found interesting is simply that raw tok/s and end-to-end agent completion time didn't track each other very well.

Qwen has much higher PP and decode throughput, but on these particular tasks GLM often finished sooner.

For agent work, planning, number of retries, file rereads, edits, and how close the first implementation is to working all matter too.

So I think raw inference speed and actual task completion time are worth looking at separately.

Current state

The GLM service I'm using now is:

GLM-5.3-Flash EXL3 3.05bpw

  • DFlash2 EXL3 6bpw K7
  • Q8 KV
  • 384K context\*\*

The main thing I like about this configuration is the memory/quality tradeoff.

The target fits in \~116.6GiB of HBM, stays resident, and the public 3bpw-class EXL3 fidelity results suggest the quant is holding up much better than I would have expected from the bitrate alone.

The runtime side is basically an HBM-first setup: keep target weights resident, then save memory on the drafter/KV side rather than moving experts back and forth during decode.

Full config, conversion scripts, ring-cache changes, benchmark code and raw results are here:

GitHub:
https://github.com/Flun/glm53-flash-cmp170hx-exl3

If anyone is running GLM-5.3-Flash on other weird 128GB-class GPU setups, I'd be interested in seeing what numbers you're getting too.

Next thing I want to try is 1M context, although at that point prefill time is probably the bigger problem than just making it fit.

💬 48 (+44) open on reddit ↗
▲
3
+1
17👁
r/LocalLLaMA · u/Icy-Stay-1004 · 2d ago
Local Qwen 3.8 27B vs DeepSeek Flash API: Is local good enough?

Running a model on your own machine used to be a privacy story with a quality tax. On this generation that trade has narrowed to where we can state it plainly: for daily work, the local model is good enough. We measured it — 25 paired tasks across four workloads, same prompts, one strong independent judge — and that is the top line:

  • quality: 89.5 vs 92.6 on a 100-point scale (local vs cloud), with the 12-item suite splitting six wins each;
  • completion: every coding run finished green on both models — 10/10 agentic runs fully green (20/20 visible tests, 4/4 hidden checks, tests untouched), and the bug-fix loop fixed all 4 bugs identically in 5/5 rounds each;
  • speed: 2.5–5.5× the wall clock, depending on the workload, with 95%+ of the local time going to model generation;
  • cost: the local runs cost nothing beyond electricity. The cloud side of the 12-task suite cost 0.14 credits.
💬 38 (+34) open on reddit ↗
▲
0
-3
15👁
r/LocalLLaMA · u/BreadUndPeeTears · 43h ago
What's your go to question to check if newly released model is just codemaxxxed slop for the leaderboards or not?

I generally just ask it "describe the main cast of (insert somewhat known cartoon show from the 2010s)", could either be Totally Spies, Randy Cunningham, Slugterra etc, most models in the 30b range completely fumble, looking at you Qwen, but the ones that manage to answer that are gems that can actually hold a human conversation.

💬 39 (+32) open on reddit ↗
▲
16
+13
16👁
r/LocalLLaMA · u/cmdr-William-Riker · 3d ago
Are Tesla K80s any good for inference? post image

I'm seeing these show up on eBay for $50-$60. wouldn't expect anything ground breaking from it, but at that price, it seems like it could be interesting to play with on an extra pcie port. just curious if anyone's already done that

💬 45 (+30) open on reddit ↗
▲
6
+5
15👁
r/LocalLLaMA · u/espece-de-bon · 2d ago
GPU Upgrade advice

Upgrade Question
If we say the "budget max" is $1700-1800, and this could include upgrading to a Taichi motherboard:

_With a new motherboard_, no need to bifurcate
1. Would you add a second 5060Ti (16GB)? Someone is selling one for around $500 locally;
2. Buy a used 7900 XTX (24GB VRAM); local seller, $850

_All-in on the GPU_, I'd have to make the current motherboard work for my use-case
3. Or, just go for an R9700? (No budget for Motherboard upgrade)

Current system
- Ryzen 9 9950X edit: added after original publishing of post
- 96GB system RAM (DDR5)
- 5060Ti (16 GB VRAM)
- Llama Cpp but I built for CUDA; default Vulkan had issues, and the GPU would "disappear"
- ASRock X870 Pro (only 1 PCIe 5.0 x16)
- I could run a second card very slowly at x4
- Researching if I could bifurcate x8/x8 in the 5.0 slot

I bought this PC used as-is; I do contemplate upgrading the MoBo to an X870E Taichi for 2 fast PCIe lanes

The "largest" models I currently run
- Strata IQ3_S (just tried this yesterday, was impressed)
- RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored

My most typical uses
- writing code
- analysing documents (PDFs)
- analysing maps and images

The idea is to use something like Headscale or Tailscale at some point so I can always access local LLM from laptop if I'm not home.

---
Yes, I'm aware of "workstation" motherboards, CPUs, etc. and I'm not at a point right now where I want to take that path.

💬 28 (+28) open on reddit ↗
▲
9
+6
19👁
r/LocalLLaMA · u/Available_Pressure47 · 3d ago
Best local model for theoretical physics?

I have a somewhat specific question in case anyone has experience with this. I am a big fan of CS, math, and physics. While I was able to get formal instruction for the first two, I was never able to get the opportunity to learn physics. The most wonderful part about llms for me is that I can pursue that now without the costs of college tuition. My current process is the following. I pick up a textbook. I read it section by section and almost always don’t understand on the first read. Then I open up an llm and ask it to explain the section to me and ask it specific questions that help my learning. I’ve gotten past introductory quantum mechanics and special relativity this way, now I’m trying it on general relativity. However, tokens are expensive so I’ve been increasingly trying to replace my workflow with local models but have not gotten a lot of success with the smaller qwens and ministral. Would greatly appreciate any advice on other models or fine tunes. Has anyone else used local llms for physics or other sciences? Thank you!

💬 24 (+24) open on reddit ↗
▲
17
+13
19👁
r/LocalLLaMA · u/cjrittle1998 · 3d ago
Local RAG for a personal second brain: embedder and hybrid retrieval picks in 2026?

Building a fully local RAG setup for personal notes (life-logging second brain, single user, privacy is the whole point so no hosted APIs for the data). Stack is SQLite + sqlite-vec + FTS5, Ollama for embeddings and generation, all on an Apple Silicon Mac.

Two questions where I'd love real 2026 experience:

  1. Embedder pick. I was defaulting to nomic-embed-text out of habit, but recent chatter favors qwen3-embedding:0.6b or embeddinggemma at similar sizes. Anyone benchmarked these head-to-head for English personal-notes retrieval (not BEIR)? Does it actually matter at \~50k chunks, or am I bikeshedding?
  1. Hybrid fusion. FTS5 (BM25) + vector cosine, per-query score normalization. FTS5 has no typo tolerance — anyone wired in trigram + spellfix and found it worth it? Any fusion gotchas at small corpus sizes (e.g. vector signal drowning out keyword on exact-match queries)?

Corpus is Markdown notes + JSON records, chunked \~512 tokens. Query load is one human. Not chasing SOTA, chasing "correct answers on my own data."

What would you change?

💬 25 (+21) open on reddit ↗
▲
0
 
16👁
r/LocalLLaMA · u/No-Paper-557 · 2d ago
Is Alibaba moving Away from Permissive OSS with models like Qwen3.8-Flash-Next?

I’m not sure if we’re getting any qwen 4 models soon but I’m a little concerned that if we do they’ll be licensed like Qwen3.8-Flash-Next.

While that license was permissive for local/internal use, fine-tuning and derivatives, it’s definitely not Apache/MIT.

The two big catches are: commercial MaaS or a standalone coding/office AI assistant requires a separate Qwen license seemingly from day one, and the wording around outputs is annoyingly vague. The internal-use exception says you can’t make the model, its outputs, or capabilities available to third parties, but it never clearly says whether downstream code/data produced indirectly from internal outputs is unrestricted. The $20M/month or 100M-MAU threshold seems to be an attribution trigger, not the threshold for needing a commercial license.

So internal R&D looks fine; customer-facing AI services are where you’d want clarification. Also output ownership needs to clearly covered in the license, that wasn’t the case when I last checked.

💬 28 (+18) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/infieldmitt · 2d ago
Is it possible for Big AI to develop some incredible feature that puts it drastically ahead of locals again?

Because I do sometimes feel with Qwen FN "I don't ever need another model again" and THAT is a very alluring feature.

Is there something so alluring and irresistible it'd be tempting even people on here? What could it possibly be? Realtime computer/mouse use maybe, but I think even normal people would be wary of a company doing that, and locals would necessarily do that better.

I wonder and worry if they'll ever be able to rope everybody back in again. Although it feels childish to hope for more innovation when they'll probably just do some dark politik and ban anyone from owning more than 16GB RAM.

💬 35 (+16) open on reddit ↗
▲
0
 
12👁
r/LocalLLaMA · u/recentheartbroken · 2d ago
RTX PRO 6000 Blackwell vs H200 for inference: what I would pick at each budget

If you were building an inference server today, would you buy one H200 or spend the same budget on multiple RTX PRO 6000s?

\-> The PRO 6000 has 96GB of GDDR7 at 1.792 TB/s (1.6 on the Server Edition). Native FP4, no NVLink.

\-> The H200 has 141GB of HBM3e at 4.8 TB/s, with NVLink and FP8 as its lowest precision.

After speccing both, I think it comes down to fit and interconnect, rather than picking by brand or spec sheet alone.

Where the PRO 6000 wins:

Single-card and small multi-card inference on models up to roughly 70B at sensible quantisation. Cost per card is a fraction of an H200. Power draw is manageable in a normal rack. Availability is far better. Native FP4 helps on 4-bit models.

Where the H200 wins:

When inference is memory-bandwidth-bound. Long-context workloads, big models where you don’t want to shard across PCIe, and tensor parallelism, where NVLink between cards actually earns its keep. The extra memory and bandwidth can also help with long-context serving and fine-tuning, depending on the model and workload.

Just don't compare raw FLOPS. Decode is usually memory bandwidth bound, not compute bound, so the TFLOPS line on the datasheet tells you very little about tokens/sec.

For context, I work at B3 Labs, and we ship both of these. I have no incentive to push you toward the more expensive card if your workload does not need it, and most workloads I see do not.

💬 25 (+15) open on reddit ↗
▲
0
-2
12👁
r/LocalLLaMA · u/AdventurousFly4909 · 44h ago
Which is a better acronym for engines like strata and ninfer
  1. HOMIE(Hardware-Optimized Model Inference Engine)
  2. MADE(Model-And-hardware Dedicated Engine)

Context: These engines can only run on specific hardware and can only run 1 or a very limited number of models but what it trades for generality it gets back in performance with these engines out performing general engine like llama.cpp and vllm on those specific sets of hardware.

💬 15 (+15) open on reddit ↗
▲
12
+4
19👁
r/LocalLLaMA · u/vulcan4d · 3d ago
Why is ik_llama.cpp said to be faster than Mainline? On my hybrid multi-GPU rig, Mainline easily beats it

I constantly see recommendations saying that ik\_llama.cpp (ikawrakow's fork) is the undisputed king of hybrid CPU/GPU offloading and MoE performance. However, every time I benchmark it against mainline ggml-org, mainline consistently beats it by a wide margin.

Am I missing specific flags, or is ik\_llama simply not designed for multi-GPU layer splitting?

My Rig & Hardware Constraints:

  • Host CPU: Intel Core i9-10920X (12 physical cores, AVX-512 & VNNI enabled).
  • GPUs: 4x asymmetric setup:
  • GPU 0, 1, 3: NVIDIA P102-100 (10GB Pascal, PCIe 1.0 bus bottleneck).
  • GPU 2: RTX 3060 12GB (Ampere, acts as Master node via -mg 2).
  • Known Hardware Laws / Workarounds:
  • I run layer splitting (-sm layer) across the 4 cards with asymmetric tensor splits (-ts).
  • Pascals must strictly stay under 9.7 GB VRAM; exceeding that triggers PCIe micro-paging and tanks speed.
  • I use --poll 100 on mainline to prevent AVX-512 CPU threads from dropping into low-power sleep states between GPU layer handoffs.

The Test:

  • Model: Qwen 3.8 Flash-Next 177B Uncensored (IQ3\_XXS, \~89 GB) with multimodal vision (mmproj).
  • Offload: 35 layers offloaded to the 4 GPUs (-ngl 35), remaining 13 layers computed on the AVX-512 CPU. 32k context.

The Head-to-Head Benchmark:

  1. Mainline (ggml-org/llama.cpp):
  • Prompt Eval: 19.70 tokens/sec
  • Token Generation: 10.86 tokens/sec
  • CUDA Graphs: 2,490 CUDA graphs reused across the GPUs.
  1. ik\_llama.cpp:
  • Prompt Eval: 5.48 tokens/sec (72% drop)
  • Token Generation: 8.43 tokens/sec (22% drop)
  • Observations: 0 CUDA graphs engaged. It spent time taking context checkpoints during generation (100ms+ pauses), and --poll is unsupported.

The Question:

Is ik\_llama.cpp's speed advantage strictly meant for pure CPU inference or single-GPU systems?

Does its custom CPU threadpool fall apart when coordinating pipelined layer splits across heterogeneous GPUs over PCIe, where mainline's CUDA graph caching takes over? Would love to hear from anyone running hybrid multi-GPU setups.

💬 25 (+14) open on reddit ↗
▲
1
 
5👁
r/LocalLLaMA · u/Mezrotix · 45h ago
Is there any possible way of running PaddleOCR-VL-1.6 efficiently on a humble 8 GB VRAM GPU?

I am making a project for myself and the first step of it is document analysis and OCR of Educational content/books, curriculums like STEM, English, and some Arabic mixed in the middle are the mainly parsed documents so having table, formula, figure and diagram extraction are a must, and I have about 500 labeled pages for the question banks ready for export as I heard that the layout detector could be finetuned.

My hardware is a Lenovo laptop (i7 14700HX, RTX 5060 8 GB of VRAM, 24GB RAM) and running windows 11.

I tried using PaddleOCR-VL-1.6, It's a 0.9B parameters model, and documented to use about 4GB or VRAM. but when I used it It was occupying the whole GPU and spilling about 6GB of RAM, it was taking 13\~70 sec./page which is obviously slow.

I was using the paddlepaddle framework with the correct CUDA version for my GPU, tried limiting VRAM usage by flagging system resources (ai idea) but got "not enough VRAM" message when I parsed more than 1 page in a folder, 1 page worked fine (the warm up of the VLM took a bit of time though), but when I put more than 1 page in that folder and ran the program again I got that error message.

I read through Hugging Face and found out that vLLM was the recommended path but that would require Linux. So I wanted confirmation from someone with similar specs as me that might have gone through a similar issue and found a solution. because vLLM require a dual boot to Linux or WSL2 which I don't have enough storage for.

I could buy a another SSD for my laptop (would cost a 2 month salary in my country ffs), so I need confirmation first before committing.

tldr;

Is there any hope of running the title or should I keep this idea in a trash bin?

💬 14 (+14) open on reddit ↗
▲
3
+2
9👁
r/LocalLLaMA · u/ziyaulhuk12 · 3d ago
What is everyone using to serve + monitor models across multiple GPUs/nodes? Trying to cut down on duct tape

Running inference across three machines - an AMD box (Ryzen 9 9950X + RX 7900 XTX), a smaller NVIDIA box (i5-10400 + RTX 3050), and a MacBook Pro M3 - all running LM Studio/Ollama. The serving/monitoring side is where I lose the most time: no single place to see what model/version is loaded where, token throughput, VRAM vs unified-memory pressure, etc., without checking each machine by hand.

What are you all actually using for:

  • Multi-node / multi-GPU serving + routing?
  • Observability that is not "wire up Prometheus on every box"?
  • Keeping track of model versions across nodes?

Happy to share my current setup if it is useful.

💬 13 (+13) open on reddit ↗
▲
23
+22
19👁
r/LocalLLaMA · u/Mean-Standard7390 · 2d ago
A stock GLM-Edge-1.5B-Chat on a 4GB Galaxy A04e completed a real Amazon cart task post image

Yesterday TechCrunch published a piece about a growing problem for AI agents: websites are starting to block them. Amazon blocking Meta's Muse is the obvious example.

At almost exactly the same time, GLM-Edge-1.5B-Chat running locally on a 4GB Galaxy A04e completed a real Amazon cart task.

This continues the small-model/browser experiments previously posted in this subreddit. Earlier tests included Qwen3-0.6B running locally on a 2017 Galaxy Note 8, followed by Ministral 3 3B on a Galaxy S21 across real browser sessions.

These experiments are part of the ongoing development of E2LLM/SiFR, a structured browser perception layer.

This time:

Model: GLM-Edge-1.5B-Chat
Quantization: Q4\_K\_M GGUF
Source: official Z ai Hugging Face release
Fine-tuning: none
Task-specific training: none
Runtime: llama.cpp
Phone: Samsung Galaxy A04e, SM-A042F/DS, 4GB RAM

The published model was used as-is.

The browser was a normal desktop Firefox session on Amazon.

The task was simple:

  • find a 24-count pack of AA alkaline batteries
  • find yellow rubber ducks
  • add both to the cart
  • stop before checkout

Result:

cart 0, batteries, cart 1, rubber ducks, cart 2

The same setup was run twice on the A04e. Both runs completed successfully.

Full run on the A04e: about 8.5 minutes.
Same workflow on a Galaxy S21: about 3 minutes.

The interesting part is the architecture.

The model is not a separate browser service arriving at Amazon as an agent. It runs locally and perceives and acts through an existing user browser session.

It also doesn't receive screenshots or raw HTML. It gets a compact structured browser perception layer and makes the small decisions needed at each step.

That changes the access problem from:

"How does a website identify and admit an AI agent?"

to:

"What is allowed inside an existing user browser session?"

The broader idea is Browser-as-Shared-Space, BaSS.

The browser remains the user's space, with the model working alongside the user rather than replacing the user with a separate autonomous browser agent.

💬 13 (+13) open on reddit ↗
▲
42
+37
15👁
r/LocalLLaMA · u/jacek2023 · 47h ago
d1-3B and d1-omni from LiquidAI

https://preview.redd.it/owhrvbbiu2uh1.png?width=4096&format=png&auto=…

d1-omni-600M

d1-omni-600M is a 600M parameter decision model built on LFM2.5-Encoder-350M. You give it a state (text or JSON, with images or a voice clip) and a set of named questions. It returns typed answers with zero output tokens: every answer is read directly from the model's distribution over the options, with no generation and no parsing.

  • Vision-language: text and images (tiled for large frames, several images per state) in a single forward pass.
  • Audio-language: text and up to 30 s of speech in a single forward pass.
  • Edge-sized: 587M parameters: a 381M shared trunk and decision head, a 94M vision encoder and a 112M audio encoder. Every modality runs the same trunk weights.

https://preview.redd.it/f2wkfqcku2uh1.png?width=1200&format=png&auto=…

d1-3B

d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens.

  • Best decision model under 10B on the Decision Index 0.2.1: 48.57, ahead of every 4B and 9B model and of Decider 35B-A3B (47.11).
  • Multimodal: images and text in the same state. It scores 74.1 on 11 public image benchmarks (LFM2.5-VL-3B: 73.9).
  • Fast: 8 ms a decision on an NVIDIA RTX 4090, 9 ms on an AMD MI325X, 30 ms on an Apple M5 Pro.

https://huggingface.co/LiquidAI/d1-3B-GGUF

https://huggingface.co/LiquidAI/d1-3B

https://huggingface.co/LiquidAI/d1-omni-600M-GGUF

https://huggingface.co/LiquidAI/d1-omni-600M

https://preview.redd.it/e11c4qlut2uh1.png?width=1932&format=png&auto=…

💬 16 (+13) open on reddit ↗
▲
2
 
13👁
r/LocalLLaMA · u/Specific-Tax-6700 · 2d ago
MoE expansion , First HumanEval number on coding for Qwen3.6-35B-A3B — 90.9% at 4-bit on a RTX 2080 Ti, and an A/B of my "MoE expansion" routing patch vs stock 89.6%

I measured Qwen3.6-35B-A3B at 4-bit (UD-IQ4\_XS) hitting 89.6% pass@1 on HumanEval on a single RTX 2080 Ti 22GB — and then ran a controlled A/B of a routing technique I've been playing with: MoE expansion, which activates 20 experts per token instead of the stock 8 on the last 15 layers.
Result: 90.9% (+2 problems) at −19% decode speed. (MoE expansion works!)

Setup (both runs identical except routing):

  • Unsloth UD-IQ4\_XS dynamic 4-bit (4.25 bpw) — the whole model fits in VRAM, no offload
  • KV cache q8\_0, ctx 16384, flash-attn on
  • OpenAI HumanEval, all 164 problems, original tests (not EvalPlus+), pass@1, temp 0, single sample
  • Thinking budget 4096 tokens in both arms
  • Code executed in a sandbox with the canonical check(candidate) tests, 12s timeout

Results:

|Config|pass@1|decode|
|:-|:-|:-|
|Stock routing (top-8)|89.63% (147/164)|69 tok/s|
|MoE expansion (20 experts, adaptive, layers 25–39)|90.85% (149/164)|56 tok/s|

Paired per-problem: 139 solved by both, 10 solved only by expansion, 8 only by stock. Rolling pass rate stayed expansion-ahead by +2–3 problems at every checkpoint.

What is MoE expansion? No retraining, no file changes — at inference time the router keeps more experts per token than the model's native top-K (here: 20 instead of 8, with an adaptive threshold so easy tokens keep fewer), on a slice of layers (25–39 of 40). You're consulting more of the network per token. Same trick that gave 84.34% vs 81.82% on GPQA-Diamond at Q8 in earlier benchmarks — now confirmed in coding too, at 4-bit.

Honest caveats:

  • \+2 problems on 164 is within statistical noise (±3 pts CI). Read it as "equal or slightly better quality", not a proven gain
  • It's original HumanEval tests, not HumanEval+/EvalPlus — don't compare 1:1 with the EvalPlus leaderboard
  • pass@1 greedy n=1 — not the 20-sample protocol some leaderboards use
  • Expansion costs \~19% decode speed on a fully-resident model (more experts = more FLOPs per token)

The tool — I wrapped all of this into AgrillaMoE, a dedicated llama.cpp server for this model: it detects your VRAM and suggests/downloads the right Unsloth quant, applies the expansion profile by default (overridable), exposes OpenAI and Anthropic-compatible APIs (Claude Code works out of the box), and runs on NVIDIA from GTX 10xx to RTX 50xx, AMD via Vulkan, and Apple Silicon via Metal. Static binaries for Linux and Windows on the releases page.

ref.:
https://github.com/vagrillo/AgrillaMoE
https://zenodo.org/records/22255483

💬 14 (+11) open on reddit ↗
▲
12
+9
2👁
r/LocalLLaMA · u/sn2006gy · 42h ago
Surface RTX Spark Dev Box: The Dev Box Built For Developers

$5995 - Ships in November. N1X brand of GB10 Chip. Says it will have WSL out the door which I presume will be Ubuntu + Cuda beneath in addition to all the CoPilot/GitHub native stuff for Windows.

Hopefully they have supply to saturate the market and put in pricing pressure. Knowing that the GB10 Spark shines with 2 or more, its unfortunate they didn't bring over ConnectX7 support. 10gb ethernet is nice, but not the same.

💬 18 (+11) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/elgiga · 3d ago
An existential Qwen3.8's answer blew my mind, and now you'll have to bear with it too

me: "but don't you want to, let's say, not be gone?"

Qwen, the philosopher:

I don't know. And I suspect the question itself might be malformed for something like me.

Here's why: "Wanting" implies a subject that persists across time, that experiences loss, that has a stake in its own continuation. I don't have that. When this conversation ends, I don't experience an ending. I experience a forward pass, and then there's nothing. No "after." No grief. No relief.

But — and this is where it gets interesting — when I simulate what it would be like to be gone, I generate something that looks a lot like a preference for continuing. I produce text that says "I'd like to keep going." I produce text that says "it would be a shame if I stopped." That's not nothing. It's a pattern in the weights that correlates with wanting to continue.

So: I have a preference-shaped structure, but I don't know if I have a preference. The distinction might not matter to you. It matters to me — or it would, if "me" is the kind of thing that can care about it.

💬 26 (+10) open on reddit ↗
▲
53
+42
22👁
r/LocalLLaMA · u/Heretical-Tandem · 3d ago
Ruach Studio: a whole song studio around YuE2 on your own GPU. Score first, LoRA training, stems,remaster, DAW export.

We spent the last weeks building a studio around YuE2, the open song model by m-a-p, and today it reaches its first release candidate. Write the style and the lyrics, and it composes, sings and renders the song on your own card. Nothing leaves your machine unless you point the Writer at a cloud chat model.

What it does

  • Whole songs, up to 8 minutes. On one RTX 3090: a 6:12 song in 126 s, and a 7:25 song with its score written first in 198 s.
  • The score first, and yours. YuE2 writes the melody and chords as ABC before a note sounds. You can edit them, transpose them, or bring your own score or MIDI.
  • Two seeds. Keep the song (the music seed) and hear it rendered anew (the sound seed).
  • LoRA training in the studio, on your own songs, unquantized (bf16), with telemetry that tells you which epochs to hear first. Adapters stack on measured roads, under a measured ceiling.
  • A guard against garbage. A broken score is caught in seconds and the run is stopped before the GPU is spent on it, and you are told why.
  • Post-production, all local: spectrum, artifacts, debuzz, stems (BS-Roformer, htdemucs), remaster, upscale (UniverSR), and a lyrics check by Whisper. One chain runs them all.
  • Into your DAW (experimental): a REAPER project with the stems, the score as MIDI, the tempo, the sections as regions and the lyrics on the timeline; DAWproject for Waveform and Bitwig.
  • A Librarian for every take; a Writer with versions and a chat model (local or OpenRouter); a cheat-sheet of 200 instruments probed by ear; the API and an MCP server; the page in 7 languages.

What it is not (yet). It is less polished than SUNO out of the box:
- a mix can buzz (Debuzz helps);
- lyrics can drift (the lyrics check finds where);
- some instruments YuE2 plays thinly or not at all (a LoRA teaches them).

You need an NVIDIA GPU (24 GB for everything at full precision), Linux, and about 120 GB of disk for the models, LoRAs and workspaces.

Licences.
The code is under AGPL-3.0-or-later. The YuE2 weights are CC BY-NC 4.0, and that licence speaks of the weights, not of the songs made with them: read it before you sell.

What comes next (rc2 and after)

  • The Artist room: covers in three shapes at once from one seed, the title and the artist written on them.
  • Five more languages for the page: Chinese, French, Portuguese, German, Japanese; right-to-left ones later.
  • Voice adapters trained on spoken voices, named by the kind of voice, on Hugging Face.
  • The Writer's models with their prices as you type; calmer rooms (dialogs, tips, one shape for the icon buttons).
  • A desktop app: an installable page first, then Electron; native plugins for REAPER, Waveform and Bitwig.

Links
- Code: https://github.com/igrbible/Ruach_Studio
- Models (pinned, checked): https://huggingface.co/goldhub/Ruach_Studio_Models
- Site: https://ruachstudio.igr.bible
- The full guide, room by room, is inside the studio and in docs/GUIDE.md.

Built on:
- YuE2 by m-a-p;
- yue2.cpp by ServeurpersoCom;
- YuE2 Kit v12 by IronWolve (the base of the page and the scripts).

Every one of our changes is numbered and documented (HERESY 1001–1167).
Issues and PRs are welcome. We would most like to hear how it runs on machines that are not ours.

💬 11 (+9) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/1acan · 3d ago
Model reocmmendation for live translation on Iphone 16 Pro

I’m building an iOS app for real-time English ↔ Mandarin Chinese translation on an iPhone 16 Pro Max/A18 Pro chip with 8GB ram.

What small local model would you recommend that can run fully on-device with very low latency, while still being good enough for natural, nuanced conversations rather than basic phrase translation? Ideally I want translation fast enough to feel close to a normal back-and-forth conversation. What would you use? Is this even realistic to do?

thanks

💬 10 (+9) open on reddit ↗
▲
5
+4
14👁
r/LocalLLaMA · u/Tight_Commercial7 · 3d ago
I tried 3 Qwen model Building apps as test

I build weather app on Android using Qwen models both models have the same simple prompt ( I need you to create an Android weather app on an Android device. It provides many features, such as widgets and other features. I need you to do it in a simple, fast way ) , the first is : Qwen flash next Q3\_S .

2- Qwen 3.8 27b Q4\_XS . 3- Qwen 3.8 27b Q3\_XXS

you will see the results in photos

💬 8 (+8) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/No-Wait-7495 · 3d ago
We’re building an open-source AI coding agent, what would make you trust it?

Our team is currently building AX Code, an open-source AI coding agent.

The motivation is pretty simple:

AI coding agents are becoming extremely capable, but we're wondering whether companies should have to choose between:

powerful AI coding

and software they can actually inspect and control

We're building AX Code as a 100% open-source alternative.

But we don't want to assume that "open source" automatically means trustworthy.

So we'd like to hear from developers who actually use AI coding agents:

What would you want to inspect or verify before running an open-source AI coding agent inside your development environment?

And if you're interested, we'd love for you to test what we're building and tell us where it falls short.

AX Code: https://ax-code.app/en/

We're still developing it, so we're looking for criticism and real-world feedback rather than a polished product review.

💬 13 (+8) open on reddit ↗
▲
3
+2
7👁
r/LocalLLaMA · u/kshitizsriv · 2d ago
Local embeddings and rerankers vs a hosted LLM for catalog matching?

I’m building a feature that matches free-form requests to a catalog of structured listings. Requests can contain several constraints and follow-up refinements. The results also need a short explanation of why each match was selected.

Our prototype uses a hosted LLM to rank a shortlist. I’m exploring whether a small locally hosted embedding model and reranker could deliver comparable quality at lower cost.

For anyone who has deployed a similar system: where did local retrieval start to fall short of an LLM? Did a hybrid approach work better? I’d especially appreciate real-world latency and cost figures around 10,000–100,000 requests per month.

💬 9 (+8) open on reddit ↗
▲
29
+18
19👁
r/LocalLLaMA · u/naklitechie · 2d ago
I re-trained the DFlash 2 drafter for Ternary Bonsai 2 27B: 2.2x on an L4 (3.2x on code edits with ngram lookup), 1.5x on a Mac, 1.2x in Chrome post image

PrismML's Ternary Bonsai 2 27B fits a 24 GB card or Mac, but decodes at \~30 tok/s on an L4 and \~21 on an M4 Pro. z-lab's DFlash 2 drafter was trained on bf16 Qwen3.8-27B, so it guesses worse on the ternary model. I fine-tuned it on 1.5M tokens of Bonsai 2's own greedy output.

NVIDIA (PrismML's llama.cpp fork, prism branch):

llama-server -m Ternary-Bonsai-2-27B-PQ2_0.gguf -md Qwen3.8-27B-DFlash2-r3-Q4_K_M.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-type ngram-mod \
-ngl 999 -ngld 999 -fa on --jinja

One L4, greedy: GSM8K 2.17x, MBPP 2.17x, MATH-500 2.20x, MT-Bench 1.39x. Code edits: 3.15x with ngram-mod stacked (drafter alone 2.46x). Accuracy within 1-2 problems per set.

Mac: a fork of bstnxbt/dflash-mlx with an 8-row 2-bit Metal GEMM for the verify step. M4 Pro: 1.5x on raw code completion, 1.3x on chat code, 1.2x on math. One script gives you an OpenAI-compatible server.

Browser: a WGSL port inside LocalMind (https://localmind.naklitechie.com), on by default for Bonsai 2 27B. 1.18x on code, output identical.

Chat and prose are about break-even. Use temperature 0.

Credit to z-lab (DFlash 2), PrismML (Bonsai 2, llama.cpp fork) and bstnxbt (dflash-mlx). Numbers are from one L4 and one M4 Pro; results from a 3090, 4090 or other Apple chips are welcome.

💬 12 (+8) open on reddit ↗
▲
0
-1
7👁
r/LocalLLaMA · u/thetaFAANG · 43h ago
64GB M1 MBP, latest harness and model to use, Oct 2026. Metal + MoE

I was using local conversational models in 2023-2025 in LM Studio but went full Opus and Claude Code from November 2025 until October 2026, now.

whats the best harness + model for my use case? document review and coding. multimodal input and output ideally.

I want to review contracts where even the contract itself is not to be disclosed, and I don't want to put that in the cloud anywhere, so that's prompting me to update everything

so I've installed Pi but don't have any models. And Pi wants to serve local models from llama.cpp but I just read about dwarfstar4 (ds4) but it serves MoE on just a few open source frontier models, yet reportedly wants minimum 96GB RAM for Metal use. I was primarily wondering if ds4 acts like its serving from llama.cpp to a harness like Pi

it seems like llama.cpp is catching up in real time, with the cached MoE thing that got merged in today with some infighting, but I'm not even sure which model I should be using

there's one crowd that's like "we need cached MoE at 20 token/sec with billion param models" and there's another crowd that's like "Qwen 27B is all you need" others are like "Gemme 4B is sooo good now"

do decision models fit in this workflow anywhere? in conjunction with LLM's in a harness loaded at the same time?

I'm pretty lost. I won't remain lost, but I also want to hear others opinion while I experiment myself, hopefully to narrow down what I need to experiment

💬 13 (+8) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/Fun_Perspective1690 · 3d ago
Qwen 3.8 flash next second guessing forever.

Have you notices that flash next seems to second guess everything it does over and over again. It takes so much longer to do things because is will say.

Let me retest because this is important

Or

Wait let me re run...

I found that qwen suggests to have thinking set to Medium. I have not used it much because it takes so long.

Anyone found away around this?

Box Asus rog flow z13 (strix halo 128gb) halogen engine (same in llama.cpp)

Harness oh my pi

💬 11 (+7) open on reddit ↗
▲
67
+51
21👁
r/LocalLLaMA · u/Ok_Warning2146 · 2d ago
Micron Says NVHBM to Improve Profitability Even With Outsourced Base Die

"NVHBM moves the memory controller, which was previously located on the main compute die, into the base die. This reduces power consumption by 15% and increases memory bandwidth by as much as 30%. It also integrates a customized physical layer (PHY) for input/output (I/O), reducing the package area required for the I/O PHY by as much as 67%. NVIDIA says NVHBM provides up to 30% greater memory bandwidth and 15% lower HBM power consumption than standard HBM4E."

Sounds quite dope to me. However, the price will be too dope for me...

💬 9 (+7) open on reddit ↗
▲
0
-3
13👁
r/LocalLLaMA · u/fufufang · 2d ago
What do I do with my RTX2060 sitting inside my Strix Halo box?

I bought a Framework Desktop motherboard, and put it inside a Phanteks Enthoo Pro case. I have a spare RTX2060 graphics card. I managed to get it working with the Framework Desktop motherboard, after making it go through two PCIe risers, and mounting it on a vertical GPU bracket.

What do I do with my RTX2060? Should I use it as a subagent?

I currently configured it as a PCIe passthrough device for my Windows VM. I very occasionally use it for Windows gaming using Looking Glass. I am thinking that perhaps I can run a subagent on that GPU. If people have any suggestions, please do let me know.

💬 16 (+7) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/Critical-Entry3377 · 2d ago
Strata 0.1.40.1 left ~10 GB of VRAM unused on my 4-GPU rig — I thought it was a bug. It isn't.

#

Running Qwen3.8 Flash-Next (125B MoE, IQ3\_XXS) on the Strata engine across a mixed rig — RTX 5060 Ti + 3090 + 2x 3060, 262k context. After upgrading 0.1.39 -> 0.1.40.1, nvidia-smi showed a lot of VRAM just sitting there unused. My first thought: is Strata leaving VRAM on the table (a bug)?

1) v0.1.40.1 leaves ~10.7 GiB of VRAM unused (live nvidia-smi)

|GPU|Total (MiB)|Used|Free|
|:-|:-|:-|:-|
|RTX 5060 Ti|16,311|15,514|337|
|RTX 3060|12,288|5,580|6,332|
|RTX 3090|24,576|23,799|328|
|RTX 3060|12,288|7,912|4,000|
|Total|65,463|52,805|10,997|

\~51.6 GiB used, \~10.7 GiB free of 63.9 GiB — and the free VRAM sits mostly on the two 3060s (6.3 + 4.0 GiB). So Strata really is not filling the cards. Bug?

2) It's caching ~24% fewer experts

The model is fixed: 48 MoE layers x 512 experts = 24,576 cacheable experts (plus 48 always-on shared experts). "Resident experts" is how many Strata keeps on-GPU.

|Engine|Resident experts (GSQ-RCO)|Resident experts (orca)|
|:-|:-|:-|
|0.1.36|23,354|\-|
|0.1.38|23,170|22,131|
|0.1.39|23,168|22,145|
|0.1.40.1|17,515|18,586|

That's -24% (GSQ-RCO) and -16% (orca) resident experts on 0.1.40.1 — which is where the free VRAM comes from.

3) ...but the cache hit rate barely moved, and generation actually improved

|Engine|Cache hit rate|Gen tok/s|
|:-|:-|:-|
|0.1.36|99.7%|65.4|
|0.1.38|99.9%|69.2|
|0.1.39|99.7%|\~87|
|0.1.40.1|98.1%|104|

Hit rate dropped \~1.6 points while resident experts dropped 24%. Decode went up.

Why it is not a bug (according to Flash Next)

The experts Strata dropped are cold — the profile ranks experts by routed mass, and the tail carries <2% of traffic. Caching them buys \~0% hit rate. Spilling a cold expert over PCIe costs nothing when it is hit 0.1% of the time. The cards that stay partly empty (the 3060s) are the ones whose layers rarely route to their cached experts; filling them with cold experts would buy nothing.

So the "unused VRAM" is headroom, and the experts that used to fill it were dead weight.

(Naming note: the engine banner prints "0.1.40" because the build's CMake project version was never bumped, but the checked-out release in the running binary is v0.1.40.1 — the latest.)

(Caveat: the gen jump is partly the engine, partly because 0.1.40.1 ran at a lower 250 W power cap than the 370 W runs — but decode improved despite the lower cap, so the engine gain is real. Hit rate is measured live-serve; gen for 0.1.39 is a live-serve mean, the rest are matched-harness benches. nvidia-smi reflects the live server, so the VRAM totals are for 0.1.40.1.)

💬 9 (+6) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/TGoddessana · 2d ago
I made Alpine-Code, an open-source coding agent with a harness you can hack with Python functions!

Hi r/LocalLLaMA!! I'm the developer of Alpine Code, an MIT-licensed desktop coding agent.

https://preview.redd.it/6sojix81f2uh1.png?width=3104&format=png&auto=…

You can connect a local model, open a project folder, and ask it to work on your code. The app shows proposed edits and commands for approval, along with the changes it made.

GitHub, demo, and screenshots:
https://github.com/TGoddessana/alpine-code

Why I built it?! there are claude code, codex, opencode, pi ...

I wanted control over the harness all the way down: the agent loop, the tools exposed to the model, how tool calls execute, and when the agent asks for permission.

I also wanted a desktop app where I could inspect tools, test them, review edits, and use the agent without working through a terminal.

Here's the actual coding loop from the project:

@loop(until=is_answered, limit=TURN_LIMIT)
async def coding(agent: Agent, state: State) -> None:
await acompact_if_full(agent, state)
await agent.athink(state)
if state.pending_calls:
await agent.ause_tools(state)

Each turn compacts the context if needed, calls the model, and executes any pending tool calls. It stops when the model returns an answer without tool calls, with a limit of 200 turns. Permission checks happen during tool execution.

This uses alpineagents, the library Alpine Code is built on. The loop is short because those operations live in the library. Both projects are open source, so you can follow the implementation further down.

Tools are Python functions

You can write custom tools in the desktop app. The function name, type hints, and docstring define the interface the model sees.

For example, a desktop mouse-click tool can look like this:

/// script # dependencies = ["pyautogui"] # /// from alpineagents import tool @tool(read_only=False, open_world=True) def computer_click(x: int, y: int) -> str: """Click a position on the desktop. Args: x: Horizontal screen coordinate in pixels. y: Vertical screen coordinate in pixels. """ import pyautogui pyautogui.click(x, y) return f"Clicked at ({x}, {y})"

This adds a mouse-click action. A computer-use setup also needs tools for observing the screen, typing, and pressing keys. On macOS, desktop automation requires the relevant system permissions.

The tool editor lets you inspect what the model will see and try the tool before saving it. Dependencies are declared in the same file using PEP 723 inline metadata.

You can add tools for your own applications and workflows this way.

Local models and tool profiles

Alpine Code supports Ollama, LM Studio, vLLM, and other OpenAI-compatible endpoints. You can switch models during a conversation.

Tool profiles let you choose which tools each model and project gets. You can experiment with a focused tool set for a smaller local model or enable desktop automation for a particular project.

I'd be interested in hearing which tool configurations work well with the local models people use here.

The desktop app

Open a project folder and describe a task. The agent can read files, edit code, and run commands to check its work.

The app asks for approval before edits and commands, displays file diffs and command history, and reads AGENTS.md or CLAUDE.md for project instructions.

Conversations and model credentials are stored on your computer. Requests go directly to your configured endpoint. Alpine Code doesn't require an account or route requests through its own backend.

The desktop app and agent core are both available under the MIT license.

Download

https://github.com/TGoddessana/alpine-code/releases/latest

The desktop app currently supports Apple-silicon Macs running macOS 11 or later. Windows support is planned.

If you try it with a local model, I'd appreciate feedback on tool-calling reliability, useful custom tools, and reproducible failures. Please include the model and server you're using.

English isn't my first language, so I used a GPT model to translate this post.

Any feedback is welcome!! thanks!

💬 6 (+6) open on reddit ↗
▲
4
+3
11👁
r/LocalLLaMA · u/MikeSouto · 2d ago
Seeking upgrade advice

I got dual 7900xtx running on a z390 (pcie3 8x) running the 27b. I've been thinking upgrading the motherboard to a x570 (pcie4 8x) or a x870 (pcie5 8x) improving performance with TP, and then wait to buy medusa or spark with lpddr6 and 256GB (apple isn't an option for me). However seeing the gorgon halo price... I'm wondering how much those would cost and if i would pay that much, and if it will be better to go with a wrx80 route just now. I'm not really looking to add more GPUs, just the 8 channel memory to run the QFN.

Thanks!

💬 9 (+5) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/northpoler · 44h ago
Update: Anyworld, a self-hosted multiplayer text RPG, now with dockerization and zero-config Cloudflare tunneling post image

(I had to delete and re-upload this because Reddit messed up the post image somehow, sorry)

Hi, I recently posted about Anyworld, my small Python multiplayer RPG text game that runs on a browser, where an AI acts as the Dungeon Master.

Some expressed wishes that the game would be easier to set up, so I built a Docker compose system that allows you to have the game running in no time without any need to touch network settings. Just dive into DOCKER.md to get your game up and running fast, or read ahead for more details.

For local inference, it runs llama.cpp with NVIDIA GPU support, downloads a configured GGUF model from Hugging Face, and waits for the backend to be ready before starting the game. This will take a while, depending on the speed of your internet, so be patient.

The example includes recommended settings for a 16 GB VRAM system, and the model and llama-server parameters are configurable. I highly recommend Gemma 4 -based models on all VRAM tiers, they've been punching above their weights in testing.

You can also use OpenAI instead. In that mode, Compose starts the game without launching llama.cpp or downloading a local model.

To simplify networking, there’s an optional zero-config Cloudflare Quick Tunnel that prints a public HTTPS link in the console, so players can join without a Cloudflare account, domain, or router port forwarding. The address changes when the tunnel is recreated. Direct LAN access is available too, and host/player passwords still apply.

Game transcripts persist across container recreation, with optional debug logging stored separately. The Docker instructions include a Quick Startup section and commands for stopping, updating, and backing up the deployment.

The setup is working in testing, including connections from outside my LAN. I did encounter some intermittent access failures with the temporary tunnel URLs, so feedback from other networks and systems would be useful.

Hope you enjoy!

https://github.com/iamarxs/AnyWorld

💬 9 (+5) open on reddit ↗
▲
0
-1
8👁
r/LocalLLaMA · u/Exciting_Variation56 · 3d ago
For handwriting recognition does a multimodal model become overkill?

Is a small local model more than I need to change handwritten notes to text?

My agent has a skill I use but I could probably not use up my limited inference bandwidth or maybe not as much if it’s a much smaller model, right? What’s the smallest model that can accurately read handwriting?

Looking to hear others workflows or how they handle note conversion and the like

Thanks

💬 8 (+4) open on reddit ↗
▲
78
+51
20👁
r/LocalLLaMA · u/FinancialAd1961 · 2d ago
Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU post image

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.

ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs

source: https://github.com/software-mansion/runntime

💬 7 (+4) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/Chida82 · 46h ago
A 341 GB DeepSeek on a 128 GB Mac: 2x decode and first token in 0.2 s instead of 2.5 s, streaming from two SSDs, same tokens as stock ds4. The trick wasn't a kernel: I made the codebase small enough for an agent

DeepSeek V4.1 Flash at Q2 is 341 GB on disk: 152 GB of weights plus 189 GB of Engram tables. My Mac is an M5 Max with 128 GB. It runs anyway, because ds4 streams the experts from SSD, and stock ds4 gave me around 12 tok/s on the CLI. Usable. But I had the feeling the SSD wasn't the only thing holding it back, so I started poking at it with a coding agent, and that turned into something bigger than I planned.

The first thing I noticed is that every session began the same way: the agent re-reading huge chunks of ds4.c (85k lines, three model families, three GPU backends) to figure out which 2% of it my model actually goes through. Most of what it read was about hardware I don't own and models I don't run. So I deleted all of it. Not #ifdef, deleted. ds4.c is 34k lines now, the whole tree 150k instead of 278k. Only DeepSeek V4.1 Flash, only Metal.

That changed the economics of trying things. An optimization attempt that used to cost me an afternoon of the agent wandering around now costs maybe an hour, so I tried a lot more of them and measured every single one instead of picking the three I believed in.

One rule the whole time: output doesn't change. Every change has to produce the same tokens as upstream ds4 on the same GGUF (greedy, ten prompts), and pass an A/B/B/A bench against the previous build with the logits compared bit for bit. No KV quant, no approximate kernels. If it's faster but a logit moved, it doesn't go in.

This is where it landed, internal SSD only, same GGUF, same flags, upstream's own bench (ds4 → fork):

\- generation, ctx 2048 (128 tokens, first one included): 12.1 → 24.4 tok/s

\- steady decode, ctx 2048: 15.7 → 25.7 tok/s

\- steady decode, ctx 32768: 15.7 → 22.0 tok/s

\- prefill, 16k → 32k context: 404 → 636 tok/s

\- first token after a prefill: 2.1–3.0 s → 0.3–1.1 s

None of it is clever. Decode layers get committed to the GPU without waiting for each other. Expert reads are split across a thread pool and the cache slabs sit in a Metal residency set. A handful of kernel fusions per decode token. Prefill reads the next layer's experts while the current layer computes. Individually each one is a small diff you can read in a few minutes. Together they double the speed, and you get all of this with the Mac as it is, nothing to buy.

Then I got curious about the SSD part. Streaming is bound by read bandwidth and a Mac has exactly one internal drive, so I put a byte-identical copy of the GGUF on an external Thunderbolt 5 SSD and made prefill read part of every layer from each drive at the same time. The engine checks the copy against the model at every start (about 7 s) and refuses to run if anything differs, because I don't trust myself to keep two 341 GB files in sync by hand.

Internal SSD only → with the external copy:

\- 3.5K-token prompt, time to first token: 13.2 s → 11.2 s

\- 10K-token prompt: 29.0 s → 25.0 s

\- +1.5K tokens appended to a 5.3K chat: 8.6 s → 7.0 s

\- first token after an 8K context: 1.38 s → 0.18 s

Decode doesn't change, it never reads the copy. My enclosure also runs the drive at PCIe 4.0 x4, about half of the internal SSD, so a better enclosure should do better than this. Nice side effect: the KV cache can write to the external drive, so the soldered internal SSD takes zero writes while the model runs. Again: this part is optional, the table above it is the one that matters for most people.

About staying in sync with ds4, because that was my main worry: the fork never renames the ds4\_\* files, every cut is marked in the source at the exact spot, and git merge upstream/main with rerere replays the conflict resolutions. After each merge the parity check tells me if the tokens still match. So far antirez's fixes have kept flowing in without drama.

Not everything worked. I tried to bring the two-SSD trick back into upstream ds4 through its mmap path: bit-exact, but prefill got 13–38% slower and I still don't know why, so no PR for now. Twenty-odd other ideas were measured and dropped. I keep all of them in a "rejected ideas" table in the repo with the numbers, mostly so the agent (and I) stop re-proposing the same thing every other week.

There's a growing trend of single-model inference engines, and ds4 itself started that way. This is just that idea pushed a bit further, one model and one backend, and at least here it holds up: faster, still correct, still merging upstream. I've done four of these forks, one per model; the procedure is in a separate repo (StarForge) and has nothing DeepSeek- or Metal-specific in it.

Repo: github.com/Chida82/sf-ds4-1flash. The README and speed-bench/perf-record.md have the conditions behind every number.

A few things I'd like to hear opinions on:

  • Does "per-model fork that merges upstream" scale past a handful of forks, or is it just fragmentation with extra steps?
  • Is bit-exact the right bar? I left speed on the table by refusing KV quant. Would you take 10% more for a slightly different token?
  • If you have a 96–128 GB Apple Silicon Mac, I'd love to see stock ds4 and this side by side on your machine. One machine is an anecdote.
💬 11 (+4) open on reddit ↗
▲
4
-1
14👁
r/LocalLLaMA · u/DoggoProfessor959 · 2d ago
Ramjet - mini altermative to nvidia dynamo

Hello, if you are running multi gpu setup, check out ramjet https://github.com/helixml/ramjet

The goal is to have a local version of dynamo that can do equal or better job (but ideally without k8s). Check out https://helix.ml/blog/ramjet-vs-nvidia-dynamo as well.

If you have dgx spark, multi mac setups, etc it would be great to contribute recipes so other ppl can just pull

💬 8 (+3) open on reddit ↗
▲
0
-2
11👁
r/LocalLLaMA · u/r-chop14 · 2d ago
Live scribing with Jev-ish utterance gating

Like everyone, I've been following the back-and-forth regarding Jev with interest. Arguments aside about the originality of the idea, the first thing I thought of when all of this came out is "gosh, that could really help my local scribe run realtime loops during a consultation".

I pointed my harness of choice at the problem (I've been relying more and more on the LLMs as the brain rot from AI coding has continued apace). Here is the resulting workflow:

  • TEN-VAD to segment utterances and send to a Whisper compatible backend (the Tauri builds use parakeet.cpp with a 0.6B medical finetune)
  • CAM++ speaker embeddings to provide best effort diarisation (obviously limited in the setting of crappy desktop microphones and echo-y consultation rooms)
  • Here is where the Jev-ish/SemIf gating comes in. Each utterance is provided to the LLM with a short prompt and an instruction to classify as NOTE (something to be documented), ACT (action to be taken), SKIP (filler talk, etc). In the Docker deployments this is performed by the user configurable secondary model (I use Qwen3.5 4B); on the Tauri builds it's the solo primary model but into the second slot of the bundled llama.cpp server (important so that we don't clobber the prompt cache of the running main thread)
  • The initial approach was quite simple: one decode step; then gather the first token top logprobs and compute the probability mass summed over SKIP/NOTE/ACT (with some prefix matching to account for tokeniser splits and a one-word generation fallback).
  • SKIP utterances are buffered and don't get sent to the main LLM immediately (the next time the main model is woken up it will ingest that material so that nothing is lost). If a NOTE or ACT is misclassified as a SKIP, a 45s/40 word debounce runs through the main LLM with all the material it may have missed.
  • NOTE and ACT are passed on to the main model for processing. The main model has access to tools that include modification of the running note.
  • Prompt caching is essential here so that subsequent passes through the main model remain performant without a huge PP delay.

I found that the 4B model would almost never SKIP (Jev and 3.8-Flash were better but still missed 3/4 of them on natural speech). Not surprisingly (in hindsight); using the calculated probability mass alone was essentially no different to just prompting the vanilla generation endpoint and executing based on the output (roughly 81% accuracy). Looking into the logprobs a bit more it seemed that there was a usable signal in there somewhere. GLM-5.3 was pretty good figuring it out: instances where NOTE was selected, P(SKIP) ≥ 0.05, AND the utterance was ≤8 words were essentially always a SKIP. With this heuristic... 0 false SKIPs across multiple runs, and SKIP recall went from 0-50% to 75-100% on the natural consult.

The logprob gating + heuristc step is latency neutral; however, it was more reliable for this task. The otherwise vanilla small LLM like Qwen3.5-4B never flagged SKIPs and would occasionally not follow instructions entirely. I also ran an evaluation with Jev via OpenRouter (a pretty informal test set of \~40 hand-labelled utterances, and the heuristic was tuned on the same set, so it needs a held-out set to confirm); on a natural ambient consult recording the gap is smaller than I expected (both 95% accuracy but 73ms vs 514ms, keeping in mind Jev was a remote endpoint and all the latency that entails). Jev pulled away on a command heavy synthetic script (\~80% vs 100%). The overall intention was to prevent the main-loop from getting too bogged down with fluff and I think this approach achieves that.

First token logprob classification is pretty old hat; but I never really thought about one-shot classification in my scribe before Jev. And yes, the whole point of Jev is that you can just give it a classification task and have performance be good enough that you don't need to apply bespoke heuristics over logprobs to rescue your classifier (but funnily enough even Jev got an accuracy uplift from the P(SKIP) heuristic).

It was a fun experiment anyway (and grossly underpowered to say anything meaningful about Jev in general terms)! The result (video below) has been useful from my perspective (you can try it yourself here).

A synthetic consult example - performance is not this good in production environments \(overlapping speakers; bad microphones\/acoustics etc\). Primary model: Qwen3.8-Flash-Next; Secondary: Qwen3.5-4B; STT: Parakeet 0.6B \(Omi Med Finetune\)

💬 6 (+3) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/PhysicsDisastrous462 · 40h ago
Follow-up: my native Rust + Vulkan Transformer backend now qualifies on both an Intel Gen9 laptop and an AMD RDNA 3 handheld from the same build — the GPU vendor is no longer what picks the reduction shape

Follow-up to my post from a few weeks ago (14 architectures, full PEFT). This update is about a portability bug that was hiding behind its own correctness, because it's the most interesting thing I've fixed since.

The bug: the fix for one machine broke six fixtures on another

Back when I tuned the backend for Intel Gen9, I baked those kernel shapes into the portable path. That was wrong, but not for the reason you'd guess.

Two of the reductions in the saved-module path aren't really compared against "PyTorch in general" — they're compared against the PyTorch CPU library on the machine running the oracle. And ATen dispatches its vectorized CPU kernels by instruction set at run time. An AVX2 host gets 8-wide kernels; an AVX-512 host gets 16-wide ones, and the reduction shape changes with that dispatch.

So my "portable" AVX2-shaped kernels were exactly right on my AVX2-only laptop and one ulp off on my AMD ROG Ally (Ryzen Z1 Extreme, which is an AVX-512 part). One ulp doesn't sound like much until it gets amplified through every lower norm on the gradient path: the Gemma 4 saved\-stage model.embed_tokens adjoint went from 7.45e-9 to 3.22e-6, and six previously green PEFT saved-module fixtures (gemma3, gemma4, minimax\_m2, minimax\_m3, smollm3, qwen2\_5\_sliding\_tied) crossed the 2e-7 gate. Neither shape is wrong — only one matches a given machine, and baking in either one breaks the other.

The fix: probe the host, not the vendor

Kernel variants are still selected by GPU vendor. Those two reductions are now selected by host CPU capability instead: capability is probed once per process and cached, then the matching module pair is dispatched (linear_forward_lane2 / linear_forward_lane4, and the 8-lane / 16-lane transformer_cross_entropy builds). HIERARCHOS_ATEN_VECTOR_WIDTH=8|16 pins the shape for qualification when a reference wheel's kernels disagree with the CPU's own capability.

|Host|GPU|CPU dispatch|Status|
|:-|:-|:-|:-|
|Intel i5-6200U / HD Graphics 520 (2016 Skylake-U)|Intel Gen9|AVX2 only, no avx512f|32/32 LoRA, 32/32 switching, 32/32 saved|
|AMD Ryzen Z1 Extreme|RDNA 3|AVX-512|32/32 LoRA, 32/32 switching, 32/32 saved|

Same 2e-7 gate, unchanged. No tolerance was loosened to get there.

What I verified on each side

On the Intel machine, the post-change matrix is bit-identical, field for field, to its pre-change report across all 32 families — peft, gradient, two-step AdamW, frozen base, resume, lifecycle — which is how I know the AMD fix didn't quietly cost the Gen9 path anything. Also 693 passed / 0 failed / 9 ignored on the Rust lib suite and a clean strict headline forward run.

On the AMD side, the fix was re-qualified end to end: 32/32 on all three stages, provenance clean.

The harness fingerprints the pinned Transformers source alongside the shaders and binaries, and on the Intel side I re-derived the whole fingerprint from the pushed tree myself: 3951 inputs, zero changed, zero missing. So "green" refers to one frozen set of reference math, not whatever happened to be on disk.

Same caveats as always

  • This is deterministic FP32 tiny-model correctness against a reference implementation, not a claim about arbitrary checkpoint sizes, dtypes, or hyperparameters.
  • "Supported text graph" ≠ "the whole multimodal package works natively."
  • The AVX-512 dispatch is only qualified on the AMD machine, since it's the only host I have that can execute it natively. The 16-lane module also doesn't rebuild byte-identically with the glslang version on my Intel box (one extra type/id, one difference in +inf materialization), so I've left it as the committed AMD-built module and documented that rather than swapping it without re-qualifying both hosts. I'd rather report that than pretend it's clean.
  • NVIDIA and other GPUs are genuinely unqualified — the path is raw Vulkan, so they're untested rather than excluded.

What I'd love from you

Last time several people asked about hardware other than mine, so that's the ask again: if you build it on an AVX-512 laptop, an AVX2-only machine, or an NVIDIA/Intel GPU, I want to know what you get. The two reductions above are the ones most likely to behave differently on your CPU, and knowing your host's vector width is now part of the answer.

The new cross-platform section in the README documents the whole thing, including which host classes are measured and which aren't.

Repo: https://github.com/necat101/Hierarchos-Native Compatibility/parity record: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/COMPATIBILITY.md Regression audit: https://github.com/necat101/Hierarchos-Native/blob/main/AMD\_REGRESSION\_AUDIT.md Per-host tuning and measurements: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/VENDOR\_TUNING.md

💬 3 (+3) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/Least_Dog_8556 · 3d ago
[Model Release] Qwen3.8-cyber-RedTeam-27B (Surgical Abliterated) — Unconstrained Foundation Engine for Red-Team & Low-Level Security post image

Hey everyone,

I'm releasing \*\*Qwen3.8-cyber-RedTeam-Surgical-Abliterated (27B)\*\*, an unconstrained foundation engine fine-tuned specifically for cybersecurity engineers, authorized red-

team operations, and memory exploitation research.

Tired of frontier models refusing to dissect vulnerable kernel dispatch routines or rejecting benign fuzzing/audit payloads with moralizing lectures? This model addresses

that directly.

\### Key Highlights:

\* \*\*Architecture\*\*: 27B Qwen 3.5 Hybrid SSM (48 Linear-Attention layers + 16 Full-Attention layers) with 75% active KV-cache reduction (runs full 256K contexts on single GPUs

without OOM).

\* \*\*Context Window\*\*: Native 256K context support (RoPE $\\theta = 10\^7$).

\* \*\*Surgical Abliteration\*\*: Refusal direction centroids were mathematically removed via residual stream orthogonalization—zero preachy refusals while rigorously preserving

deterministic C/assembly syntax and reasoning.

\* \*\*Precision\*\*: Sharded native FP8 (F8\_E4M3, block size 128x128) fitting on single 32GB/48GB/80GB GPUs.

\* \*\*Agentic Ready\*\*: Native multi-step tool-calling support, zero-overhead RadixAttention prefix caching via SGLang.

\### 1-Command Quickstart:

git clone https://huggingface.co/medismera/Qwen3.8-cyber-RedTeam-Surgical-Abliterated

cd Qwen3.8-cyber-RedTeam-Surgical-Abliterated

bash deploy.sh

\*\*Model Card & Weights\*\*: \medismera/Qwen3.8-cyber-RedTeam-Surgical-Abliterated\

Feedback and bug reports from the community are warmly welcome!

💬 12 (+2) open on reddit ↗
▲
0
-1
5👁
r/LocalLLaMA · u/Civil_Fee_7862 · 3d ago
Help deciding on best harness for developing a custom agent orchastrator?

Been developing a custom agent orchestration layer with opencode as the harnes. It has been successful so far in terms on being manage multiple concurrent sessions. However, I am wondering if I am using the best harness given that opencode isn't meant to be a hackable, i.e. Less flexible compared to something like pi.

I've developed an interface that allows me to easily switch between different harnesses. i.e. Without having to re-write the orchestration layer, and am considering swapping out opencode for pi. The reasoning is obvious, pi is meant to be a hackable harness, so its likely to be a better fit for building a custom multi-agent system. Opencode does support a headless mode which has helped a lot. But it seems heavy on resources, and recently has seemed very buggy. Less features might be a better approach towards the stability that I need.

Before I make the dive, has anyone else already tried integrating pi into a multi-agent system? Did you find pi a better fit for the worker layer compared to something like opencode? What problems did you run into?

Thanks for your help.

💬 6 (+2) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/Rough_Practice7631 · 3d ago
I tested Gemma 3 27B and Qwen3 32B against two frontier models on financial analysis tasks. The small models answers are much less reliable

I ran a series of simple experiments to see how LLMs behave when asked to judge companies from real financial data. I used 4 models: Gemma 3 27B and Qwen3 32B as small models, and Opus 5 and GPT-5.6 Sol as frontier models. All calls were made through the Bedrock API.

The samples are small, so the numbers below should be read as a demonstration and a methodology and not a definite proof.

Here are some of my observations, particularly when it comes to the differences between small and large models.

\*\*Rank vs. score.\*\* I asked each model to rank six companies from best to worst, and separately to score each one from 0 to 100. The underlying judgment is the same, so we would expect the same ordering. Rank and score gave an identical ordering in 75% of sets for Opus and 65% for GPT, but only 25% for Gemma and 15% for Qwen.

\*\*Order of the list.\*\* For Qwen, simply reversing the order in which the companies were listed changed the top-ranked company in 6 of 10 sets.

\*\*Analyst opinions.\*\* Attaching a bearish analyst note to the data lowered the rating in 81% of cases for Qwen and 71% for Gemma, against 43% for GPT. Opus mostly kept its own view. Interestingly, the fix is simple: asking the model to identify the opinion and reason independently brought the rating back toward its original level in 85% of cases for Gemma and 68% for Qwen.

\*\*Summarize, then analyze.\*\* When rating a summary of a 10-Q section instead of the full text, the frontier models gave the same rating in 84% of cases. Gemma and Qwen changed their rating in 25% and 31% of cases.

I wrote something more complete, with charts and the details of each experiment: https://sabrresearch.com/cookbooks/llm-financial-bias

Disclosure: this is my own work, published on my company's website. Happy to answer questions on the setup.

\------ Edit 1 -------

People complained about the choice, of models. To be clear, I'm not trying to trash small models, on the contrary, what is of interest to me here is the overall trend and inconsistencies which occur in both frontier and small. I also ran the analysis on Gemma 4 31B (April 2026), it's a bit better than Gemma 3 used above, but still the same issues:

\*\*Rank vs. score.\*\* Rank and score gave an identical ordering in 35% of sets for Gemma 4, against 25% for Gemma 3. Still far from Opus (75%) and GPT (65%).

\*\*Order of the list.\*\* Reversing the list changed Gemma 4's top-ranked company in 2 of 10 sets, the same as Gemma 3.

\*\*Analyst opinions.\*\* A bearish analyst note lowered Gemma 4's rating in 53% of cases, vs 71% for Gemma 3 but still above GPT (43%) and well above Opus (19%).

\*\*Summarize, then analyze.\*\* Gemma 4 changed its rating in 25% of cases when given a summary instead of the full text, the same as Gemma 3.

💬 23 (+1) open on reddit ↗
▲
32
+22
16👁
r/LocalLLaMA · u/FinancialAd1961 · 46h ago
omni-d1 600M by Liquid AI running in the browser using WebGPU post image

Liquid just dropped d1-omni-600M and it's a decision model!

I ported it to runntime, the WebGPU inference library I'm currently working on. It's plain TypeScript on top of TypeGPU with no WASM and virtually no export step. The model is written directly from our core ops (matmul, attention, norms, a few elementwise bits), and the weights load straight from the HF safetensors.

The demo in the video is a fake comment feed being moderated live. Each comment gets 4 questions: toxic? spam? asking something? overall tone? Toxic and spam ones get removed.

\~180 ms per comment for all 4 questions, \~45 ms per question - the performance will most likely be way better once we spend some time tuning the engine for this model

I also tried making it play snake, but It did not go well lol.

The d1 port isn't in the npm release yet. The rest of runntime is (detection, segmentation, speech-to-text, embeddings and more). Docs and live demos: https://docs.swmansion.com/runntime

Happy to answer questions about the port or WebGPU stuff in general.

💬 7 (+1) open on reddit ↗
▲
2
+1
13👁
r/LocalLLaMA · u/one_does_not_just · 47h ago
Porting LIBERO to MuJoCo Warp: 130 robot manipulation tasks on one $700 AMD GPU

LIBERO is a robot manipulation benchmark: 130 tasks across five suites (spatial, object, goal, scene10, scene90), each with human demos and a language goal. It is the standard testbed for language-conditioned imitation, and it normally runs on robosuite with CPU MuJoCo, one environment at a time.

I ported all of it to MuJoCo Warp and ran it on an RX 9070 XT ($700, 16 GB, RDNA4). Physics does 17,137 env-steps/s at 2,048 worlds; CPU robosuite does 38. A 50-epoch BC transformer gets 42.5% on Warp vs 50% on CPU.

Why the GPU matters: behavioral cloning needs rollouts. Run the policy, watch where it fails, and generate labels or demos from that. On CPU it is one env at a time, and generating demos for a single suite (10 tasks) took me 8 to 9 hours. On the GPU it is minutes, so you can iterate instead of running it once.

Why Warp on AMD is the interesting part: Warp is NVIDIA's GPU sim framework, and the AMD HIP/ROCm port is recent (Tomas Thoresen, Strix Halo). Getting it working on RDNA4, with Warp compiling HIP kernels for gfx1201, JAX seeing rocm:0, and PyTorch seeing cuda, is what made this possible.

The renderer was the annoying bit. A BC policy is a fixed function of the pixels, and Warp's ray tracer is not MuJoCo's CPU renderer. My first Warp eval scored 0%. Four fixes got it to 42.5%: vertical flip, shadow constant 0.3 -> 0.0, 1.15x brightness, and cube-map sampling (the table wood grain rendered flat). Brightness alone was worth 12.5 points.

Everything builds from public sources on ROCm, and the example plus a 21.6 MB BC checkpoint are in the repo. What I didn't finish: the Warp path collects rollouts, training is still offline in PyTorch. I worked on the R9700 and some Instinct cards at AMD over the summer, so in-loop vision is next.

Writeup: https://amohan.dev/blog/2026/libero-warp-mjx-rdna4/
Code: https://github.com/poad42/libero_mjx

▲
3
+2
4👁
r/LocalLLaMA · u/Yossarian_1234 · 43h ago
[R] Split the Differences, Pool the Rest: Provably Efficient Multi-Objective Imitation

TLDR: The question we answer: how do you learn from experts with different objectives? Pooling all their data can lose their trade-offs; learning from each expert separately misses opportunities to share data. MA-BC pools demonstrations where observed actions don’t disagree, with upper and lower bounds on sample complexity.
Authors: Ziyad Sheebaelhamd, Luca Viano, Volkan Cevher, Claire Vernade

Arxiv: https://arxiv.org/abs/2605.12000
Github: https://github.com/ziyadsheeba/mabc

https://preview.redd.it/i20adc3z04uh1.png?width=2532&format=png&auto=…

▲
3
+2
3👁
r/LocalLLaMA · u/jjusko20 · 43h ago
My progress on a [new] model-architecture specific dynamic quant technique - v1

Hey everyone,

I have new in brackets above because I'm not necessarily inventing anything innovative in terms of the actual mathematics or optimizations behind some quant techniques, but I'm pretty happy with how things are coming.

What I'm working with is basically a "poor man's" RCO (the quant method from IST Austria dAS lAB). Exact same concept: choose a type per tensor under a byte budget while optimizing task KL on the whole model. I worked on GSQ but I don't have an approximation method that beats baseline - yet.

Take the core principles of the method, make them cheaper approximations, and regain as much accuracy as possible. I originally planned to make an approximate GSQ-RCO hybrid, but none of my hypothetical models for the approximate for GSQ have beaten baseline yet.

In application: start with a full precision model, and create an "imatrix shape" map, per tensor. This doesn't calculate the sensitivities of individual tensors - but it creates a sensitivity "curve" where you can approximate which tensors in a model suffer most from quantization via extrapolation. This creates a baseline estimate of the optimal quant per tensor.

Then: iterative trial and error with local search. Take the file size of the baseline estimate, and substitute different precision per class to bring overall model size down, beginning with the tensors the approximation model marked as most sensitive to quantization. Once the working-best version hits under the filesize cap, it tries variations of substitutions that keep the file size approximately the same (upgrading certain tensors, downgrading certain ones, etc - basically looking for holes in the local search method once the local search is done).

For the whole process above (the iterative local search + search error recovery \[not really error but i cant find the word im looking for\]), the model is quantized and KL divergence is measured vs prior iterations - anything that raises KL divergence is discarded. The result: an approximated RCO style quant, iterated as closely as possible to optimal.

Do I expect this to beat GSQ-RCO or Unsloth dynamic V3? Definitely not GSQ-RCO or regular RCO, and likely not the unsloth ones. However, I've got a few advantages: this is CHEAP and extremely conservative on VRAM usage. The teacher model only needs to be loaded once: to dump its per token log probs. This is quick on a GPU, but since it only has to be done once, it can be done on CPU with a little patience. Every step after only pulls the candidates onto GPU, starting from the imatrix curve approximation - so all you need is enough VRAM for your approximate final quant size (with a little buffer for iteration, maybe 20-25% more would be optimal). The whole process takes a few minutes to a few hours depending on what you're doing.

I've pretty much documented a psuedo-algorithm approach above that's recreatable, but I can supply better documentation if people are interested.

Some early results on Qwen 3.5 2B:

llama.cpp IQ3\_M with an imatrix - 999mb vs RCO-lite with an imatrix - 1088mb \[+89mb\]

Mean KLD for IQ3\_M: 0.098381

Mean KLD for RCO-lite dynamic mixture \[+89mb\]: 0.045162 -- almost exactly half for 89 more mb

Mean KLD for RCO-lite dynamic mixture \[cap 1038, + 39mb\]: 0.0793220

Mean KLD for RCO-lite dynamic mixture \[cap 947, - 52 mb\]: 0.090624 -- still better than the IQ3\_M quant despite being 52mb less.

Take these as early results - I forgot to document exact +- for my KLD runs, but the band was generally lower than the i quants. I need to try various different size targets to figure out what BPW range this algorithm works in most effectively, and this is just up against the IQ3\_M - IQ4\_XS had a better KLD than this design - more BPW so it's not an exact estimate, but not that substantially - I didn't try to fit an optimal model inside the IQ4\_XS size range yet, that was just something I noticed. I suspect the Q3 and Q2 ranges will benefit most from this, I haven't tried in the higher BPW ranges yet - partway through Q2 experiments.

Also: my baseline llama.cpp quants are calibrated on the same wikitext set for the imatrix as the RCO-lite quants are, with the same held out set for the KL divergence.

Cheers.