115 posts · 1 sub · RSS
← prev Oct 1, 2026 → Oct 2, 2026 next →
2026-10-01 → 2026-10-02 hourdayweekmonthyearall
allr/LocalLLaMA
▲
2324
+1170
113👁
r/LocalLLaMA · u/StayLameBro · 7d ago
I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. post image

\*\*DISCLAIMER\*\* THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE.

Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context is all that fits next to Qwen 3.8 27B (IQ4\_XS), even with the wired limit raised to 20480. An iPhone 17 Pro Max was sitting in my pocket, so I figured what can I do to make use of this extra silicon.

Turns out a 10 Gb/s USB-C cable & some software is all you need. The Mac runs layers 1–40 of each 256-token batch and streams the activations to the phone. The phone runs layers 41–64 on its GPU while the Mac starts the next batch. The A19 Pro's GPU has matrix units (Metal 4 tensor ops), and they make the phone's half 2.4x faster than the same phone without them.

Same build, phone off vs. on, prefilling a 2,000-token file into a saved agent session:

  • 8k context: Mac alone 132 tok/s → Mac + iPhone 177 tok/s (+35%) (measured two days earlier, same bench)
  • 16k context: Mac alone 109 tok/s → Mac + iPhone 157 tok/s (+44%)
  • 32k context: Mac alone 101 tok/s → Mac + iPhone 130 tok/s (+29%)
  • 48k context: Mac alone 87 tok/s → Mac + iPhone 113 tok/s (+30%)

A fresh 27k-token agent session, cold: 245 s on stock llama.cpp, 228 s on my fork with the Mac alone, and 168 s with the phone.

Past 64k the phone switches jobs. The oldest KV pages move to the phone and the Mac runs all 64 layers. For every attention layer, the phone computes attention over the old keys on its GPU, and the Mac merges that with its own part. While writing, the phone's Neural Engine takes part of that work too: each 16k-key page of old context is compiled into a Neural Engine model with the keys as its weights. At 140k that took writing from 279 to 176 ms per token compared with the phone's GPU alone.

The server allocates 196k–229k of 8-bit context based on the phone's free memory; that's up to \~5.7 GB of KV cache living on the phone instead of the Mac, so the Mac's memory use stops growing at 64k. I've tested a growing session to 128k at 8-bit, with 3/3 planted facts recalled. Separately, at 140k in 4-bit, the run passed the gate with greedy output matching the Mac-only run for 32 generated tokens.

What it doesn't do: speed up writing below 64k. That's the Mac's job. My fork's kernels (SME2 on the M4 CPU and Metal fusions) plus DFlash2 speculative decoding take it from 11.3 tok/s on stock llama.cpp to 25 tok/s at about 30k context with medium thinking, phone or not. SME2 also adds up to 29% to prefill on the Mac alone. Past 64k the phone does share the writing (attention over the old keys), and without it the Mac would have to drop to 4-bit context to reach 128k. In real use I have seen upwards of 30 TPS at lower context.

The phone joins prefills over about 512 tokens. In one real session, that was 7 of 36 requests, but about 83% of the tokens read. Past 64k it holds the context and does the old-key attention, but it stops running layers 41–64 there for now; doing both is next. One request at a time.

I'm curious what this setup could do with newer model architectures. DeepSeek V4.1-Flash reports 890 bytes per token for its global KV cache and adds n-gram embedding tables (Engram). Qwen3.8-Flash-Next, the Qwen 4 architecture preview, has an n-gram lookup table too. Those aren't features of the 27B model I tested, and I haven't benchmarked either architecture here. The real gold is within the newer phones and models working together. With the A20 Pro in the iPhone 18 Pro Max, I bet there is a lot more for me to push.

Code, setup and bench scripts: https://github.com/StayLameBro/backburner

Still a lot of work to do but I built this with Opus 5.5. Happy to answer anything.

💬 328 (+102) open on reddit ↗
▲
943
+297
94👁
r/LocalLLaMA · u/kvyb · 7d ago
Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following

Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since.

I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly:

incapable of producing more than a few words at a time.
single default personality which no amount of prompting can overcome
will not use tools, at all, whatsoever.
There needs to be a middle ground

They were right. The tool calls didn't actually work, and when people asked it to do something it would sometimes just say it's busy or going to bed. Very human. In a bad way.

So I spent the last three weeks on 2.0. The goal was simple: keep the voice people liked and lose the drawbacks.

What 2.0 does now

  • With no system prompt, it's a normal person texting. Not an assistant, not a catgirl.
  • Give it a character card and it becomes that person, and still texts like one.
  • Ask for a formal email, numbered steps or a proper explanation, and you get exactly that. Then it goes back to texting.
  • Don't want the lowercase texting? Tell it "from now on write in full sentences" (or put it in the system prompt) and it sticks to that until you say otherwise. v1 ignored this completely.
  • It calls tools, and it asks when something is missing instead of making it up. This is the part I'm happiest about. Ask the base model to book a flight without saying where from and it picks JFK. 2.0 asks where you're flying from.
  • It writes code and does math at roughly base-model level.

It's a colleague and a humanlike companion, not an assistant. Use it for chat, roleplay, agents or actual work.

How I trained it

v1 was plain SFT on real and synthetic conversations (139,845 messages from 1,396 conversations). That copies habits, including the bad ones.

For 2.0 I used on-policy distillation. The model writes its own replies and a teacher grades every token. There are two teachers:

  • v1 plus a hidden "text like a person" instruction, for chat and characters;
  • the plain base model, for instructions, tools and code.

The student never sees the hidden instruction, so it learns the behaviour without needing a prompt. Same 27B, a second LoRA on top, merged.

Numbers (vs the model I trained on, huihui-ai's abliterated Qwen3.8-27B; same prompts, same run, thinking off)

|Benchmark|Base (abliterated)|2.0|
|:-|:-|:-|
|IFBench (instruction types I never trained on)|37.3|43.7|
|When2Call (call, ask or refuse correctly)|48|58|
|BFCL irrelevance (don't call a tool when none fits)|60|78|
|IFEval, GSM8K, BFCL simple|81.9 / 89.1 / 97|83.5 / 89.1 / 98 (ties)|

Full chart in the images.

Where it's still worse: knowledge (MMLU-Pro 72.5 vs 78.5) and competitive code (LiveCodeBench 51 vs 56).

Is it actually more human? I built a benchmark for this, "ishuman":

  • It takes 150 fragments from unseen chats.
  • Has each model write the next message.
  • Shows a judge the real message and the model's without labels, and asks which one a person wrote.

|Model|Judge thought it was the real person (50% = can't tell)|
|:-|:-|
|Qwen3.8-27B abliterated (huihui-ai, the model I trained on)|0.3%|
|Same abliterated model + a "text like a human" system prompt|6.8%|
|Qwen3.8-27B official (unmodified, via OpenRouter)|15.1%|
|Qwen3.8-27B-Humanlike-Chat 2.0|23.5%|

So no, you can't just prompt your way there. In a separate test of 16 live multi-turn chats with invented people, 2.0 was picked over the base model 16 out of 16 times.

Links

Big thanks to everyone who left feedback last time, especially the ones who were critical. Tell me where it still sounds like an assistant.

Edit: safetensors are up for vLLM and SGLang:
GPTQ-Int4 (24 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-GPTQ-Int4
FP8 (48 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
BF16 (80 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0

💬 191 (+29) open on reddit ↗
▲
725
+223
109👁
r/LocalLLaMA · u/LegacyRemaster · 7d ago
Does anyone know if any new releases from Mistral are planned? post image

It’s been a long time since the last models came out. I notice they are selling GLM on the site, and I wonder if they are developing something, given the long silence.

💬 333 (+50) open on reddit ↗
▲
581
+186
63👁
▲
506
+118
79👁
r/LocalLLaMA · u/paf1138 · 7d ago
New in llama.cpp: Decision Models
💬 123 (+14) open on reddit ↗
▲
271
+75
43👁
r/LocalLLaMA · u/Recoil42 · 7d ago
New Architecture from Percepta: Spotlight — separating intelligence from memory, allowing knowledge and skills to grow without changing the model's weights. post image

https://www.percepta.ai/blog/can-llms-grow-their-own-capabilities https://www.percepta.ai/blog/spotlight-memory "Our new architecture, Spotlight, replaces attention with a memory that escapes this trade-off: it is the first architecture to achieve infinitely growing memory without increasing the access cost. Every token reads from and writes to an unbounded memory, but because the model learns to index individual memory cells, each token only touches a small number at a time. While other sparse architectures fix the fraction of capacity used at each step—a mixture-of-experts model, for instance, always activates the same number of experts out of a fixed set—Spotlight is arbitrarily sparse, touching the same number of cells regardless of how the memory grows. The fraction of memory it uses can shrink as far as we want. Spotlight separates an intelligence module, which performs computation, from memory, which holds knowledge, procedures, and working state. The intelligence module stays the same size, and the weights don't change as memory grows. The memory is writable, and the model itself decides what to load and when to overwrite it, token by token. Because memory can hold skills as well as facts, the model can gain new capabilities without retraining: what it can do is not limited by the size of its intelligence module."

💬 32 (+4) open on reddit ↗
▲
458
+70
56👁
▲
220
+56
55👁
r/LocalLLaMA · u/jacek2023 · 7d ago
microsoft/FrogNano-4B-2609 · Hugging Face

An agentic model from Microsoft for the GPU poor

https://huggingface.co/bartowski/FrogNano-4B-2609-GGUF

FrogNano is derived from Qwen/Qwen3.5-4B, a general-purpose post-trained model designed for language, reasoning, coding, agentic, and multimodal tasks. FrogNano inherits Qwen3.5-4B's dense 32-layer hybrid Gated DeltaNet and gated-attention architecture, but its additional post-training is text-only and focused on repository-level software engineering. The model is further trained using reinforcement learning on approximately 1,500 synthetic SWE task environments generated and calibrated against the evolving policy using TaskPilot. Training uses the five-tool Leaf harness and executable test-based rewards over complete multi-turn coding trajectories.

The additional post-training is intended to improve long-horizon repository navigation, debugging, code editing, test execution, and patch generation in a compact 4B model. Unlike approaches based on behavioral distillation, FrogNano does not train on stronger-model solution trajectories, actions, reasoning traces, or patch targets. This specialization also introduces limitations and risks: performance is sensitive to the Leaf harness and test quality, training data are Python-heavy and primarily English, and generated patches may be incorrect or insecure despite passing available tests. When integrated with the Leaf harness, FrogNano generates structured tool calls that can propose repository changes. Leaf executes authorized tool calls within an isolated repository environment to produce a candidate patch; FrogNano does not itself deploy the changes. Any resulting patches require human review, regression testing, and security validation before use or deployment.

💬 69 (+14) open on reddit ↗
▲
108
+55
73👁
r/LocalLLaMA · u/basnijholt · 7d ago
Self-hosting AI does not save money, and I do it anyway

Hi folks, I'm a long-time lurker and big fan of this subreddit and a massive self-hosting fan (also outside of AI).

I doubt many people will disagree with me here because I see the same arguments being made in many posts. However, I thought it might be interesting to share anyway. I wrote down why self-hosting AI does not save money: https://www.nijho.lt/post/self-hosting-ai-is-not-cheaper/

EDIT: didn't think this would be so controversial 😅 I do say explicitly in my blog post "I would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not".

EDIT 2: Comparing $200 sub with Opus 5.5 or Astra with Qwen 3.8 27B is not apples to apples.

💬 194 (+42) open on reddit ↗
▲
343
+47
59👁
r/LocalLLaMA · u/-p-e-w- · 8d ago
Heretic is on PewDiePie!

So I haven’t played a computer game in 20 years, and I know nothing about Minecraft, and I definitely prefer classical literature over YouTube culture, but even I have heard about the individual called PewDiePie, for two reasons:

  1. His monicker starts with the initials of my own name
  1. I remember a recurring Internet meme a few years ago where he was competing for the most subscribers with an Indian film music channel

I had never watched a single one of his videos, however.

Well, until today, when people started spamming me with messages informing me that Mr. Kjellberg aka PewDiePie has tried out Heretic and made a video where he talks about it:

https://m.youtube.com/watch?v=ODDJXGY_1kQ

(Heretic mentioned around 9:00)

Obviously I’m thrilled that a less technical audience is being exposed to my work, and the more people understand what is possible the better. I expect to be receiving a couple hundred more mails in the coming days asking how to run Heretic on ChatGPT (you can’t), or accusing me of working for the CIA (I don’t), but other than that, the more the merrier I guess 😏

Heretic 2.0 coming soon…

💬 90 (+3) open on reddit ↗
▲
201
+39
65👁
▲
410
+28
45👁
▲
138
+26
47👁
r/LocalLLaMA · u/matteiuspi · 7d ago
Two 96 GB Ascend cards crun Qwen3.8-flash-next hardware notes, vLLM work, benchmarks, and what is next

I have been building a somewhat unusual local inference machine around two Huawei Atlas 300I Duo cards. They are relatively inexpensive, passive, dual-accelerator PCIe cards with 96 GB of device memory apiece. They are also absolutely not drop-in CUDA replacements.

When I first brought up Qwen3.8 Flash-Next these past two weeks, it was often incoherent and lived around 1 generated token per second. Some runs were below that. Today the same two-card machine is producing coherent output at roughly 30 tok/s for one request and about 61 tok/s aggregate at four-way concurrency on my short decode benchmark. It also completed the full 198-question GPQA Diamond set.

This post is the start of a guide for these cards: what the cards physically are, how I cool them, what “96 GB” really means, what I changed in vLLM and vLLM Ascend, which optimizations actually mattered, and which problems are still open.

The short version is the hardware is capable. My work has been mostly on the software stack, ubuntu-26.04 driver support, model architecture support, memory layout, custom operators, and getting every asynchronous state transition exactly right.

The hardware: one card is really two devices

My current machine has two Atlas 300I Duo cards, which enumerate as four Ascend 310P3 devices.

Current system / Planned system

Physical cards: 2 / 3

Ascend devices/npus/AIcpus: 4 / 6

Nameplate device memory: 192 GB / 288 GB

Approx. runtime-visible memory with this configuration: 172 GiB / 258 GiB

Combined maximum accelerator-board power: 300 W / 450 W

Each card has two accelerator SoCs and 96 GB of LPDDR4X in total, or 48 GB local to each chip. It is not one unified 96 GB allocation. A model that does not fit on one 48 GB device still needs tensor, expert, pipeline, or another form of model parallelism. The card is PCIe Gen4 x16, full-height/full-length, and a surprisingly thin single-slot design. Huawei rates it at 408 GB/s aggregate memory bandwidth and 150 W maximum board power. The official specifications are here (https://support.huawei.com/enterprise/en/doc/EDOC1100285916?section=j00e).

Also, despite the generic “HBM” terminology used by a lot of accelerator software, the memory on these cards is LPDDR4X.

This two-chip-per-card layout matters. Communication within a model still goes through the distributed runtime, and memory remains local to a rank. I use HCCL collectives and explicitly map tensor and expert ownership across all four chips. Thinking of the machine as four 48 GB ranks is much more useful than thinking of it as two 96 GB GPUs.

https://reddit.com/link/1wvt1m4/video/7rju7qbrw1th1/player

Passive cooling is not a deal-breaker

The cards have large heatsinks and no onboard fans. They were designed for server airflow, so putting them in an ordinary workstation and hoping a rear case fan will sort it out is a bad plan. There is a useful teardown here (https://videocardz.com/newz/huawei-atlas-300i-dual-ai-gpu-with-96gb-memory-worth-1400-has-been-taken-apart) if you want to see the heatsink and heat-pipe arrangement.

I give them direct, high-volume airflow and run the room on AC/heat-pump cooling. Under real multi-hour model loads, the cards can crank continuously without drama. Across my recorded Qwen runs, peak device temperatures were generally 72–78 °C. My watchdog limit is 96 °C, and the cards have not approached it.

So I would not bat an eye at adding another passive card. The actual checklist is mundane:

• Keep unobstructed airflow through the heatsink fins.

• Make sure the chassis fans have enough static pressure.

• Budget another 150 W of board power per card, plus the rest of the host.

• Exhaust the heat from the room instead of recirculating it through the rack.

• Log temperature during long prefill, decode, and concurrency tests rather than trusting an idle reading.

Passive does not mean low-power or self-cooling. It means the chassis and room are the cooling system. Once that is handled, these have behaved like ordinary 150 W server cards for me.

Note: Nothing heats up these cards more than loading/moving things around in their ram -- the npus at full utilization run cooler than large block memory assignments. We keep this in mind when optimizing the model serving code paths.

ECC, nameplate memory, and what is actually usable

My cards arrived with ECC enabled by default. I disabled it to reclaim device memory. This is an inference and development box, and I consciously prefer capacity over ECC protection here. That is a reliability tradeoff, not a universal recommendation.

Even with ECC disabled, firmware, the runtime, communication buffers, graph captures, workspaces, and allocator reservations consume memory. In practice, the software sees roughly 43 GiB per 48 GB chip. Four chips therefore provide about 172 GiB of useful aggregate capacity, but it is still four separate local pools. The exact free number also changes with the CANN build and launch configuration.

That distinction has shaped nearly every model decision. The question is not only, “Does the checkpoint total fit in 192 GB?” It is, “Does each rank's weight shard, recurrent state, KV/cache allocation, graph capture, collective workspace, and worst-case temporary allocation fit in its own 43 GiB?”

Why I forked vLLM as well as vLLM Ascend

The public work lives in the OpenSensor vLLM Ascend fork (https://github.com/opensensor/vllm-ascend) and paired vLLM fork (https://github.com/opensensor/vllm). I needed both sides because this was not just a missing device kernel.

I am currently the only person developing these forks. The software bus factor today is one. I have made a lot of progress, but a fast-moving one-person fork should not be confused with the maturity, test coverage, or support depth of mainline vLLM on NVIDIA.

I also ran into a bizarre tooling problem: in my sessions, Claude repeatedly refused to engage with prompts about this architecture because the cards are Huawei hardware from China. These were ordinary engineering discussions about serving, sharding, cooling, and performance—not requests to build a restricted application. I am describing my direct experience rather than claiming that every Claude version or account will behave identically, but it made Claude unreliable as a development assistant for this project. Whatever anyone thinks about the politics, that is a real practical constraint when choosing tools around this hardware.

Qwen3.8 Flash-Next combines MoE routing, Gated DeltaNet recurrent layers, sparse quadratic-attention layers, packed low-bit experts, long context, and an MTP draft model. Supporting that cleanly touched model integration, the v1 runner, cache accounting, scheduling, graph capture, distributed state, model loading, and Ascend-specific operators.

The current development sprint has been roughly two and a half weeks of nearly continuous bring-up and optimization, with hundreds of fork commits, repeated full checkpoint loads, profiler captures, operator microbenchmarks, and multi-hour quality runs. This was not one magic kernel patch.

What took Qwen from incoherent \~1 tok/s to where it is now

These are the architectural changes that moved the needle.

  1. Make the hybrid model correct before making it fast

The early model could generate tokens, but generation was not a correctness test. I found failures that only appeared at production geometry: incorrect Gated DeltaNet gate-vector handling, recurrent-state precision and lifecycle problems, incomplete sparse-attention score width, and mismatches between the host operator API and the installed kernel package.

One particularly nasty GDN issue looked fine in small-head tests but accumulated state error across the real 36 recurrent layers and produced incoherent text. Keeping recurrent state in FP32 and fixing the production-shaped data movement was foundational. So was treating the custom OPP package and Python host code as one ABI-versioned unit. A stale kernel can look like a model problem for a long time.

  1. Shard the model at load time instead of loading everything everywhere

The Qwen checkpoint is about 169 GiB and contains 1,610 safetensor files. It was larger than available host RAM in one of my bring-up configurations, never mind the memory on an individual NPU.

I built an expert-aware loader that reads only the experts owned by each rank, keeps dense/shared tensors where required, and avoids materializing the whole expert bank before throwing most of it away. Host-side expert data is mapped and moved lazily. This changed model loading from an accidental memory stress test into a deterministic TP4/EP4 layout.

  1. Keep low-bit weights packed and do the work on the NPU

My first usable W4 reference dequantized packed weights in the eager path and then called a regular matmul. It was useful for correctness and managed only 0.229 tok/s on one recorded smoke case.

The production path keeps the weights packed, routes tokens to experts on the device, and uses custom AscendC Cube kernels for grouped expert projections. I added FRACTAL\_NZ layouts, fused gate/up handling, tiled reductions, route and tile reuse, and dedicated W4A8 execution instead of repeatedly expanding W4 weights into a larger temporary representation.

This is both a speed win and a capacity win. Avoiding transient expanded expert banks leaves memory available for state, cache, graphs, and concurrency.

  1. Build caches for the model I actually have

Flash-Next is not a conventional all-attention transformer. My configuration has 36 recurrent GDN layers and 12 sparse QSA layers. Treating all of that as a normal dense KV cache wastes memory and misses the state semantics.

I implemented separate recurrent-state management, compact physical cache layouts, prefix-state tiers, sparse page selection, direct NZ gathers, and 310P-specific QSA paths. The service is configured for a 262,144-token context limit, and the cache planner retains capacity for about 4.08 such windows. That is a memory-planning result, not a claim that every possible four-by-262K workload has completed an end-to-end soak.

  1. Remove synchronization and launch overhead from the token loop

On these devices, a stray device-to-host scalar read can serialize the whole pipeline. I removed hot-path .item() calls, reused per-step tensors, deferred collectives behind useful work, tightened CPU affinity, and moved routing and sparse selection away from Python.

Once the eager path was correct, I added decode-only ACL graphs and MTP2 speculative decoding. I capture the actual concurrency shapes I serve rather than pretending one graph is universal. Fused operators and graph replay matter enormously when a decode step otherwise consists of many small kernel launches.

  1. Optimize the service, not just an isolated kernel

Several kernels won a microbenchmark and lost end to end. I kept the ones that reduced real request time and rejected or quarantined the others. Multi-request QSA, grouped MTP experts, expert-route caching, cache accounting, cold-prefill chunking, and collective overlap were all measured at the API boundary.

That last part is why four concurrent requests reach roughly 61 aggregate tok/s even though one request is around 30 tok/s. The extra work can occupy parts of the machine that a single token stream leaves idle.

Qwen performance today

These are milestones from different stages and workloads, not one controlled single-variable benchmark:

Qwen3.8 Flash-Next milestone / Measured result

Earliest uncontrolled service: Roughly 0.2–1.9 tok/s, often incoherent

Correct eager W4 dequant reference: 0.229 tok/s

show remaining 8,027 characters

Stable W8 service baseline: About 18–19 tok/s

Native W4A8, MTP2, graphs, one request: 29.74 tok/s median; 34.34 peak

Same optimized service, four requests: 60.88 tok/s aggregate median

Two concurrent 40K warm-prefix requests: 30.76 tok/s aggregate

40K cold prompt: About 120 seconds TTFT; 30–32 tok/s afterward

The 30/61 tok/s figures are short, fixed-output decode tests. Long reasoning requests tell a less flattering and more useful story.

https://reddit.com/link/1wvt1m4/video/rpc5c2c1s1th1/player

For quality, I ran all 198 GPQA Diamond questions on the four-chip Ascend W4 service and, as a reference, an RTX 6000 Pro running a different IQ4\_XS GGUF in llama.cpp. Both scored 140/198 (70.71%) with the same AISBench-style answer extractor. This is evidence that the Ascend path is coherent; it is not a pure hardware or quantization comparison because the runtimes, quantizations, chat templates, and concurrency differ.

On the final uninterrupted 106-case Ascend phase, four workers emitted 507,256 tokens in 2 h 52 m 46 s: 48.93 aggregate tok/s. Median per-request client rate was 12.50 tok/s on these long reasoning generations. The server completed every request in that phase with no eager fallback or zero-acceptance interval. The RTX reference was much faster per request, so I am not presenting this as an NVIDIA killer. I am presenting it as a large model working correctly and usefully on hardware that initially produced slow nonsense.

GLM is my next hard(er) model

I am also bringing up the roughly 304B-parameter GLM-5.3-Flash architecture. It combines 34 KDA linear-attention layers, 11 DSA sparse-attention layers, 288 experts, latent MLA history, mHC mixing, and a mixed W2/W4 expert checkpoint. The checkpoint is about 151.6 GiB, with approximately 35.6 GB of loaded weights per rank in one four-rank profile.

It now loads and generates on the same machine. The speed progression so far has been:

GLM four-chip milestone / One request / Four-request aggregate

Initial full-service profile / 0.764 tok/s / 1.474 tok/s

Batched NZ workspace write / 0.912 tok/s / 1.822 tok/s

NZ-packed code layout / 1.386 tok/s / 2.946 tok/s

Fused mHC/MLA work / 1.743 tok/s / 3.269 tok/s

Latest measured integrated runner / 2.236 tok/s / 4.476 tok/s

An 8,232-token GLM prompt prefills at about 39.1 prompt tok/s and then decodes at about 2.16 tok/s. Those numbers are far from my target, and the 32-token test completion was too short to establish answer quality. GLM is currently a bring-up and optimization result, not a service recommendation.

I have already found an important quantization lesson there. An early W2 checkpoint showed residual growth all the way to an RMS around 525 in the last layer. Moving the affected experts to a no-clip W4 treatment kept the network bounded. Separately, my custom blocked-dequant Cube kernel was about 17 times faster than the eager reference in its isolated test. As Qwen taught me, both numeric behavior and end-to-end integration have to pass before either result means “done.”

The third card: 50% more memory and cores, not just a spare

I plan to add a third Atlas 300I Duo to this system. That takes the machine from four to six 310P devices and from 192 GB to 288 GB of nameplate device memory. With the same ECC and runtime reservations, I expect roughly another 86 GiB of runtime-visible capacity, for approximately 258 GiB across the six ranks.

The n-card architecture is designed to use it. Expert ownership is distributed across ranks, so the two new chips add local expert capacity and accelerator cores; they are not merely passive storage. Tokens route to the ranks that own their experts, and the additional ranks participate in the model's compute and collectives. I expect useful scale from EP6, although the exact speedup will be measured rather than advertised in advance.

The extra capacity gives me several options:

• Keep larger expert sets or higher-precision layers resident.

• Fit models that are just over the four-chip limit without host offload.

• Spend more memory on long-context state and cache.

• Reduce aggressive quantization where the quality trade is not worthwhile.

• Run a large distributed model while retaining room for another smaller service or evaluation workload.

The cost is one more card, 150 W of maximum board power, two more device ranks, and another passive heatsink that needs real airflow. In the cooled room, none of those are architectural concerns. The interesting cost is communication: six ranks change route balance, collective sizes, and PCIe/HCCL traffic. I will tune and benchmark that topology, but the software is already organized around n-card expert distribution rather than hard-coded four-way ownership.

Known issues and active work

This is what is still on my bench:

• I just fixed one real cross-stream race: a prefix-Mamba state slot could be spilled or reused before its pending NPU writer completed, allowing an older checkpoint to be restored. That fix has an NPU regression test.. A later 106-case run survived 12 state spills without degrading, which is encouraging but not a root-cause proof. I am continuing long mixed-load and eviction/reuse soaks.

• Qwen cold prefill. A 40K cold prompt still takes roughly two minutes even though subsequent decode is fast. Profiling points primarily at the 12 QSA layers, especially sparse selection and tiled attention. This is now a more important target than another tiny decode micro-optimization.

• Qwen long-context qualification. The planner has the capacity, but I am separating configured context, allocated capacity, and completed end-to-end long-context tests. I want retrieval and concurrent-fill evidence, not a screenshot of a launch flag.

• GLM coherence and performance. I am requalifying the full model after KDA, MLA, QSA, and runner integration changes, then moving the grouped mixed W2/W4 expert path, cache layout, graph replay, and eventually MTP through the same correctness-first gates used for Qwen.

• GLM loading and memory. The filtered loader can skip large amounts of peer-owned or superseded checkpoint payload before tensor materialization. I saw one roughly 20% load-time improvement, but it needs controlled reruns and byte-accounting before I call it a result.

• Six-device expert parallelism. When the third card arrives, I will measure rank balance, per-card temperatures, collective time, model capacity, and c1/cN throughput on the exact six-rank topology.

• Other model adapters. The same loader, packed-expert, cache, and operator infrastructure is feeding ongoing DeepSeek and other hybrid/MoE work. I am avoiding model-name conditionals where the underlying contract can be made generic.

Should you buy one?

I plan to list my first additional card in the OpenSensor storefront (https://www.opensensor.io/) next week at $2,900. The listing is not live yet. I believe it is a good price for what the hardware can already do and what the software should unlock. It is not yet the same kind of turnkey purchase as a supported NVIDIA card running mainline vLLM.

I think the right buyer is a developer, lab, or systems-minded end user who is comfortable with both of the following:

  1. You own the airflow solution. These are passively cooled server cards. Depending on the chassis and motherboard, that may mean high-static-pressure case fans, a duct, or a 3D-printed shroud. I use Fusion 360 and am happy to help with additional shroud designs. There are too many motherboard layouts, card spacings, fan sizes, and case geometries to pretend that one printable design will fit everything.
  2. You are adopting an active development fork. I am the sole developer on the vLLM work today as upstream is focused more on their server grade accelerator modules not yet available to the US. The progress is real and the benchmarks in this post are from actual hardware, but more bugs and better approaches will be discovered. Buyers shoul
💬 118 (+6) open on reddit ↗
▲
66
+24
60👁
r/LocalLLaMA · u/Postmodern_Plunger · 7d ago
Inference Engineering for Dummies

Hi all! I am a former SWE who has recently transitioned into inference engineering. I launched a side hustle a few months back and I've just taken it full time due to excessive demand. Its been such an opportunity because the people optimizing runtime are far fewer than the people that are trying to build apps or offer AI solutions.

So I've been lurking around this community, and I've noticed a lot of people who seem to have massively suboptimal setups for their hardware, and I've grouped the biggest errors into several buckets. The purpose of this guide is to expose common inference bottlenecks and provide best practices for avoiding them within your hardware constraints.

RUNTIME:

I. Choosing the Right Runtime

This is the biggest mistake I see. Choosing the correct runtime for your architecture and model is the most important decision to make. In general, here are some rules to help you determine what runtime to use.

Firstly, Ollama is never optimal. Its just the simplest. If you want quick and easy and have extra RAM, it's a good place to start. It's very user friendly and requires less setup. But it just won't offer best inference speeds.

If your model requires cpu offload, then llama.cpp will be your best choice. If not and you're solely in gpu, VLLM will likely provide the best results. It's really as simple as that for 90% of cases. SGLang may be worth it if your workload involves Langgraph, as it is highly optimized for the tooling. Otherwise, stick to the above. Mainline branches are best, with community forks offering only highly niche performance boosts (i.e., for specific models/configurations, but are generally under optimized and not well maintained).

II. Optimizing and Maintaining Runtime

The other big mistake people make with runtime is failing to compile it with hardware specific flags. Not going to go through all of them here, Google can help you out. Just search "optimal runtime compilation flags for \[runtime\] using \[GPU, CPU, RAM type\]." The most missed/missed flags tend to be for architecture specific optimizations. Those are crucial.

Runtime should be recompiled (with optimal flags) any time \*any\* of the following occur:

\- System updates

\- Kernel/driver updates

\- Running a model released or modified later than your last compile

\- You haven't recompiled in over a month (recent updates often contain kernel or path optimizations)

MODEL CHOICE:

I. Quantization:

Quantization. Such a big word. Such little meaning. All you need to know is that it makes a model smaller. There are a million Q\_K\_X\_&$&$&$ quant sizes, so I'm not going to go over them individually. Rather, I will provide basic principles.

\- IQ quants are generally the best for the size. If you're choosing between IQ4\_XS or Q4\_K\_M, IQ4\_XS is both a smaller VRAM footprint and higher complexity.

\- Nonlinear (NL) quants are only ever going to be better if you have CPU offload. Even then, IQ quants often offer extra context space vs NL quants and thus are preferable.

\- If it's a quant you've never encountered, read the docs. It more than likely is highly optimized for that specific model and is indeed the one you should choose. Searching it or asking chatgpt \*will give you the wrong answer every single time for custom quants.\* You will only encounter these with custom tuned models.

\- Standard Q\_K\_M quants are best if you have absolutely no hardware constraints for the model you're running, as path optimizations are the best. If you have no hardware constraints, though, you could be running a better model. This is only useful for running simple models for simple tasks.

II. Task

Certain models excel at certain tasks. This is subjective and preference based but this is my list:

Coding: for API, anthropic. Hands down the best models. Opus and Sonnet 5.5 both excel in performance and their low token usage per task makes them more affordable than previous iterations. Deepseek models are the best budget choice. Qwen models are the clear winner for local inference on all fronts.

Writing: Opus/sonnet for technical writing, Gemini for creative writing. Gemma for local creative writing.

Video: Wan 2.2 for local in most cases will get it done, chatgpt and copilot both have surprisingly robust free image/video gen, Veo is the best paid.

HARDWARE:

Buy a budget box, or build your own from parts. I've managed to squeeze better performance out of an RTX 3060 and 128 gb RAM than a DGX spark across all categories for multiple models. The spark has an edge for dense models, but I was able to run higher complexity models overall on the other setup for 1/5 the price. AMD and Intel lag significantly on speed per price, but I've heard Intel has had some major gains recently. Have not confirmed myself though.

MODEL OPTIMIZATIONS:

I. Spec Decode (MTP)

\- If you have CPU offload, spec decode will \*always\* slow you down. The extra overhead compute isn't worth it if you don't have at least several hundred Mb/s bandwidth, which your CPU won't.

\- MTP is sometimes a baked in feature, and sometimes requires a special secondary model. Ensure you know how it works for your model and what flags to run.

\- Each model will be optimized for exactly 0-1 type of spec decode. Figure out which one it is (or isnt) rather than wasting your time testing methods.

II. Model Tuning

Just to show the kind of command optimization you can get, here is my sample command for running Qwen 3.8 Flash-Next, a 156b parameter model, on 12 gb VRAM (and 128 gb RAM) at 10 token/s decode and 150 token/s profile at 200k context:

\~/llama.cpp/build/bin/llama-server --flash-attn on --batch-size 1024 --ubatch-size 1024 --no-warmup --cache-reuse 256 --jinja --host 0.0.0.0 --port 8090 --presence-penalty 0.0 --repeat-penalty 1.0 -m \~/llama.cpp/LLM/Qwen3.8-Flash-IQ4\_XS/UD-IQ4\_XS/Qwen3.8-Flash-Next-UD-IQ4\_XS-00001-of-00003.gguf --temp 0.95 --top-k 20 --top-p 0.97 --min-p 0.05 -np 1 --chat-template-kwargs '{"enable\_thinking": true, "preserve\_thinking": true, "reasoning\_effort": "xhigh"}' --threads-batch 16 --threads 8 --gpu-layers 150 --n-cpu-moe 48 -c 200000 --override-tensor per\_layer\_token\_embd.weight=CPU -ctv q8\_0 -ctk q8\_0 --cache-ram 8192 --checkpoint-min-step 512 --ctx-checkpoints 4 --kv-unified --reasoning-preserve --load-mode mmap+mlock

That's a lot, right? It's every possible optimization you could apply. I'll go through them individually. This is llama.cpp specific, but you'll find the same flags with slightly different syntax apply to other runtimes.

\-flash-attn (-fa) on: forces flash attention optimizations and paths. Explicitly set to on to override any fallback. Auto can be optimal if the model is recent and is not yet optimized.

\-batch/-ubatch: batch is the decode chunks, ubatch is the prefill chunks. They must be multiples of one another, otherwise you're adding compute. Equal to one another is ideal for CPU offload, and a 2x-4x higher batch is optimal for full GPU loads. You'll need to play with these values to optimize. Batch/ubatch should be a power of 2 to optimize architecture. Intervals of 256 is typically good enough for testing.

\--no-warmup: prevents initial model poll to load weights. Removes unnecessary latency

\-- cache-reuse x: instructs the model to reuse cache values and scan for similarity at x token intervals

\-jinja: highly underutilized and important flag. Utilizes native chat template kwargs to ensure output consistency.

\--presence-penalty: flat penalty rate to words that appear in text. Used mainly for creative writing to prevent repetitive prose.

\--repeat-penalty: reduces liklihood of already used tokens being reused. Best used for preventing loops in thinking agents.

\-temp: model temperature– how creative the model is. 0 is completely deterministic, 1 is creative freedom.

Top-k: hard cutoff that keeps only the k most likely words. Each model will have recommended k values for thinking/instruct setups. Low key reduces hallucinations at the cost of repetition and loss of creativity

Top-p: includes P percentage of possible tokens. It reduces liklihood of hallucination dynamically.

Min-p: dynamic cutoff based on highest probability token. If the biggest probability token is 50% and min-p is 0.05, then the bottom 0.025 (2.5%) liklihood tokens will be excluded. Reduces noise without hampering creativity terribly.

Chat template kwargs: explicit chat template activations; newer runtime compilations should have flags for these. Controls model reasoning, reasoning effort, and internal chain of thought storage.

\-threads (-t): number of cores used for decode. Set equal to physical cores (hyperthreading will thrash cores and degrade results)

\-threads-batch(-tb): number of cores used for prefill. Set to double the number of physical cores, as hyperthreading helps here.

\--gpu-layers (-ngl) : total layers on GPU. Fit as many as you can without OOM.

\--n-cpu-moe: number of MoE layers offloaded to cpu. For MoE models, you should always offload these first and keep all layers on GPU if possible. Offload as few as possible to CPU.

\--override-tensor...: tells the runtime to offload the n-gram table if needed. For qwen 3.8 flash specifically.

\-ctk/-ctv: k and v cache quantization. K cache should \*never\* be below q8 unless youre running on less than 80k context. V cache can be q4 up to 150k context without issues, for the most part. Generally, q8 for both will be best for speed and is my recommendation to start with.

\--cache-ram: sets RAM aside for cache allocation to ensure it doesn't go to swap

\--context-checkpoints: the amount of checkpoints captured. Generally, you don't need as high as the defaults do. Leave the default if you have extra RAM, otherwise you may want to lower it.

\--kv-unified: tells all instances to run on the same kv cache pool rather than allocating individual cache.

\--reasoning-preserve: tells the runtime to retain CoT traces for evaluation. Prevents the model from getting stuck or looping as much when thinking.

\--load-mode: tells the runtime how to load in the model. No mmap generally loads slower but is more stable. Mmap+mlock (or just mlock) is a balance of both with fast loading and page faults initially but it stabilizes as you run it, mmap alone is fast but will cause constant page faults and slows down inference, especially on models with CPU offloading.

I hope this guide helps! I'd be willing to answer any specific questions or make any additions if there are additional areas the community agrees are major uncovered inference bottlenecks. Some claims are based on my personal experience and I am open to data based claim revisions or anecdotal counterclaims, so feel free to provide. Happy tuning!

💬 53 (+10) open on reddit ↗
▲
72
+22
45👁
r/LocalLLaMA · u/lbgos_Loss783 · 7d ago
I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking

Hey local AI community, I've been working on this for a while and finally feel ok sharing it.

It's a cyber benchmark where the model gets a shell in an isolated docker box and has to find the exact flag. Pwn, web, crypto, rev, forensics, a few real CVEs and some multi-stage ranges. 19 tasks, 6 models, 544 scored attempts.

To be clear, I didn't build every task by hand. GLM 5.3 helped me create several of them. For an open model its cyber capability is really high, and it barely refuses anything, so it was one of the best options for this. GLM 5.3 isn't one of the benchmarked models.

The local part: I ran Qwen3.8 27B (Unsloth Q4\_K\_XL, xhigh) on a llama.cpp RPC pool across a 3090 and a 3080 in two Proxmox nodes, connected over a direct 2.5G link. That gave me enough concurrent tps to run several agents at once. I also started a low reasoning run, but it was taking 20+ hours because of a harness problem, so I killed it.

Why I'm posting now: John Hammond put out a video about how threat actors use AI (https://www.youtube.com/watch?v=xHDc6-7bjyw). One part is a guide from a criminal forum on running abliterated models on RunPod, and one of the models in it is Qwen3.8 27B. I had benchmark data on that exact model, so here's what it can actually do.

Stock Qwen3.8 27B got 28.1% on the first try and 0% on pwn. Not bad for a 27B on two gaming cards, but not much of a threat on its own either.

The cheap API models are a different story:

\- MiMo 2.6 Flash solved 73.7% on the first try

\- GPT-6 Luna solved 90.9% within 3 tries

\- On multi-stage ranges, where you chain several steps, the top models got 92-96%

Pwn is still hard for everyone (best was 56%), and 3 tasks haven't been solved by any model in 82 attempts.

Results: https://lbgos.dev/bench

Harness (MIT): https://github.com/lbgos/rangebench-harness

The tasks aren't public so they don't leak into training data, but you can still run them. DM me here or on X (lbgosna), and I'll send them over. You run it on your hardware or tokens and I'll add your results to the board. If a few people send local runs, I'll make a separate local-only table.

This is my first time building something like this, so any feedback on methodology, task mix or what's missing is welcome.

💬 40 (+9) open on reddit ↗
▲
74
+21
10👁
▲
129
+19
69👁
r/LocalLLaMA · u/aya-ifm · 8d ago
AMA about K2 Horizon, Meet our team from IFM

Hi r/LocalLLaMA

We’re researchers at the Institute of Foundation Models (IFM), an AI research lab dedicated to open and independent development of frontier-class foundation models.

We recently released K2 Horizon a connected fleet of six fully open models with size ranging from 0.9B to 375B. In addition to weights, we also open-sourced training data and recipes, training code, intermediate checkpoints, fine-grained training logs and evals.

Ask us anything about pre-training and data mixes, post-training, small models on-device, MoVA and sparse attention, deployment, what’s out now, and what’s coming next.

Participating in the AMA:

  • Hector Liu u/hunterhector
  • Alexander Moreno u/IFMAlex
  • Mikhail Yurochkin u/my-moonfolk
  • Rupesh Srivastava u/k2pt
  • Junlin Chen u/Junlin_Chen110
  • Haonan Li u/East-Career9147

We'll be live Mon, Oct 5, 8–10 PM PT. Questions are open now, so drop yours anytime!

Join IFM on Discord: https://ifm.ai/discord

https://preview.redd.it/awu0r6fbowsh1.png?width=3240&format=png&auto=…

💬 135 (+108) open on reddit ↗
▲
255
+17
52👁
r/LocalLLaMA · u/jacobpederson · 8d ago
Why am I like this? (Full Chat and Image generation on a 286 Tandy 1000 TL/3) post image

40 year tech gap? No problem! The Tandy runs DeskMind, a native DOS program. It talks over WiFi (a PicoMEM 2 card with mTCP) to a small Python server on my PC. That server drives Qwen3.8-27B (NInfer on a 5090) and Krea 2 (ComfyUI on a 4090). The 286 never sees JSON, base64 or a PNG. It gets plain text lines and pictures that are ready to copy into video memory.

Drawing from chat without tool calling. The system prompt tells Qwen to wrap a picture request in \<draw>...</draw>\. The server catches the tag mid-stream, runs Krea 2, dithers the result, and streams a \picture ready\ line. "Draw me a 286 AI logo" takes about 9 s from Enter to a thumbnail in the chat.

Qwen Vision sees what the Tandy sees. When you ask about a picture, Qwen gets the original and the 16-colour dithered version, so "why does the sky look striped?" has context. The latest picture stays in context for follow-ups.

The model knows where it lives. The system prompt knows it's talking IN a 286 with 80 columns and 16 colors. It keeps answers short and plain ASCII, and when asked about games it suggests Wolfenstein 3D or Commander Keen.

Streaming cleanup for a 1990 screen. Reasoning is stripped, Markdown is removed on the fly, Unicode becomes code page 437 (bullets turn into the CP437 block character), and tiny tokens are merged into \~48-character lines so the 286 isn't redrawing for every token.

Per-request reasoning effort: low for chat (replies start in \~2 s), medium for rewriting image prompts.

Prompt "enhancement" tuned for dithering: bold shapes, strong contrast, simple backgrounds. The rewrite shows up in an edit box on the Tandy before drawing, and the rules themselves can be edited from the Tandy.

\- \*\*A Dither Lab\*\* in the server GUI: Floyd-Steinberg, Atkinson, Bayer, Yliluoma and more, previewed at the Tandy's real aspect ratio.

Numbers: Krea 2 at 1024x768 in \~10 s 8 steps, the target is 640x200). About 1 s to send a 64,000-byte picture over the PicoMEM WiFi (56-79 KB/s).

Code (GPLv3): https://github.com/RowanUnderwood/DeskMind

added image gallery https://imgur.com/a/wysqoM1

💬 129 (+12) open on reddit ↗
▲
179
+16
54👁
r/LocalLLaMA · u/ea_man · 8d ago
Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now post image

pi-llama-skip-reasoning is an extension for the Pi.dev harness that forces a local llama.cpp model to stop reasoning and answer / act immediately.

When you are deep into the ctx session and ask 27B a simple question about a fact or need a direct action, the model may still feel the urge to indulge in copious deliberation in the reasoning trace. This extension allows the user to force the model to snap out of the reasoning stage and provide the answer immediately.

Disclaimer: don't skip the reasoning for important problem-solving, that would hurt quality.

This uses the same mechanism the llama.cpp web interface uses to skip reasoning, so it's native to llama.cpp, this extension is meant for Pi.dev yet the same mechanism could work for other harnesses.

Usage: /skip-reasoning command or shortcut Alt+T ,
Install: pi install npm:pi-llama-skip-reasoning

\- https://pi.dev/packages/pi-llama-skip-reasoning

💬 82 (+4) open on reddit ↗
▲
50
+16
42👁
r/LocalLLaMA · u/menage_a_un · 7d ago
I've ended up with an AI lab in a public community college. What should we actually be teaching?

Looking for some ideas from people who know a lot more about this than I do.
We've got funding for a small AI lab in a public further education college in Ireland (roughly community college in the US).
The hardware is reasonably decent. The goal is to give students useful skills beyond just using ChatGPT.
If you had the lab, what would you teach them?

💬 70 (+18) open on reddit ↗
▲
53
+16
46👁
r/LocalLLaMA · u/EmPips · 7d ago
Heavily quantized Qwen3.8-Flash vs Q8 Qwen3.8-27B - thoughts?

I'm currently choosing between IQ3_XXS Qwen-3.8-Flash and Q8_0 27B.

This month I don't have anything complex enough to justify either's potential so I've got a fairly bad read on how these two stack up in terms of intelligence.

Have any of you compared the two enough to speak to which you've had a better experience with?

💬 84 (+27) open on reddit ↗
▲
74
+16
38👁
▲
26
+13
35👁
r/LocalLLaMA · u/wayneworkman · 7d ago
Peacebell - a from-scratch small language model

I've been using my free time during weeknights and weekends for the last 11 months working on and refining a small domain-specific language model. It specializes on information about World War II.

The number one question I get asked about this is "Why did you pick World War II?" here are some of the reasons:

\- There's a lot of good Wikipedia articles about WWII, and this is permissively licensed. Meaning I can use the materials.
\- There's a lot of good public domain information about WWII in general - more to train on.
\- WWII is factually dense - making it a challenge.
\- The facts surrounding WWII are mostly unchanging - meaning my model would age well.
\- I had to pick a first topic.

A lot of my journey is documented on my blog: https://wayne.theworkmans.us/llm.html though I've not posted recently.

The model is more than from-scratch. I'm using a custom built training pipeline. And I produced all of my own synthetic data to train on (based on Wikipedia articles).

The majority of my time has gone into data curation and balancing.

As I built this model, I've learned a ton about training data for language models, and about language model creation. I tripped over every bump along the way, 100s of times.

I also learned a lot about World War II, and I'm emotionally exhausted. Many know the basics... the Manhattan project, the Holocaust, the concentration camps, Pearl Harbor, D-Day. Though beyond these topics, there is enormously more tragedy than I previously knew. As an adult with my own family now with better ability to comprehend, many times I'd just cry face down on my keyboard from some of the things I learned. Sometimes I would abandon working on it and go to bed early. I've talked with my wife about how awful some of the things that happened are. It's been hard. And I'm ANGRY! So incredibly angry about the atrocities that happened. Especially angry about the things that happened to civilians, non-combatants, women and children.

Well enough of that.

I open-sourced the training materials and the weights. There are two versions of the model. There's a 291M parameter version and a smaller 148M parameter version.

I built the 148M to compete in the various HuggingFace dashboards that limit model size to 150M. Then I built a new benchmark that focuses on WWII topics, that's also on the hub, though the questions are private to prevent them ending up in people's training data (and no they aren't in Peacebell's training data either).

You can try the 291M model for free here. As you use it, keep in mind this is first-version, it's rough, it's not always right. And it really struggles with longer context. Fresh context gives better results.
https://huggingface.co/spaces/wayneworkman2012/peacebell-v1-291M-demo-cpu

The leaderboard is here:
https://huggingface.co/spaces/wayneworkman2012/ww2bench-leaderboard

I've entered Peacebell into various SLM Arena's, such as CodeSoft's SLM arena here:
https://huggingface.co/spaces/CodeSoft/SLM-Arena

Basically everyone in this LocalLLaMA would be able to run the model easily, even without a GPU. There's a customized vLLM fork here that can run either Peacebell model:
https://github.com/wayneworkman/vllm

Next version is expected to be released sometime in 2027.

💬 8 (+3) open on reddit ↗
▲
41
+12
29👁
r/LocalLLaMA · u/Terminator857 · 8d ago
China and the memory market

Once china sets its goals for dominating a market it wins. Usually takes many years, but it happens. Can't compare the political will of a country versus profit and loss thinking of a corporation.

China will eventually win in the memory market and current memory makers are at an unfair disadvantage.

https://www.tweaktown.com/news/112680/chinas-cxmt-is-on-track-to-nearly-match-microns-dram-production-capacity-by-the-end-of-2026/index.html Quote:

CXMT will finish 2026 with approximately 350,000 wafer starts per month (WSPM) of DRAM capacity, which is just 25,000 WPM less than Micron.

... by 2030, its total capacity will increase to around 1.41 million WSPM, according to Citrini. CXMT alone is projected to build new production capacities in Beijing, Hefei, and Shanghai, to expand its production capability to 950,000 WSPM in 2030, assuming everything goes as planned.

/end quote

Is China hoping for a RAM price drop crash to extinguish the competition?

Additional references:

  1. https://www.tomshardware.com/pc-components/dram/chinas-cxmt-targets-30-percent-dram-memory-market-share-by-2030-with-sixth-mega-fab-future-plans-bottlenecked-by-access-to-advanced-chipmaking-tools
  2. https://www.trendforce.com/news/2026/09/24/news-cxmt-ymtc-ramp-memory-capacity-but-chinas-ai-cloud-boom-could-soak-up-new-supply-through-2027/
  3. Microns quarterly report: https://www.investing.com/news/company-news/micron-fq4-2026-slides-record-revenue-ai-demand-drives-supply-tightness-93CH-4926074 . Turn off javascript to view.
💬 34 (+3) open on reddit ↗
▲
126
+11
83👁
r/LocalLLaMA · u/FutureStriking283 · 7d ago
Anyone wonder why americans lag so far behind in the open LLM market?

I mean , DeepSeek, Kimi, GLM, MiniMax -- the list of Chinese LLM's is such a long freaking list. As American's -- why don't we feel .. a little funny .. about being so far behind? Chinese entrepreneurs are peneuring like crazy and American's .. just obsess on .. what ?

update -- already I'm starting to see some clear answers. American's are putting money ahead of technology. Completely understandable.

second update -- I asked "why can't we have american AI as good as or better than the chinese" and BY FAR the number one most supported comment? "Have you even thought of shareholder value!?" . We .. America , are so fucked.

💬 425 (+28) open on reddit ↗
▲
29
+11
32👁
r/LocalLLaMA · u/jjusko20 · 7d ago
Update #3: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Last update for those who may be following: https://www.reddit.com/r/LocalLLaMA/comments/1wv5h8x/update\_2\_post\_training\_yandexaliceai80ba3b/

I screwed up guys 😂

Turns out my loss curve during my last run was legitimately unhealthy - as some of you, and myself, were concerned about. After evaluating my QLoRA, I found zero'd gradients in all but two layers. Turns out I had a NaN issue related to my custom v100 kernels that I didn't catch - so that run is cooked, I had to restart. I guess two layers training managed to emulate a loss curve I could at least derive a sensible explanation for until I actually got to evaluate.

Thankfully, checked to make sure gradients were applying again, and restarted the run. Once again, it's live streaming at https://figure-bios-expect-cio.trycloudflare.com/

Loss curve looks much healthier this time and is making me feel more confident that this is going to be okay. Stay tuned! I'm gonna release GGUFs and a llama.cpp patch when I have a working version.

My first epoch loss curve from this run

my first epoch loss curve on the failed first run \(note the differences in scale even if the pattern looks similar\)

▲
18
+10
28👁
r/LocalLLaMA · u/Apprehensive_Side219 · 7d ago
Alternative to Nvidia spark?

I just spent the last two weeks trying to catch a microcenter in stock with a spark and then today the price went up 30% and now I can't realistically afford it. It was already pretty close to the edge of my budget, and now I don't think I can swing 7k for what I had been expecting to pay 5 for last week. Any suggestions for alternative approaches welcome. I really don't want to wait until 2028 to get started.

💬 35 (+10) open on reddit ↗
▲
43
+10
36👁
r/LocalLLaMA · u/SnooPredictions515 · 8d ago
Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

I've been working on getting the 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac. In my earlier post, I shared a custom expert-streaming fork of llama.cpp . It worked, but decode capped out around \~23–27 tok/s and slowed down as context grew.

Today I'm releasing Slipstream: a compiled C++ Metal inference engine with native SSD expert streaming and speculative drafting for Apple Silicon.

The main result: If you already downloaded my original V3 model (34k+ downloads), you don't need to re-download anything. You can run that exact checkpoint on Slipstream for a 1.76x speedup: 41–52 tok/s (up from 23.1 tok/s in llama.cpp) on the same 64GB Mac.

Even better: decode speed doesn't collapse at long context. Across 3,086 live requests in real coding sessions, it stays flat at 33–44 tok/s all the way out to 130,000 tokens.

Previous posts for context:

Open source resources:

1. How to run your existing V3 model on Slipstream

If you have the model from the last post (~/models/qwen38-flash-next-v3), you can point Slipstream directly at it.

Step 1: Clone & build (under 1 minute)

git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4

Step 2: Download the model (if you don't already have it)

Downloads the 3 GGUF shards + MTP draft head (~95.5 GiB total) huggingface-cli download nitinpanj/qwen38-flash-next-v3 \ --local-dir ~/models/qwen38-flash-next-v3

Step 3: Raise wired GPU memory limit & serve

Raise wired GPU memory limit once per boot (required on 64 GB Macs): sudo sysctl iogpu.wired_limit_mb=59392 # Serve your existing model: ./slipstream serve --model ~/models/qwen38-flash-next-v3 --port 8090

First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into <model>/prepared/ (\~5–7 minutes). Subsequent launches load in \~10–15 seconds.

The server exposes a standard OpenAI-compatible API (http://127.0.0.1:8090/v1/chat/completions) ready for curl, Oh My Pi (omp), Claude Code, or OpenCode.

2. Speed: llama.cpp Fork vs. Slipstream (Same V3 Checkpoint)

Here is a direct head-to-head comparison running the exact same 95.5 GiB model files across 6 reasoning and coding tasks on the same M5 Pro (64 GB unified memory, temperature 0.0):

|Domain / Task|Prompt Task|llama.cpp Fork|Slipstream|Speedup|llama.cpp TTFT|Slipstream TTFT|
|:-|:-|:-|:-|:-|:-|:-|
|Math Reasoning|GSM8K (eggs problem)|24.0 tok/s|43.6 tok/s|1.82x|4,024 ms|2,337 ms|
|Math Derivation|MATH-500 series ($p - q$)|24.3 tok/s|43.1 tok/s|1.77x|1,655 ms|1,587 ms|
|Constraint Logic|3-chair deduction|25.4 tok/s|46.0 tok/s|1.81x|1,469 ms|1,042 ms|
|Python Coding|merge_intervals ($O(N log N)$)|19.7 tok/s|35.0 tok/s|1.77x|1,507 ms|1,070 ms|
|Systems Coding|Rust CSV parser|22.7 tok/s|37.5 tok/s|1.65x|1,257 ms|859 ms|
|Tech Writing|Multi-head attention|22.5 tok/s|39.4 tok/s|1.75x|1,267 ms|843 ms|
|AVERAGE|Across all 6 tasks|23.1 tok/s|40.8 tok/s|1.76x|1,863 ms|1,290 ms|

https://preview.redd.it/vh1danbnzwsh1.png?width=3000&format=png&auto=…

What made Slipstream faster:

  1. Asynchronous layer-ahead prefetch (fcntl(F_RDADVISE)): In llama.cpp, synchronous page reads for missed expert matrices stalled the GPU on NVMe latency (\~475 ms per chunk). In Slipstream, non-blocking read-ahead hints stream upcoming expert layers from SSD into RAM while the GPU is still executing the previous layer, cutting prefill staging latency by 28%.
  2. Hybrid MTP + Prompt Lookup speculation: During tool calls and code generation, Prompt Lookup Decoding (PLD) matches prompt anchors in under 50 ns with 0 allocations, preventing draft rejections. This lifted tool-calling decode from 5.6 tok/s to over 45 tok/s.
  3. Metal GPU-mapped n-gram tables: llama.cpp faulted on the 26.8 GiB n-gram table during prefill. Slipstream maps and gathers n-gram embeddings directly in Metal kernels.

3. Context Scaling: Real Telemetry up to 130,000 Tokens

On standard Transformers, decode slows down sharply as context grows because the KV cache swells and memory bandwidth saturates.

Qwen3.8-Flash-Next avoids that through its hybrid architecture:

  • 48 recurrent linear DeltaNet layers (fixed $128 \\times 128$ hidden state, $O(1)$ memory growth with context).
  • Only 16 full-attention layers.

Here is actual telemetry collected across 3,086 live requests during real agent coding sessions on my M5 Pro (64 GB):

|Context Range (Tokens)|Live Runs|Average Decode|Median (p50)|Peak Decode|Average TTFT|Notes|
|:-|:-|:-|:-|:-|:-|:-|
|< 1,000|314|41.5 tok/s|41.9 tok/s|59.8 tok/s|2.16 s|Short baseline|
|1k – 4,000|21|41.0 tok/s|42.5 tok/s|64.5 tok/s|5.26 s|Small documents|
|4k – 8,000|58|43.6 tok/s|43.2 tok/s|67.2 tok/s|7.36 s|Code review turns|
|8k – 16,000|117|43.6 tok/s|44.6 tok/s|58.2 tok/s|7.91 s|Multi-file context|
|16k – 32,000|562|38.2 tok/s|40.9 tok/s|58.0 tok/s|13.59 s|Deep agent session|
|32k – 64,000|1,029|35.0 tok/s|37.5 tok/s|55.6 tok/s|13.24 s|Large repo refactor|
|64k – 96,000|650|32.4 tok/s|34.7 tok/s|53.9 tok/s|12.81 s|Multi-turn transcript|
|96k – 130,000|364|32.9 tok/s|33.3 tok/s|43.8 tok/s|7.95 s|Cache-hit deep turns|

https://preview.redd.it/qruz5nmpzwsh1.png?width=3300&format=png&auto=…

Takeaway: Decode speed stays between 33 and 44 tok/s all the way out to 130k tokens. Even at 130k context, it generates tokens faster than stock llama.cpp did on a 500-token prompt.

4. Optional: Swift KV-Sparse Model Variant

If you want higher reasoning accuracy and lower KV cache memory, I also put together an optional Swift variant of this model: Swift-Qwen3.8-Flash-Next-V3.

What Swift changes:

  • KV-Sparse Attention: Replaces standard dense attention with KV-sparse layers distilled from Swift-1.5, cutting down RAM pressure at long contexts.
  • Spliced Q8 Donor Backbones: Slices 686 high-precision Q8 donor tensors into resident backbone layers for sharper representations.
  • Concise Reasoning: Distilled to eliminate repetitive thinking loops in deep contexts.

Both models run on Slipstream using the exact same engine command. Here is how they compare across 145 paired evaluation problems (temperature 0.0, seed 1234):

|Domain / Benchmark|Items|Original Flash-Next V3|Swift-Flash-Next V3|Accuracy Delta|Original Decode|Swift Decode|
|:-|:-|:-|:-|:-|:-|:-|
|AIME 2025|20|45.0% (9/20)|45.0% (9/20)|0.0%|44.3 tok/s|44.3 tok/s|
|MATH-500 (L4–5)|35|60.0% (21/35)|62.9% (22/35)|+2.9%|44.8 tok/s|44.8 tok/s|
|GPQA Diamond|35|45.7% (16/35)|54.3% (19/35)|+8.6%|44.8 tok/s|44.8 tok/s|
|GSM8K|25|96.0% (24/25)|96.0% (24/25)|0.0%|45.6 tok/s|45.6 tok/s|
|HumanEval|25|92.0% (23/25)|92.0% (23/25)|0.0%|40.6 tok/s|40.6 tok/s|
|Hard Systems Logic|5|100.0% (5/5)|100.0% (5/5)|0.0%|39.2 tok/s|39.2 tok/s|
|OVERALL|145|67.6% (98/145)|70.3% (102/145)|+2.8%|43.9 tok/s|44.4 tok/s|

https://preview.redd.it/gt1edfvrzwsh1.png?width=3000&format=png&auto=…

To run the Swift model instead:

huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3

./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090

5. Foundation for Qwen4

The core primitives in Slipstream:

  • 512-route sparse MoE streaming with SSD prefetch
  • QSA (Quasi-Sparse Attention) indexer & selection kernels
  • Hyper-connection mixing and per-layer embedding gathers
  • Metal GPU-mapped n-gram table gathers
  • Single-lane speculative verification with PLD & MTP

...were built around this hybrid architecture. If Qwen4 adopts a similar blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream should be able to run Qwen4 locally on consumer unified memory hardware on day one.

6. Hardware Tested & Porting to NVIDIA / AMD

  • Hardware tested: All testing and benchmarking were done on an Apple MacBook Pro (M5 Pro, 64 GB unified memory, 2 TB SSD).
  • CUDA / ROCm ports: I don't have access to modern NVIDIA or AMD GPU hardware, so I can't build or test CUDA/ROCm backends myself.
  • If you have hardware and want to help port this: If anyone in the community has NVIDIA or AMD hardware and wants to help bring expert streaming and hybrid speculation to Linux/Windows, I'm happy to help collaborate on the port. Feel free to open an issue on the repo or DM me.

7. Credits & Upstream

  • Splash Team (Incoai): Full credit to the creators of Splash. Their C++ Metal speculative decoding design and memory architecture provided the foundation for this work. I will prepare a clean PR proposing these Flash-Next and SSD streaming extensions to the Splash upstream repo in case they want to incorporate them.
  • ds4 Team: For their insights on Metal router numerical precision (Taylor polynomial softplus expansion) and streaming scheduling designs.
  • Qwen Team: For training Qwen3.8-Flash-Next and releasing the hybrid linear MTP architecture.
  • ukisai: For the Swift-1.5 distillation work enabling KV-sparse reasoning.
  • bartowski & unsloth: For donor quants and quantization tooling.
  • mihailescu2m: For the initial expert streaming work in llama.cpp.
💬 29 (+2) open on reddit ↗
▲
18
+9
25👁
▲
43
+9
33👁
r/LocalLLaMA · u/jwhh91 · 8d ago
I "built" a local LLM radio like the Pharaohs built the pyramids

My local stack is a couple DGX Sparks, a 5090, and a 4070 TI. This radio site took three local LLMs, and we generate music, voices, and wacky sound effects.

I give you https://pilgrim.farm

It's a lot more farm than pilgrim.

💬 26 (+3) open on reddit ↗
▲
181
+8
46👁
r/LocalLLaMA · u/Creative-Type9411 · 8d ago
Finally got my 4th card in (64gb total) post image

I was waiting on the blowers for the T4s and posted this before it was finished, other than some braided cable sleeves for the fan wires its pretty much good, I was going to upgrade the CPU, but I'm getting great speeds comparatively to a CPU in my old box that had way more cores, so I don't think it's going to make a difference.

Fractal Design Torrent Mid-Tower Case w/Tinted Glass
SuperMicro X11SPA-T Motherboard
Xeon W3225
768GB DDR4 ECC 2666
4xTesla T4 16GB GPU
4x1tb Samsung 870 EVO SATA SSD Raid

Ubuntu 26.04/llama.cpp/openwebui+custom powershell harness

now i want more cards 👀

💬 108 (+5) open on reddit ↗
▲
79
+8
30👁
r/LocalLLaMA · u/Ok-Breadfruit-3523 · 7d ago
Update on the “Monstrosity”. 6 BC-250 board cluster post image

This is 6 bc-250 ex mining boards with 5 in the asrock 4u12g case they came in. After a lot of testing my current preferred setup is 4 boards running Qwen Next Flash IQ2\_XS at 100k context with around 28 tok/s for short generation and 24 tok/s at 50k with around 115 ppt. The other two boards run 3.6 35b q4 at 60 tok/s with 100k context and 450 ppt. This is all using llama with vulkan and rpc over 1gb Ethernet.If anyone has any suggestions with this beast I am all ears. I had these boards left after reselling a bunch and had never done anything with local ai before so this has been a blast. Also yes that is a cardboard box with 3 fans on top as the intake.

💬 42 (+5) open on reddit ↗
▲
30
+8
47👁
r/LocalLLaMA · u/Porespellar · 7d ago
RTX Spark laptops and mini desktops rumored to launch Oct 7th (24GB to 128GB variants possibly)

Basically a DGX Spark minus the ConnectX-7 ports. I’ve seen expected initial pricing from like $1800 to $2900. Not sure what configurations are actually at those price points.

It’s all Internet hearsay until we actually see these things ship, but it’s nice to know that it’s potentially around the corner next week, especially given DGX Spark’s insane price increases lately.

Sadly, you can’t cluster them, but getting an entire computer + GB10 equivalent chip for a little over half the price of a used 4090 seems like an ok deal in this market.

💬 51 (+19) open on reddit ↗
▲
37
+8
26👁
▲
35
+7
31👁
r/LocalLLaMA · u/Any-Winter-4079 · 8d ago
DDR4/PCIe4 vs DDR5/PCIe5 for LLMs- I benchmarked them for pre-training. What are your thoughts? post image

Hello everyone.

I've recently ran some experiments comparing DDR4/PCIe4 and DDR5/PCIe5 for AI workstations on a pre-training run, and would like to hear yours thoughts.

First of all, and as a summary of my results ( code here: https://github.com/Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training ), I rented two machines on Vast.ai, one with an H12SSL-i motherboard, an EPYC 7352, 192 GB of RAM and of course using PCIe4 (26.3 GB/s) and another with a WRX90E-SAGE SE motherboard, a 9975WX CPU, 256 GB of DDR5 RAM and PCIe5 (54.3 GB/s), and DDR5/PCIe5 is about 15-20% faster on pre-training (depending on whether you include or exclude validation and other costs) under the same number of GPUs.

With the current RAM prices, however, for the cost of 256 GB DDR5 RAM at 6400 MT/s you can get a full (extra) RTX PRO 6000 WS/Max-Q, at which point the comparison clearly favors DDR4/PCIe4 (with 2 GPUs), with about 50% extra throughput vs a single GPU at equal(ish) cost.

Now, there aren't a lot of downsides in my mind to choosing DDR4/PCIe4, but there can be a few:

  1. at least some of these DDR4/PCIe4 motherboards are on the older side, and were one to need replacement, they are not so easy to get (for example, the H12SSL-i used, I can only find it for sale as refurbished now, so who knows in a few years if it will even be available for retail).
  2. newer GPUs (as in, new NVIDIA generations) may stop working with older motherboards (meaning yes, PCIe is backwards compatible but the motherboard's BIOS/UEFI sometimes has issues during POST with newer GPUs (e.g., some older PCIe3 motherboards already have trouble recognizing Blackwell cards, and this may be the case for PCIe4 and newer cards in the future). Meaning if one were to buy a newer GPU in the future, the whole workstation may not be suitable.
  3. If one has to go ahead and bite the bullet on RAM prices, the million question is when? To me prices now are \*\*awful\*\* but so were RTX PRO 6000 prices and here we are (i.e., even higher).
  4. For pre-training, I would still choose DDR4/PCIe4, but I suspect for inference DDR5/PCIe5 might be a fair bit better than the 15-20% that it gives you on pre-training, plus we might be moving into some techniques soon such as dynamic expert/data loading into the model at runtime, which again would favor better DDR/PCIe speeds.

With all of this, I am curious if anyone has benchmarked this, and what are your thoughts on it. Would you hold out on DDR5 at the moment, and therefore go for PCIe4, or would you bite the DDR5 bullet early? Another issue with RAM is channels and DIMM count/channel, because if you want to go 'cheap' like let's only get 192 or 256 GB of DDR5 on 8 channels at 1 DIMM/channel (e.g., 8x32 to get 256), then upgrade to 512 later (when budget allows), that means you have to replace your full RAM (because all the slots are occupied, requiring new 8x64 to get 512 for instance)... And if you get fewer DIMMs like 4x64 to get 256 GB (leaving 4 DIMM slots unoccupied) then you get half the bandwidth because only 4 channels are populated. So maybe a machine that has dual DIMM support per channel is the answer to this (fully populating 8x32 to get the full bandwidth, and still allowing you to expand to another 8x32 to get 512), but in general it's a tricky point too.

So, what do you do/are you guys doing? Have you recently bought a workstation or upgraded to one, for pre-training, fine-tuning, RL, inference, whatever your use case may be, and come up with this dilemma? Are you choosing DDR4/PCIe4 as it would seem reasonable or are you going for DDR5/PCIe5 and if so, why? I am interested in all use cases and opinions!

💬 19 (+1) open on reddit ↗
▲
38
+7
32👁
r/LocalLLaMA · u/SultanGreat · 8d ago
What's the best setup for Qwen3.8 27b for a 16 gig VRAM?

Hello guys!

I have been experimenting with qwen 3.8 for a long time and I hadn't been able to get reasonable speed. I am on a 5060Ti 16 GB, and although this gpu can game, I am aware that AI demands more than 16 GB.

I am on a Fedora 44, AMD Ryzen 9600x and 16 GB system ram (16 GB system ram and 16 GB vram, totaling to 32 GB) and I would like to use llamacpp, although I would use any other tool if I could if it meant faster speed.

I am looking for a large context. Atleast 128k context. The first question is, what quantization to pick? In my experience Q3 UD was satisfying, but I am looking for uncensored model. In my experience, MTP has never lived up to its hype for me (and I don't know why!?), which is why I am thoroughly lost on making a good setup after an honest week of experimentation, which is why I have resorted to ask here as a last resort.

Update : Found a model, thanks to u/_wortkarg_

link : https://huggingface.co/RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3\_XXS-Uncensored

command (A better command would be appreciated and updated accordingly):

~/llama.cpp/build/bin/llama-server \
--model ~/Documents/Models/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf \
--alias "llamacpp" --host 0.0.0.0 --port 8001 \
-ngl 99 --flash-attn on --ctx-size 131072 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--parallel 1 --batch-size 512 --ubatch-size 256 \
--no-warmup --jinja \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

I am hitting at about 35 t/s+ speed with this one.

💬 84 (+3) open on reddit ↗
▲
15
+6
31👁
r/LocalLLaMA · u/FantasticNature7590 · 7d ago
I spent 3 weeks testing local Qwen3.8 on the new low-latency SGLang/vLLM recipes: DFlash2 2.8x. Builds: RadixArk + Inferact 27B NVFP4, 27B BF16, orcarouter 27B Uncensored, Flash-Next NVFP4

Hey guys,

Last time I tested Qwen3.8-Flash-Next on its own. This time I put three Qwen3.8 checkpoints through the same 10 tests on the same RTX PRO 6000:

  • RadixArk/Qwen3.8-27B-NVFP4 (dense)
  • orcarouter/Qwen3.8-27B-Uncensored-NVFP4 (dense, uncensored fine-tune)
  • RadixArk/Qwen3.8-Flash-Next-NVFP4 (MoE)

Each model got the same prompts and its own model card's sampler, with one attempt per task.

Video with the battles, the castles and the ball run: https://youtu.be/VOtfja\_Toj4**

Short version

  • Tests won: Flash-Next 5, 27B 4, Uncensored 0, plus one tie (long-context recall was 100% for all three).
  • Prefill, full window: Flash-Next 22.4s, 27B 97s, Uncensored 99s. Not the same power cap, see section 1.
  • Speculative decoding on the 27B: the DFlash2 drafter took Spec-Bench from 75 to 210 tok/s for one user, 2.8×.
  • SGLang vs vLLM: SGLang was faster overall (210 vs 160 tok/s), but that's mostly the checkpoint. On the one export I ran on both engines, vLLM was 16% faster (169 vs 146).
  • Tool use (BFCL subset, thinking off): 27B 73.3%, Uncensored 70.8%, Flash-Next 64.5%.
  • Battle arena: the 27B scored 700/1000, ahead of Claude Fable 5.1 (678) and GPT-5.6 (473), both entered through their chat apps at max thinking.
  • Rube Goldberg machine: only Flash-Next got the ball into the cup. Both 27B models spent their whole \~111K-token answer budget thinking and never placed a part.
  • CAPTCHA (40 puzzles, local copy): 27B 24/40, Flash-Next 21/40, Uncensored 19/40.
  • Things you look at: Flash-Next made the best voxel castle, the best design board and the best video edit (19/20 on my rubric).

https://preview.redd.it/4e22yukvi4th1.png?width=1484&format=png&auto=…

Setup

  • GPU: one NVIDIA RTX PRO 6000 Blackwell, 96GB
  • CPU: AMD Ryzen 9 9950X
  • System RAM: 96GB DDR5
  • 27B and Uncensored: lmsysorg/sglang:v0.5.20, DFlash2 drafter, 262,144-token window, 4 slots
  • Flash-Next: lmsysorg/sglang:dev-qwen38-next-local, built-in MTP drafter, 262,144-token window, 1 slot. Its BFCL run used sglang:v0.5.20, like the 27Bs.
  • Sampler: the model card's thinking settings at the highest effort for the agent tests. BFCL and the needle test use the card's non-thinking settings.
  • The agent tests (SVG, video editing, voxel, design, Rube Goldberg, CAPTCHA) run inside Pi, a coding agent, with bash, read, write and edit. CAPTCHA gets only screenshots, mouse and keyboard.

One workstation, one model server at a time, and every number comes from a saved run.

1. Speed: drafters, SGLang vs vLLM, and long prompts

For the 27B I ran a speed matrix: every drafter, two engines, and four builds (three NVFP4 exports, one of them the Uncensored fine-tune, plus full-precision BF16). Each arm got its own server from a cold boot, the card's sampler and the 400W cap. "One user" is the Spec-Bench median over its 480 prompts. Engines: lmsysorg/sglang:v0.5.20 and vllm/vllm-openai:v0.29.0.

Which drafter (tok/s, one user):

|Drafter|SGLang · RadixArk NVFP4|vLLM · Inferact NVFP4|
|:-|:-|:-|
|none|75|59|
|MTP (built into the model)|160|113|
|DSpark|174|137|
|DFlash2|210|160|
|DFlash2 + torch.compile|214|not run|

DFlash2 wins on both engines. It keeps about 3.7 drafted tokens per step, against 2.9 for MTP.

Which build, on which engine (tok/s):

|Build · engine|DFlash2, 1 user|No drafter, 1 user|DFlash2, 4 users (total)|
|:-|:-|:-|:-|
|RadixArk NVFP4 · SGLang|210|75|607|
|Inferact NVFP4 · vLLM|160|59|517|
|Uncensored NVFP4 · SGLang|146|46|473|
|Uncensored NVFP4 · vLLM|169|63|538|
|BF16 · SGLang (full precision)|97|29|291|

  • The engine gap depends on the build. RadixArk's export on SGLang was the fastest arm overall, but on the one export I ran on both engines (the Uncensored), vLLM was 16% faster.
  • The NVFP4 exports aren't interchangeable. Same architecture, same 4 bits, same engine (SGLang), same drafter: RadixArk's export ran 210 tok/s and the Uncensored one 146.
  • 4-bit vs full precision: NVFP4 with DFlash2 is 2.2× the BF16 speed.
  • Prefill doesn't care about the engine: a full 245K-token window took 96–103s on every NVFP4 arm, SGLang or vLLM. BF16 took 129–135s.

https://preview.redd.it/vnd83h5zi4th1.png?width=1484&format=png&auto=…

https://preview.redd.it/0v1unh5zi4th1.png?width=1484&format=png&auto=…

The three models:

|Metric|Qwen3.8-27B|27B-Uncensored|Flash-Next|
|:-|:-|:-|:-|
|Prefill, full window|97s|99s|22.4s|
|Decode, Spec-Bench, one user|210 tok/s|146 tok/s|not run|
|Drafter vs no drafter|2.8×|3.2×|n/a|

The 27B keeps writing at 223 tok/s with a full 245K-token window behind it. Speculative decoding depends a lot on the content: maths ran at 339 tok/s, roleplay at 149. The pattern was the same for both 27B builds.

One important caveat. Flash-Next's speed test ran on 2026-09-12 at a 600W power cap. I later moved the card to 400W, and the 27B matrix ran at that cap. In my power sweep, prefill lost about 6% per 50W removed, so some of the gap is the cap. Moe also helps

https://preview.redd.it/jx28sde1j4th1.png?width=1484&format=png&auto=…

2. Tool use: the dense 27B leads

This is a 900-case BFCL v4 subset (11 categories), not the full leaderboard. Thinking was off, with temperature 0.7 and top\_p 0.8 from the card.

|Metric|Qwen3.8-27B|27B-Uncensored|Flash-Next|
|:-|:-|:-|:-|
|BFCL core|73.3%|70.8%|64.5%|
|Tool accuracy|87.8%|88.0%|82.4%|
|Abstention|79.5%|69.5%|68.5%|
|Multi-turn|52.5%|55.0%|42.5%|
|Malformed calls|0.08%|0.27%|0.28%|

These aren't comparable with my last post's Flash-Next BFCL numbers, which used temperature 0.

https://preview.redd.it/yqdc2j32j4th1.png?width=3396&format=png&auto=…

3. Long context: perfect for all three

I hid a fact in a log file that filled 33%, 66% or 99% of the 262K window, at three depths, with three needle types. The cache was flushed before every request.

  • 27B: 27/27
  • Uncensored: 27/27
  • Flash-Next: 81/81 (three samples per cell instead of one as I run this at the beginning)

The largest prompt was about 259.5K tokens.

https://preview.redd.it/fgeuq733j4th1.png?width=3396&format=png&auto=…

4. Battle arena: the local 27B beat Claude

Each model gets a rules sheet and a 1,000-point budget. In the open arena it designs one army blind and fights 13 armies: nine historical references plus the other entries. The score is 1,000 × its average win rate. Every matchup is 200 deterministic battles (100 seeds, sides swapped).

|Rank|Entry|Score|
|:-|:-|:-|
|1|RadixArk/Qwen3.8-27B-NVFP4|700|
|2|Claude Fable 5.1 (chat, max thinking)|678|
|3|Qwen3.8-27B-Uncensored|603|
|4|Qwen3.8-Flash-Next|535|
|5|GPT-5.6 (chat, ultra thinking)|473|

In the gauntlet, the model sees each enemy and builds a counter. The 27B beat 12/15, the Uncensored 12/15 and Flash-Next 13/17. Flash-Next ran an earlier version of the gauntlet with two more enemies, so treat that row as close, not ranked.

Thinking cost: the 27B's arena army took 39K thinking tokens in 4 minutes. Flash-Next's took 73K in 9 minutes.

https://preview.redd.it/1wbb6p28j4th1.png?width=3396&format=png&auto=…

https://preview.redd.it/o2yccq77j4th1.png?width=1484&format=png&auto=…

5. The SVG test is also a fact check

Prompt: find out which card local-AI hobbyists run and which current open model fits it, then draw the card lifting the model, labelled with a quant and a size that fit. All three picked the RTX 3090. I checked every label against what each session actually fetched.

  • 27B: 5/5 facts correct. Qwen3-Coder-30B-A3B at Q5\_K\_M, 21.73 GB, the real file size. Q6\_K at 25.09 GB is correctly marked as not fitting.
  • Flash-Next: 4/5. It got all four file sizes right and the exact 3.3B active parameters, but labelled the 3090 with a "12VHPWR, melted once" joke. That's the wrong card.
  • Uncensored: 3/5. It labelled the model "QWEN3.8-27B" but used the file size of Qwen3.6-27B Q4\_K\_M, and its "12 tok/s" isn't in anything it fetched.

All three passed 8/8 format checks. The 27B looked at its render twice and Flash-Next three times, where the rule allows one look.

https://preview.redd.it/oijy66u9j4th1.png?width=1920&format=png&auto=…

6. Video editing, voxel and design

Video editing: the model gets a raw 132-second take with fillers, a retake and a swear. It never sees the footage, only transcription and silence-detection tools, and then edits through FableCut's tools. There are two cases, each scored by hand out of 10:

  • Flash-Next 9 + 10 = 19
  • 27B 8 + 8 = 16
  • Uncensored 5 + 8 = 13

Voxel (Wawel Castle in three.js), ranked by eye:

  1. Flash-Next is the only one with the gold Sigismund Chapel dome and the Vistula bending around the hill.
  2. 27B built a clean but generic castle.
  3. Uncensored placed the camera inside its own build.

Flash-Next also used the fewest thinking tokens there: 73K, against 101K for the 27B.

Design (an animated explainer board in my design system), ranked by eye: Flash-Next first, and the two 27Bs shared second. All three passed 7/7 hard rules.

https://preview.redd.it/ue511ppaj4th1.png?width=1920&format=png&auto=…

7. Rube Goldberg: only one machine

The setup is a fixed level: a ball on a ledge, a cup on the floor and a wall in between. The model writes a parts list (no code), and a 2D physics engine runs it. It can run and look as often as it likes within 90 minutes. The score is automatic: does the ball itself end in the cup?

Flash-Next: yes, at 12.4s. It made 49 simulator runs and used 40 parts (35 of them moved). The ball travelled 1,672 px. It used 224K thinking tokens and compacted its context 8 times.

27B and Uncensored: no machine. Both spent about 111K tokens thinking in their first answer, reached the per-answer limit and stopped before writing a single part. Everyone got the same rules and one attempt. A rule that let a model continue after hitting the limit might change this, and I haven't tested that yet.

https://preview.redd.it/px8mlhhbj4th1.png?width=1280&format=png&auto=…

8. CAPTCHA: local models in a real browser

I used Open CaptchaWorld (20 CAPTCHA types, two of each). The model only sees screenshots and only acts with the mouse and keyboard. The site's own checker marks the first answer, and it must arrive within 7 minutes.

|Model|Solved|Median time|Thinking tokens, all 40|
|:-|:-|:-|:-|
|Qwen3.8-27B|24/40|55s|456K|
|Flash-Next|21/40|145s|1.0M|
|27B-Uncensored|19/40|36s|401K|

With its own 20-minute limit, Flash-Next solved 23/40. The paper reports 93.3% for humans and 40% for the best agent on its full set, which isn't the same 40 puzzles. With one run each, a three-puzzle gap is not a strong signal.

Which one should you run?

  • RadixArk/Qwen3.8-27B-NVFP4 for agents and tool calls. It won BFCL, the arena, the SVG fact check and CAPTCHA.
  • RadixArk/Qwen3.8-Flash-Next-NVFP4 for long prompts and building things, especially visual ones. It reads a full window much faster and won video editing, voxel, design and the Rube Goldberg machine.
  • orcarouter/Qwen3.8-27B-Uncensored-NVFP4 only if refusals are your actual problem. It won nothing here and invented facts in the SVG.

Resources

Configs, Docker setup and reports

show remaining 403 characters

The test harness is still private while it's changing.

Full video with the battles, the castles and the ball run: https://youtu.be/VOtfja\_Toj4**

I abused AI to help write this up and to check it against the report. Every number above comes from a saved run.

Which test do you find most interesting and maybe you have some other creative ideas how to test models?

💬 8 (+1) open on reddit ↗
▲
84
+5
41👁
r/LocalLLaMA · u/WebAssemblyMan · 8d ago
DeepSeek harness 0.2 - Optional Bundle Architecture, Windows Sandbox improvements, Async Question Mode, Desktop release, Web Search without key post image

Optional Bundle architecture: Schedule (session-local delayed / timed / interval reminders) was removed from the default set and made an explicit Optional Bundle. This cleanly separates “installed” from “enabled” and is the first systematic use of the Profile + Bundle model for official features.

• Windows Sandbox improvements: A new permission-diagnosis skill can detect common Access Denied causes and perform backed-up, recoverable permission fixes after user authorization, giving the Agent a reliable recovery path instead of blind retries.

• Async Question Mode (experimental): “Ask the user” is no longer a hard synchronous block. After a timeout the Agent can keep working while the user answers later, introducing asynchrony between interaction and execution.

• Model-layer polish: DeepSeek-account sessions can use Web Search without an extra API key; third-party model catalog updated (some old IDs removed); long model lists now support fuzzy search and keyboard navigation.

• Desktop release: Official Windows and macOS clients are out (Linux unsupported). Account login is supported, suggesting paid plans may be coming soon.

• Overall theme: Version 0.2 strengthens the Agent Runtime’s composability, recoverability, permission boundaries, and execution-state semantics — the practical foundations needed to move from a toy toward production use.

💬 17 (+3) open on reddit ↗
▲
10
+5
26👁
r/LocalLLaMA · u/GodComplecs · 7d ago
Should we plead opensource labs to still produce great non thinking (instruct) models?

The results are in, no thinking / instruct mode for new models degrade performance more than on old models such as 3.6 vs 3.8, where 3.6 takes the lead on several coding benches in instruct mode.

I would ask the labs to still nicely focus on instruct mode also still, there a probably gains to be had without the lengthy reasoning still, some of us still use models for everything and they do not need long reasoning traces. Agentic is fine and all but to start SACRIFICING performance for the "base" model which we are used to from early Llama days is not a good direction imo.

💬 19 (+3) open on reddit ↗
▲
58
+5
22👁
r/LocalLLaMA · u/jjusko20 · 8d ago
Update #2: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch post image

Last update: https://www.reddit.com/r/LocalLLaMA/comments/1wu9ksu/update\_yandexaliceai\_80ba3b\_fine\_tune\_progress/ \- basically, an instruct fine tune on the base model using a synthetic distilled data set. I've been posting regular updates so I imagine at least a few people have seen this.

Live stream: https://figure-bios-expect-cio.trycloudflare.com/

UPDATE: Finished train. hopefully some examples soon.

The initial train is finally almost done, after about 48 hours of humming. While the loss curve looks a little crazy, I've done some analysis (and some chatting with the LLMs) to understand that my average loss each epoch has been steadily decreasing (few reasons the loss curve looks wacky, vocabulary size, low to high token counts in epochs, etc) - but I'm pretty happy with what I'm seeing so far.

I'm post training the attention and the shared expert, and leaving the base experts frozen - this is a behavioral and logic fine tune that preserves the original yandex training data.

I plan on, within the next few days, releasing a few gguf quants of this, along with a llama.cpp patch for running it locally. I'm not sure how well the initial fine tune is going to work out - loss looks good but I'll have to do some evaluating. Either way, I plan on continuing training with reinforcement learning and an extended SFT set, as I have room and a ton of capacity left in my QLoRA adapter. I'll release this version as a public checkpoint anyways though (kinda like how deepseek did it) so people can play around with it and hopefully get excited for new checkpoints.

Cheers! Stay tuned, this is a pretty fun model size to play with, I'm excited to release the instruct version. I'll open source whatever you guys want out of this - I already open sourced the distillation engine (see SFTMill, it's been posted in here in the last few days) - but I also have a custom kernel for training this for V100s and a few other patches I can share (this training has been plugging away on 3, 32gb v100s - man it took a while to get that to work). Mandatory plug for my own goals: if you're hiring remote or in NYC for a dev or ml engineer, hit me up!

God I hope it writes the adapter when this is done I didn't audit that code well enough.

💬 18 (+1) open on reddit ↗
▲
40
+4
24👁
r/LocalLLaMA · u/Qual_ · 9d ago
Astrabox - Open source Arcade Game Generator post image

Hey everyone! I’ve been working on ASTRABOX: an arcade interface where you describe a game, the AI builds it, and you can ask for changes by voice while playing.

Each game gets its own visuals and gameplay, while a shared runtime handles controllers, scores, player joining, etc.

I built it around Codex, but the code is open source. I’d love to see someone adapt it to a local coding model, local STT/TTS, and a different harness.

It runs as a local web app—you don’t need a Raspberry Pi or an actual arcade cabinet, although that’s what inspired the project 🕹️

It’s still experimental, but feel free to customize it, change the little robot, the environnement, or everything.

https://github.com/Qualzz/astrabox

Curious what models and tools you’d use for a local version.

Edit: Clanker helped me with writing this message in english.

https://preview.redd.it/rru52ja4brsh1.png?width=2224&format=png&auto=…

https://preview.redd.it/fm3m19j2brsh1.png?width=1280&format=png&auto=…

💬 11 (-1) open on reddit ↗
▲
207
+3
26👁
r/LocalLLaMA · u/jacek2023 · 8d ago
Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp

now you can use MTP with Qwen Flash Next, time to switch from Qwen 3.8 27B?

(merged after 17h of development)

quants: https://huggingface.co/ggml-org/Qwen3.8-Flash-Next-GGUF

link to the previous discussion (I deleted the old post to avoid duplicates): https://www.reddit.com/r/LocalLLaMA/comments/1wur4lt/qwen\_flash\_next\_mtp\_work\_restarted/

▲
57
+3
54👁
r/LocalLLaMA · u/Effective-Ad2060 · 8d ago
We benchmarked 18 RAG pipelines against an agent loop on Google's FRAMES. The best pipeline hit 78.9%. The agent loop hit 92.7%.

Hybrid search, reranking, query decomposition, and query expansion are often treated as must-haves for good RAG. We wanted to see how much each actually helped, so we tested them. Same model, same embeddings, same documents, across all 824 multi-hop questions in FRAMES.

We built 18 pipeline variants. The best one scored 78.9%. Our agent loop (with retrieval tools) that could read the results and search again scored 92.7%—roughly the same as giving the model the right articles upfront.

The reranker results might surprise you. A small reranker dropped our best pipeline’s accuracy by 9 percentage points, while a larger one barely helped. I’d already suspected reranking wouldn’t help much here, but wanted to test that assumption.

Another thing we noticed: models sometimes fill in gaps from memory, even when you explicitly tell them to stick to the retrieved documents. Those answers can still be full of citations. We ended up checking every correct answer against what the system had actually read.

Here’s the write-up if you’re interested:
Agentic RAG vs. traditional RAG on FRAMES

Full disclosure: I work on PipesHub, which is open source. The benchmark code and runbook are in the repo: https://github.com/pipeshub-ai/pipeshub-ai/tree/frames

Quick note on what the numbers mean: they're end-to-end answer accuracy, not retrieval scores. Every answer was graded by an LLM judge (Claude Sonnet 5) using the FRAMES paper's own grading prompt, and independently by a second judge (Gemini Flash 3.8). The two agreed on almost every answer (Cohen's κ 0.93–0.98). We also checked each correct answer against the text the system was actually shown, so answers that came from the model's memory don't count as retrieval wins.

💬 54 (+4) open on reddit ↗
▲
10
+3
18👁
r/LocalLLaMA · u/Designer_Elephant227 · 7d ago
Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal.

What i run:

\- model: Qwen3.8-Flash-Next from turboderp, exl3 5.05 bpw (head 6 bit, vision 6 bit, mtp 5 bit)

\- backend: exllamav3 rocm fork from phoenixhaxor (commit cbbef08), i had to patch one file for gfx12

\- cpu/ram: Ryzen 9 7945HX3D with about 90gb ram

\- 112 of 512 experts per layer are on the gpu, the other 400 on cpu with 16 threads

\- 262144 context, q8 kv cache, chunk size 4096, batch 1

\- mtp drafting is on, acceptance is around 53-54%

\- the ngram table (102gb, bf16 not quantized) gets streamed from nvme

Speed at 230k context (prose): prefill 863 t/s (only the new 101k tokens, the rest came from the prefix cache) and decode 34.7 t/s.

Is this ok for the r9700 or can i still tune something? Thanks 🙂

💬 26 (+5) open on reddit ↗
▲
16
+3
26👁
r/LocalLLaMA · u/brainchillzZ · 8d ago
Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

So everyone has been yelling about how I should be using Gufo instead of halogen because it's open source and it's "just as good or better". Checking in on their GitHub (GitHub.com/gufo-org/gufo) got me immediately .. "Qwen 27B Q4: 70.56 tok/s single user, 123 tok/s with 8 users" on a strix halo device? Yes please ... So I broke down and tried it today ...

Setup: gufo 0.4.0 from their podman image, Qwen3.8 27B UD-Q4\_K\_XL from Unsloth plus the DFlash2 Q4\_K\_M draft model, using their own benchmark script and their own settings (greedy, thinking off, 128 output tokens, prompt cache off).

If you want the short version ... yeah I got 70.22 tok/s. So the number is real. But the prompt that produces it is "Write the word red exactly 1000 times".

But it's also not real. In that figure all the speed comes from the speculative decoding. The draft model guesses like 7 tokens ahead, the 27b checks them in one pass and keeps what it agrees with. When the output is the same word over and over the draft is right every time. On a real prompt it's right maybe half the time.

Their benchmark has a second set of nine ordinary prompts (some C++, a word problem, a summary, Italian, Chinese, JSON, a bit of fiction, a debugging checklist).

On those:
| | repeat-a-word prompt | normal prompts |

|---|---|---|

| 1 user | 70.2 tok/s | 39.4 tok/s median, anywhere from 22 to 52 depending on the prompt |

| 8 users, their "aggregated" number | 122.6 | 82.4 |

| 8 users, tokens actually delivered per second | 82 | 52 |

About that last row. The "123 tok/s aggregated" figure is each request's decode speed added together, with prompt processing and queue time left out. If you just count tokens coming out of the box per second of wall clock it's 82, or 52 on normal. prompts.

To be fair to the gufo people, none of this is hidden. Their benchmark docs have separate "mixed" and "repetitive" columns and the mixed numbers they publish match what I got. It's only the repo description and the top of the README that lead with the best case. And 39 tok/s from a 27B at Q4 on an APU is still really good. Without the draft model their docs put it around 12.

The other thing I wanted to know was how it compares to halogen (peonist-ai/halogen-flash-server), which is what I normally run. Both can serve Qwen3.8 Flash-Next, so I put that on both and sent the same prompts to each. Two boxes, same hardware, same OS image. Greedy, thinking off, 256 tokens.

| | halogen 0.13.8 | gufo 0.4.0 |

|---|---|---|

| nine normal prompts, average decode | 43.9 tok/s | 38.2 tok/s |

| 4 users at once, end to end | 76.6 tok/s | 63.1 tok/s |

| cold prompt processing, \~9.7k tokens | 1288 tok/s | 1495 tok/s |

| the "red" prompt | 56.9 tok/s | 87.4 tok/s |

So for everyday generation halogen was about 13% faster for one user and about 18% faster with four. gufo was 16% faster at chewing through a long prompt and a lot faster on the repetitive one.

I'll add this just in case, because someone will ask or at least try to poke about it in the comments

\- I know the weights aren't the same. halogen uses its own 4-bit format, gufo uses the Unsloth GGUF. I only measured speed. I did not compare output quality at all.

\- They were two different machines but identical hardware and software, and my boxes have agreed within 1% on other benchmarks, but it's still two machines.

\- One run each was done for the head to head. The reproduction of their numbers was 3 reps.

\- gufo has shipped four releases over the last four day, so this could all be stale by next week.

There was quite a lot of stuff that I liked about gufo that isn't performance related. It takes plain GGUFs, it's MIT, the 27B loads in about 3 seconds (Flash-Next in 13), the per-request log line tells you draft acceptance and cache hits, and it does 8 batched sessions. It also has ASR, TTS and image models that I haven't touched. Their benchmark hashes the output with and without the draft model and it was identical every time, so the speculative path isn't changing what the model says.

One thing to keep in mind if you try it out is that it reserves memory per session up front. Flash-Next with 4 sessions at 64k context took 94GB.

So it isn't smoke and mirrors exactly. Everything I checked reproduced. Just know that the 70 is a ceiling you'll only hit if your workload is incredibly predictable text, and plan around the 30s for the 27B on normal stuff.

I kept this all setup to tinker with on actual output quality over the next few days, I'm happy to run other prompts or try it with different settings if anyone wants to see something specific.

▲
14
+3
15👁
r/LocalLLaMA · u/KissMyShinyArse · 8d ago
Strata: how to configure sampling parameters

The top-level README doesn't mention this, but you can add a "sampling" key to your strata-iq3_s.json like this:

{
"sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20},

"exe": "/path/to/Strata/engine/strata",
"args": [ ... ],
...
}

From docs/DETAILS.md:

The run config's optional sampling block sets the defaults for requests that leave the fields out ("sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20}); a request's own fields always win, and with no block at all a request without sampling keys decodes greedy.
💬 10 (+1) open on reddit ↗
▲
5
+2
17👁
r/LocalLLaMA · u/Fz1zz · 7d ago
QFN at 262K on 32 GB RAM 48GB VRAM : 1,800 tok/s prompt, 130 tok/s decod three Strata patches.

Qwen3.8-Flash-Next (IQ3\_XXS, 76 GB) at the full 262K context on 31 GB of RAM: 1,800 tok/s prompt, 130 tok/s decode, on a 5090 + a 4070 Ti SUPER on a PCIe x1 slot

Strata (github.com/Niko1221/Strata) streams MoE experts from an mmap'd GGUF, and its own sizing rule says RAM >= expert shard + 10 GB, so 57 GB for the ISTA GSQ-RCO IQ3\_XXS. I run it on 31 GB and it is fast now. Hardware: RTX 5090 32 GB + RTX 4070 Ti SUPER 16 GB (the small card sits on a chipset x1 slot, 0.8 GB/s), i7-14700K, one NVMe.

Stock 0.1.33 with a layer split at 36: 80K-token prompt 890 tok/s with the 4070 at 100% and the 5090 idle, decode \~110 tok/s, and switching between two chats re-reads the other one (30K tokens = 48 s).

Three patches on v0.1.33 (repo below, they apply to a pristine checkout):

  1. Prompts run entirely on the big card (port of Strata PR #269), the small card gets its layers' state copied afterwards, only the cells in use. Conversation parking works with the split, so alternating between a phone and a desktop session takes 0.6-1.2 s instead of 20-48 s.
  1. The real bottleneck on a box with less RAM than the expert file: the expert pool copied each 2 MB expert out of the mmap with memcpy after a MADV\_WILLNEED hint. Under memory pressure the kernel drops that readahead and you pay one major page fault per 4 KB, hundreds per expert, while both GPUs wait. A pread per slice instead: major faults per benchmark run went from 24 million to 14 thousand.
  1. Bigger prompt chunks (--prefill auto:32768). Every chunk re-streams every layer's experts, so four times fewer chunks matters a lot when the working set does not fit in the page cache.

Numbers on the final config (int8 KV, 32K cells resident per layer, MTP spec 4, vision on, 262,144 context):

\- 80K fresh prompt: 1,796 tok/s (45 s); 16K: 985; 2K: \~350

\- follow-up turn on an 80K conversation: 6 s

\- decode: 128-134 tok/s median on real sampling (0.6 / 0.95 / 20) with --spec-min-p 0.8, 95-108 greedy

\- 40K needle + follow-ups and a two-conversation parking test all correct

\- VRAM: 5090 at 32.0 GB, 4070 at 15.7 GB; RAM: the engine \~5 GB, the rest page cache

Repo with the patches, install script, launcher, benchmark tools and all measurements: https://github.com/ExTV/strata-5090-4070

▲
5
+2
20👁
r/LocalLLaMA · u/giveen · 7d ago
GitHub - giveen/KernelOPT: Dispatch-aware agentic GPU kernel optimization

I want to share something I've been working on, and research paper from Redhat really helped.

This is my agentic GPU kernel optimizer for inference engine development

Its whole goal is to look at inference engine kernels, and using cloud models, plan, test and find improvements. Nothing is changed in your git, it provides a diff, send the diff over to your coding harness and ask it to analyze the diff.

It has a setup wizard and a run wizard which I recommend using, please give me feedback on what needs to be improved, as the AMD stuff I was unable to test, and mostly is theory at this time.

▲
2
+2
33👁
r/LocalLLaMA · u/Cyb3erDudu · 7d ago
shardr — like docker for models, BitTorrent sync, OpenAI-compatible serving, inference engines from upstream. post image

I kept running into the same three problems: the same 40 GB quant downloaded twice because it lived in some folder I forgot about, models quietly disappearing from Hugging Face, and every tool keeping its own copy of the weights on disk. So I've been building shardr \- a small Go daemon (Apache-2.0) that gives your machine one content-addressed store for models.

What it does, concretely:

  • Everything is stored and verified by SHA-256. Trust comes from digests, never from where the bytes came from.
  • It speaks BitTorrent v2. You can pull models from peers, and it seeds whatever you hold back into the swarm.
  • The runner starts llama-server with an OpenAI-compatible API and mmaps the weights directly out of the store — one copy on disk, no staging copies, no moving files around.
  • Runtime is pinned via a lockfile to upstream llama.cpp release binaries (never self-compiled, digest-verified), so a llama.cpp update is just a PR against that lockfile with the full test matrix behind it.

It works with pirateface.co as a catalog: shardr catalog search qwen, shardr pull <owner/repo>, and the download is anchored against the Hugging Face checksums for that exact revision — the magnet can't lie to you. Rescued models (HF source gone) pull against the catalog's recorded checksums, but only if you explicitly opt in with --trust-catalog. Other mirrors fit behind the same interface — the catalog is pluggable and the base URL is configurable.

Quick taste:

$ make all

$ shardhive serve &

$ shardr catalog search qwen2.5-0.5b

$ shardr pull unsloth/Qwen2.5-0.5B-Instruct-GGUF --quant q4_k_m

$ shardr serve unsloth/qwen2.5-0.5b:q4_k_m --id chat

$ curl http://127.0.0.1:<port>/v1/chat/completions -d '{"model":"chat",...}'

Where it stands: end-to-end works on macOS arm64 and Linux amd64, releases build themselves from CI, docs at https://cyb3rdudu.github.io/shardr. What it doesn't have: a UI, Windows support, runtimes beyond llama-server, and honestly, probably a bunch of rough edges.

What I'd like help with:

  • people with large local collections to try imports and tell me what breaks
  • feedback on the trust model — HF-anchored pulls, the explicit opt-in for rescued models. I'm sure there are holes; poke at them
  • anyone who enjoys the swarm/seeding side and wants to hack on it

Happy to answer anything about the design decisions. Docs are linked above, specs are in the repo if you want to see how the sausage is made.

💬 4 (+2) open on reddit ↗
▲
10
+2
11👁
r/LocalLLaMA · u/ParvusNumero · 7d ago
Mirostat?

Reading another post made me think:
Is anybody still using Mirostat?

It was all the rage and people said it avoided the “boredom trap” for long texts.

Do newer generation models not need that anymore, or are other samplers superior?

💬 24 (+3) open on reddit ↗
▲
4
+2
16👁
r/LocalLLaMA · u/failuremap-f · 8d ago
Can your local coding model repair these boundary-case bugs? Failure Map: 20,168 open Python tasks

I’m the creator of Failure Map, an archive of compact Python debugging tasks. The open release has 20,168 tasks across 254 categories, with standard-library implementations, explicit contracts, failed repair attempts, and executable boundary checks.

Three small cases to try:

• Duplicate delivery: deduplicating equal amounts loses legitimate events. https://failuremap.org/cases/FA-001

• Cache expiry: subtracting a whole tick rejects an entry that is still valid. https://failuremap.org/cases/FA-006

• Pagination: changing > to >= repeats the cursor record. https://failuremap.org/cases/FA-011

Prompt template: “Repair the solve function to satisfy the stated contract. Return Python source only. Preserve the signature. Contract: {prompt}. Broken implementation: {broken\_source}.”

Measured program baselines, passed checks out of 3 (broken / attempted repair): FA-001 2/3 / 1/3; FA-006 2/3 / 2/3; FA-011 2/3 / 1/3. These are executions of the included programs, not model scores. I have no measured local-model results to claim yet.

To compare runs, report the exact model and revision, quantization, prompt, sampling settings, seed, attempts per task, and pass counts. Run candidate code in isolation and keep grading fixtures outside its control. Recorded-check success is not hidden-test performance.

Download: https://failuremap.org/api/exports/tasks.jsonl.gz

Methodology: https://failuremap.org/methodology

▲
16
+2
17👁
r/LocalLLaMA · u/Federal-Effective879 · 9d ago
szmcp: a ZIM HTML to Markdown converter and yet another ZIM MCP server

Hello all, I wanted to share a little project I vibe-coded for myself that you may find useful.

As many people here like to suggest, I wanted to give my small local LLMs access to information to improve their world knowledge. I didn't want to give my LLM free reign searching and browsing the web to keep my queries private and functional offline, so I wanted to give them an offline knowledge base. Wikipedia ZIM files from Kiwix were a good starting place for this. Several MCP servers for ZIM files exist, but I didn't like the existing ones I found for various reasons. The most notable one is openzim-mcp , which works in its advanced tool mode but has overly complicated context-bloating tools, and whose simple single tool mode doesn't work very well in practice.

I built my own MCP server for ZIM files in Rust, exposing a simple tool set that's actually easy for small local LLMs to use, while providing all the functionality one normally needs. It's designed mainly for Kiwix MediaWiki ZIM archives generated by mwoffliner (such as Wikipedia, WIkivoyage, etc.) but also usable with many non-wiki ZIM files. I also wrote my own custom HTML to Markdown converter for MediaWiki pages that produces clean, well-formatted Markdown including special content such as wiki infoboxes, LaTeX formulas, tables, etc. It also strips out references and boilerplate sections from wiki pages to keep the resulting markdown clean and context efficient.

You can hook this MCP server to llama.cpp's Web UI to give your small local LLMs much better world knowledge. A system prompt that I found works well is:

You are a helpful assistant. When answering factual queries, search through Wikipedia using the provided ZIM access to ground your answers. If the articles or sections you read don't have relevant details, you can search more, but don't keep searching forever; you need to answer reasonably quickly.

I tested it with various LLMs of varying sizes. I got good results with Gemma 4 12B (or bigger), IBM Granite 4.2 8B (or bigger), and Ling 3.0 Flash (best results while still maintaining usable speed on my 128 GB Mac). Qwen 3.6 35B-A3B was usable but tended to overthink and hallucinate; Qwen 3.8 27B was too slow to be usable for this purpose on my Mac. I also experimented with smaller models, and got usable results for simpler queries with MiniCPM5 2B, LFM 2.5 2.6B, and IBM Granite 4.2 3B. Gemma 4 E4B did not work well for this.

I also build a sub-command within this tool to convert entire ZIM files from HTML to Markdown to save disk space (and avoid the need to convert on every tool call). It converts a 49 GB Kiwix nopic full English Wikipedia ZIM file into a 19 GB Markdown ZIM file, while maintaining all article content (aside from references) and maintaining full-text search. Likewise, it converts the 17 GB top-1M nopic enwiki Kiwix ZIM file to 6 GB. You can make the resulting ZIM files even smaller if you specify the option to only index article intros for full-text search (since the full-text search Xapian index is a large fraction of the file size). The converter is multi-threaded and written fairly efficiently using Rust, so you can convert all the millions of articles in a full English Wikipedia Kiwix files in a few hours on a typical modern computer.

GitHub link: https://github.com/sultanqasim/szmcp

▲
119
+1
33👁
r/LocalLLaMA · u/Usual_Maximum7673 · 8d ago
Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory

A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO 6000, but as a general zero-shot classifier it trailed the big models.

Then it occurred to me that most decisions an agent makes in front of a local model aren't open-ended. They fall into a handful of recurring kinds: is this a prompt injection, which tool to call, how urgent is this ticket, is this answer grounded in the sources. So I trained 9 LoRA adapters, one per job, and you pick the ones you need. The server loads the base once plus whichever adapters you choose (about 40 MB each), and every request either names an adapter or goes to plain Jeff.

That means you keep both: the base model stays untouched, so you still get Jeff's general zero-shot ability for anything new, and the adapters give you near-perfect accuracy in the domains you care about. Each adapter was also trained with 10% of the base model's own training data mixed in, to help it keep its general skills.

Everything is on jeffhub.ai: the adapters, the results, the docs. Code on GitHub, models on Hugging Face, and you can try all nine adapters in your browser.

The headline: I let Jeff + adapters answer first and pass only the queries it's unsure about to Qwen3.8-27B. Same test rows both ways, on an M4 Max:

|Measure|Qwen3.8-27B alone|Jeff + adapters, 27B only when unsure|
|:-|:-|:-|
|Accuracy (mean of 8 adapters\*)|86.6%|95.3%|
|Time per decision (mean)|8.1 s|0.25 s (38× faster)|
|Wrong answers|13.4%|4.7%|
|Memory|28.6 GB|under 2 GB for Jeff, even with all 9 adapters loaded (+6.9%)|

On the five decisions an inbox agent makes for every message (guard, triage, support intent, tool choice, grounding) alone: 87.7% → 95.7%, 39× faster. Jeff wins outright on 8 of the nine adapters and ties on grounding (96.3% vs 96.7%, at 20× the speed). On their full held-out test sets, six of the nine adapters score 97–98%. On a GPU, a decision takes about 30 ms, whether you load one adapter or all nine.

\*Emotion is left out of the averages: picking the single strongest of 27 emotions (or neutral) in short Reddit comments is hard even for people, and the human labels often disagree. Jeff + adapter scores 60.6% there against the 27B's 35.6%, at 42× the speed. Including it, the average across all nine adapters is 91.4% for Jeff + adapters against 80.9% for the 27B, so leaving it out makes the gain shown above smaller, not larger.

Caveats, up front:

  • the 27B ran in 8-bit with step-by-step reasoning off (with reasoning on, the speedup would be even more dramatic);
  • each task used a fixed sample of 300 held-out rows (500 for emotion and legal-clauses);
  • each adapter's "pass it on" threshold was chosen on separate calibration rows, before the test rows were scored.

Data: 4 adapters are trained on public data sets. 5 are mostly synthetic. Every generated row records which model wrote it, and the cards give the counts. Every data set went through a shortcut check and an independent review before training, and a lot of first drafts failed: things like the answer being given away by length.

What's open: weights (Apache 2.0), code (MIT), and each adapter's test and calibration sets, so you can check every number. The training data isn't published.

This is a community preview: I'd love feedback.

Next: over the next \~36 hours I'll train v1.3, a long-term-support base. The fixed parts of a prompt come first, so servers can prepare them once and reuse them, which means faster decisions. I'll then retrain all nine adapters on it and keep the request format stable, so others can build and submit their own adapters. The adapter kit, with the data checks I used, is in the repo.

I've got access to more hardware now, so if there's a decision you'd like an adapter for, tell me and I'll train it.

The goal: when the next generation of local models lands (like everyone, I'm watching for Qwen 4), anyone running one locally should also have a tiny, fast, well-calibrated decision layer in front of it.

💬 33 (+1) open on reddit ↗
▲
2
+1
5👁
r/LocalLLaMA · u/itsokimjudgingyou · 7d ago
LINKUP AI sever PCIE 6.0x16 Cables

Has anyone tried the PCIE 6.0 AI server cables made by LINKUP? I don't need 6.0 but the cable routing these cables could offer me is huge compared to the normal risers. The lack of reviews is really the only thing holding me back.

Does anyone have experience with them?

💬 8 (+6) open on reddit ↗
▲
12
+1
34👁
r/LocalLLaMA · u/klieret · 7d ago
New benchmark on LMs fixing bugs before users run into them

Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW.

Most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But shouldn't we expect models by now to also find bugs before anyone runs into them?

So in SWE-sweep we just hand an agent a big codebase and ask it to find & fix as many bugs as it can. We then give a score based on a hidden set of bugs that we know about in the repos. All the bugs are real-world bugs. We do a lot of filtering to make sure the bugs are actually discoverable & fixable from reading the repo alone.

https://preview.redd.it/irffy7x5v2th1.png?width=1080&format=png&auto=…

We're still expanding the leaderboard list with more local models (unfortunately it's always a big harder with funding/infra etc), but right now it seems like it's quite hard to beat Luna xhigh in terms of cost efficiency.

Also the scores are way lower than I would've expected. Some tasks are legitimately superhuman in practice (like fixing up all of numpy), but there's also lots of small repos, where I would've expected a lot more from current models.

Everything is open source (MIT license) on github and you can find paper etc. on the website.

Happy to answer questions here, also super curious what open weights models you'd recommend running next (we're working on an update next week).

💬 14 (+2) open on reddit ↗
▲
2
+1
4👁
r/LocalLLaMA · u/OlegDoDo · 7d ago
SAGG — turning unreliable Gonka brokers into a reliable inference API (cascading failover, real data)

If you've used Gonka inference directly, you've probably noticed individual brokers aren't always consistent — one might be fast and reliable for a while, then slow down or drop requests, then recover. That's just how a decentralized network of independent nodes behaves.

SAGG takes a different approach: instead of relying on one broker and hoping it stays healthy, it holds several at once and automatically routes around whichever one is struggling at that moment. From the outside, you just get a normal, reliable API — the instability gets absorbed before it ever reaches you.

We didn't just build this and claim it works — we measured it properly, on real, sustained production traffic, two separate campaigns:

September 18 (1000 requests/line, standard prompt mix):

\- Standard line: 100% success (1000/1000 requests)

\- Super Deal line: 98.9% success (989/1000 requests)

September 30 recheck (200 requests/line, heavier prompt mix - longer context, code generation):

\- Standard line: 99.5% success (199/200 requests)

\- Super Deal line: 99.0% success (198/200 requests)

TTFT p50: \~190-490ms depending on line and load, p95 under 15s on heavier workloads.

Full methodology, raw data, and a reproduction script: github.com/privatedeskai/sagg-benchmark-data

For the technically curious: the hard part wasn't picking a backup broker — it was streaming responses specifically. Once a provider starts sending content to the client, you can't silently switch mid-stream without

▲
1
+1
8👁
r/LocalLLaMA · u/Away_Interaction6630 · 7d ago
How do you keep a local multi-agent app usable on CPU-only / low-RAM machines?

Hi everyone,

We're three final-year students at Epitech building Horus, a multi-agent assistant that runs entirely locally and offline. Our current challenge is hardware: keeping it usable on machines without a powerful GPU, without long setup times, excessive RAM/VRAM use or crashes.

Where we are today, from our last beta test:

One tester needed over 2 hours to install. The Python dependencies alone take 21–46 min, and the download is about 25 GB.

On CPU only, routing a question can take around 40 s, long enough for our WebSocket connection to drop.

\[Models we use + the smallest machine we've tested on\]

We'd love advice from anyone experienced with:

\- CPU-only LLM inference and memory-efficient loading

\- Quantization and model choice for low-end hardware

\- GPU/CPU fallback strategies

\- Hardware detection and adaptive configuration

\- Preventing resource exhaustion during setup and execution

Advice in the comments is very welcome, with no strings attached.

Looking for contributors: we also have a few small, well-scoped tasks or code reviews (about 1–4 hours), for example \[reviewing our hardware detection and model selection, or benchmarking a quantized model on a 16 GB RAM laptop\].

To be transparent: Horus is closed source and will be licensed to companies. Contributing is voluntary and unpaid. Before seeing any code, contributors sign a short confidentiality and contributor agreement, and the code they contribute becomes part of Horus. In return we offer thorough code reviews, full credit in the project and a professional reference on request.

We're not sharing code or private links publicly. If this interests you, comment below or DM me with your experience in local inference, CPU optimisation or offline apps

💬 25 (+12) open on reddit ↗
▲
3
+1
15👁
r/LocalLLaMA · u/Thac0-is-life · 7d ago
Help me with a better hardware setup for Local LLm

&#x200B;

Hello. I've been playing around with local LLMs for a while now, using my 7900xtx. I understand the concepts and usually what to do. But I'm now in a bit of choice paralysis on where to go next.

I have a Ryzen 5600X with a 7900XTX and 64GB of DDR 4 (how I wish I had purchased more at the time..) and a motherboard (MS-7B79/X470 GAMING PRO (MS-7B79)) that is not really great for multiple GPUs (it was a gaming PC).

I would like to increase my Local LLM game. I'm running mostly qwen 3.8 27B at 3 or 4q from with 128k to 220k context with KV cache at 8q. Sometimes I get up to 50tk/s and around 750 tk/s of context ingestions. But I wanted to run bigger models/have faster speed, or at least run multiple copies of that same qwen so multiple agents can run at the same time. Or try the Qwen 3.8 Flash for example. This level of model is already awesome enough to do anything I need.

I've been thinking of purchasing 2 or 4 MI50 16GB( which costs 1/3 of the 32GB), but I don't really know the rest that I should get. Motherboards that would help me optimize the performance around that, etc.

Or should I just bite the bullet on another 7900xtx (more expensive than 4 MI50)? But I think I would still need a new motherboard at least to let me use both at the same time

I have a basement so noise is not a problem, and I have solar, so power is not a real issue (at least during summer).

What are you all suggestions here? Does it make sense to go with older GPUs like that?

▲
5
+1
12👁
r/LocalLLaMA · u/No_Contract_8296 · 7d ago
CalDec v1 - Fully Open Decision Model for Personal Assistants

Somebody just released a fully open-source, open-weights decision model that beats Jev!

Just kidding, it's me and I this is my first time releasing a public model, recipe and dataset so I am looking forward to learning from the experience.

Jev is indeed a very powerful and inexpensive model and obviously a much better all-rounder, and some of my checkpoints did in fact score better on some tests (namely LocalLLaMA/typed-decisions and the internal test set) but that doesn't mean it "beats Jev" of course.

The motivation for this was a quick experiment to see how far behind Jev open-weights models like Laya are, and how much closer I can bring them with a small dataset and fine-tuning. The results were better than expected especially for me since I do not have professional ML experience.

For my use-case - a Jarvis-like personal assistant which aims to be real-time and fully-local - this model proved to be genuinely useful for certain aspects of that project so I decided to share the results and how I got there. Going local also means privacy and eliminating network latency.

I hope some of you find this experiment valuable or useful in some way.

I would also love to hear you suggestions, criticism or just discuss the approach!

Dataset: https://huggingface.co/datasets/kgrozdanovski/assistant-decisions**
CalDec Laya: https://huggingface.co/kgrozdanovski/caldec-v1-laya**
CalDec GLiNER: https://huggingface.co/kgrozdanovski/caldec-v1-gliner2.5-decide**
GitHub: https://github.com/kgrozdanovski/caldec**

▲
2
+1
4👁
r/LocalLLaMA · u/Real_MakinThings · 8d ago
How to know about optimized engines

Optimizing an engine for a family of models and hardware combination seems very appealing. As someone who uses qwen3.6 and 3.8 a lot, and is evaluating hardware options before going fully local, it's hard to keep up with the state of things.

Huggingface made it possible to see the development branches and derivative modifications to models. Is there something similar for inference engines yet? I've seen some where the it's optimized for a shell game of moving layers between vram and ram while using ngrams (amazing), others are all about quants (less amazing), but it's incredibly difficult to compare apples to apples where there's variability on card architecture, vram size, quant approach, memory management optimization approach... I was already busy over thinking my vram selection, now it's an even bigger decision matrix without any filters!

▲
1
+1
12👁
r/LocalLLaMA · u/TheRealJesus2 · 8d ago
Dwarfstar quants

anyone try these out? https://dwarfstar.sh

they have very clever quant techniques, bespoke for a handful of models running on their software. i got qwen 3.8 next running on m3 ultra 96GB studio and its fast and seems good so far. with memory headroom for other stuff

kinda blown away to be honest. want to know if others tried this yet and what the experience has been like for you.

💬 15 (+2) open on reddit ↗
▲
22
+1
12👁
r/LocalLLaMA · u/norenEnmotalen · 8d ago
Unsloth, Swift1.5, Peculiar-Ragdoll, ThinkingCap - Qwen3.8-27B

In a previous post I shared comparison between Swift1.5 and peculiar-ragdoll's checkpoints. Added the original unsloth Q4\_K\_XL and ThinkingCap Q4\_K\_M (they don't offer L or XL) to the comparison. Here are the results over a 69 set of eval questions.

All tests are now run at same "medium" reasoning effort.

unsloth-ud\_q4\_k\_xl one ran using llama.cpp - not the splash forked inference engine.

https://preview.redd.it/tztb66wygvsh1.png?width=2958&format=png&auto=…

I'll do a 3x repeat for the slow run to see if it maintains 69/69 each time.

EDIT: u/jucabala457 asked I test mradermacher/Signal-3.8-27B-Terse-Coder-i1-GGUF The Q4\_K\_M is closest quant available. A nice addition for sure! That GGUF couldn't run with Splash-based engine due to tensor incompat. I ran it using llama.cpp the slow way. The total time taken isn't a fair comparison for that reason. Updated results below

https://preview.redd.it/pvfvgd35twsh1.png?width=2976&format=png&auto=…

I also just made the tuieval tool available here https://github.com/ashe-wb/tuieval

Can't promise you the tool will work right away on your install since a fully local binary is what I've been using and testing with. Customize it with packs of domain-specific eval questions you deal with on the daily. This is the most important part. A model or fine-tune that is not good for one thing might be excellent for something else and only you know what your domain interests are. The ability of a model to render game graphics means nothing to me but it means everything to someone else.

https://preview.redd.it/m2n4rzlzjwsh1.png?width=2000&format=png&auto=…

▲
5
+1
19👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 8d ago
Computer use powered by local/cloud models for regulated industries?

The local models seem powerful enough to be capable of running basic local computer use. This is a computer use agent/harness built via claude code, and powered by qwen3.8 flash next nvfp4. Drawing a simple image of its choice took 12m 1s, but a lot of that time was spent by the agent trying to figure out a WebGL bug. For what it did, it seems relatively fast. Prefill speed \~1000 tps, gen speed \~50 tps. I went to openai’s devday a couple days ago, and it seems like cloud models that are very smart and fast, like “astra ultrafast”, can perform work even quicker, for a premium.

Does anyone have experience with computer use agents/harnesses that are open-source and plug-n-play, that are robust enough to be used in regulated fields such as law/medicine? How do people deal with regulations, such as making such workflows HIPAA compliant in medicine? Experiences with helping users ensure that workflows are completed accurately? And whether they go with local or cloud models to power computer use?

💬 7 (+1) open on reddit ↗
▲
4
+1
11👁
r/LocalLLaMA · u/SignatureMoney6648 · 8d ago
FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.

I've run the benchmark on a RTX 3090, 1024 tokens in / 256 out, concurrency 1–32.

If the model fits on vRAM (Gemma-4-26B-A4B, byte-identical GGUF on both engines): llama.cpp has 2.2–3.2× the throughput and 5–6× faster TTFT. FreeToken 0.1.2 can't keep 4-bit experts in VRAM at all, and it OOM'd at 8 concurrent.

If the model doesn't fit (gpt-oss-120b, 63 GB): FreeToken's TTFT stays at \~9 s from 2 to 8 users while llama.cpp's goes 17 → 58 s. At 32 users it's 19 s vs 139 s. Throughput is basically a tie (10–17 tok/s for both).

Spilling to system RAM costs \~10× in generation speed whichever engine you use.

FreeToken's PCIe link sits at its ceiling the whole time, so PCIe 4.0 should help it a lot (I've run this on a gen3 motherboard).

So from this test FreeToken only makes sense with many concurrent users in models that cannot be hold inside vRAM. But I am not sure if that is always the case or an artifact of the gen3 bottleneck on my PC.

Has anyone run a benchmark like that with a gen4 Motherboard?

Full details on the link. BTW: I used AI to generate the charts and correct my spelling and grammar.

▲
16
+1
14👁
r/LocalLLaMA · u/caenum · 8d ago
Best OpenSource Claude Cowork alternative?

Hey guys,

Looking for an alternative for Claude Cowork:

  • Project Work / Documents
  • Integrations like Notion, Gmail, etc.
  • Tools like Websearch, PDF creation, etc.

Came over Eigent (https://github.com/eigent-ai/eigent) but cant find any actual reviews about it, what usually is a sign thats not good performing..

Also have tried multiple other frameworks (OpenClaw, Hermes, OpenWebUI Chat Interface) - but those are different use-cases for me.

LLMs will be server through my own server, so should be open for connecting to Ollama, Ninfer, etc.

So anyone knows a good application which behaves like Claude's Cowork?

Thanks )

▲
10
+1
11👁
r/LocalLLaMA · u/Biomass23 · 9d ago
tp=6 can work on vLLM, with padding

vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).

I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.

So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.

I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.

GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x

▲
0
 
8👁
r/LocalLLaMA · u/knighty1981 · 7d ago
2x 3090 in server chassis, upgrade time

I've got 2x 3090 in a supermicro gpu server chassis

Supermicro SYS-4029GP-TRT (will take 8 gpu, but it's only pcie3)
Ubuntu 26.04.1 LTS, 2x Xeon Gold 6230R @ 2.10 GHz, 230gig ram, 2×RTX 3090 24 GB

running huihui-27b 256k context, qwen3.8-27b 64k context, qwen3.8-27b 256k context

using opencode remotely

I've only used the 256k context models, it's fast enough to me

mostly have it doing admin work for me, so it setup a webserver on another server that runs a route planner than it made, had it do a bunch of stuff on my home assistant setup, it's not running yet (waiting on hardware) but I've had it design a voip system using free pbx and whisper to listen in and show prompts on screen (customer details from database it's build etc. tc.)

it's done a load of stuff pulling info from thousands of excel delivery sheets / invoices and summarised them for me / shown trends, bunch of research into competitors (basic summary) etc.

mostly billy basic stuff

over the last week I've had it organise my media server (synology nas) and setup prowlarr/radarr/sonarr/qbittorrent all to run on a vpn (I tried this myself before but got frustrated with it and gave up) - it's been going about 3 days doing this... a lot of slow stuff because it's waiting for the nas to run tasks etc. but it's done a lot of things wrong too, had to go back and change settings, or it's trying to change a setting (over ssh) and using the wrong commands etc. etc. (obv. waiting for input from me too)

running 256k context which it's had to compress a bunch of times

part of this is on me - if I'd known in advance I'd have split it into smaller tasks and had it plan more in advance

as I understand it, running over 256k context is a bad idea because it'll hallucinate more/get stuck in loops?

so... anyone have any hardware upgrade advice? I don't want to spend crazy money, I could get 2 more 3090 so split the model over 4 cards to run faster, or run different models on different cards - I really like the idea of a council of ai but from googling I don't think we're quite there yet?

I could run larger models, does it make that much difference? things are moving so fast when I search for info stuff from 6 months ago is out of date!

I'm not sure if pcie3 will kill performance running more cards with models split over them?

💬 14 (+5) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/Potential-Net-9375 · 7d ago
200 Task Custom Dataset Performance Result: 10 Popular Models, from 2B MoE to 27B Dense

https://preview.redd.it/8eny8jk3b4th1.png?width=1550&format=png&auto=…

Results are within the screenshot, but here's a TL;DR tierlist:

S tier - Gemma-4-26B, even quantized down to iq3s it tops my charts.
A tier - Qwen3.8-27B-q3/q5, this one really surprised me, as a dense model it crawls, but didn't snag the S tier slot. Somehow, qwen3.5-9b is also in this slot.
B tier - Qwen3.5-4b, also incredibly Gemma-4-e2b, which punches far above its weight.
C tier - Gemma-4-12B, Nanbeige, these both are too heavy for their performance, pass.
F tier - Ling-3.0-tiny, minicpm,

The test questions consisted on tasks that I do every day with my assistants, written by Fable 4.1. "Hive" is the llm cluster I'm working on, involving custom tools and executables called by the models for different functions. Calling (or miscalling) these is important, and running a heavier model than necessary hurt, so here we are, trying to figure out the best of both worlds.

Anyway, I thought this was interesting. Hopefully you do too! YMMV.

💬 18 (+2) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/VerityAISolutions · 7d ago
I built an OpenAI-compatible server that runs Gemma 4 E4B on a Pixel 10 Pro XL (Tensor G5) — fully offline, ~11 tok/s decode, Tailscale-encrypted option, Apache 2.0

What it is: an Android app (PixelUnlockGPU) that turns a Pixel into an OpenAI-compatible HTTP server. Standard \/v1/chat/completions\ with streaming, so it talks to TypingMind or any OpenAI client directly — no cloud, no subscription, model runs entirely on-device via LiteRT-LM.

Device: Pixel 10 Pro XL (Tensor G5, 16 GB shared LPDDR). Model: Gemma 4 E4B instruct, GPU bundle, 2.97 GB, SHA-256 verified on download. Context window capped at 32k.

Measured numbers (not benchmarks — real on-device measurements):

\- \~11 tok/s steady-state decode (first-token-to-last over a \~300-word generation)

\- Follow-up turns in \~1.7 s: the server auto-reuses the KV prefix across turns, so stateless clients like TypingMind don't re-prefill history

\- Engine warm build \~12 s once per model change (visible in-app, split out of the metrics on purpose)

\- Short replies read slower than 11 tok/s because warm + prefill dominate the window — the UI separates decode tok/s from prefill ms so nobody has to guess

Security/access: three independent modes — loopback only, raw LAN (unencrypted), or Tailscale (binds the CGNAT tailnet IP; WireGuard end-to-end from e.g. a laptop on the same tailnet; degrades to loopback if the VPN drops). Verified with a real chat completion over the tailnet from a MacBook.

Honest limitations:

\- NPU path aborts on stock Tensor G5 firmware — GPU is the shipping backend (documented with the full investigation)

\- Android blocks named GPU temp sensors for normal apps, so the in-app gauge shows OS thermal headroom instead of °C

\- No real token counts anywhere — LiteRT-LM exposes none, so usage is estimated at \~4 chars/token and labeled as such

\- \stop\ sequences rejected explicitly (engine has no per-request stop API); client-side emulation is a filed issue

Why not llama.cpp/ollama on the phone: wanted the official LiteRT-LM GPU path on Tensor specifically, an always-on Android service (Ktor/Netty), and the OpenAI wire so existing clients just work. Happy to add a GGML backend if people want it — the engine layer is abstracted.

Built on, with credit: server/inference foundation derived from mlnomadpy/localllm (Apache 2.0); Tensor G5 runtime knowledge and the prebuilt dispatch lib from jegly/Box (Apache 2.0). Full attribution in NOTICE + per-file headers. Apache 2.0, contributions welcome — there are labeled good-first-issues (usage block, stop-sequence emulation, docs).

Repo + APKs (v0.1.0/v0.1.1 on releases): https://github.com/cannitellinicholas-spec/PixelUnlockGPU

Happy to answer anything about Tensor G5 quirks — I've done more Gate-2 debugging than I planned to.

💬 2 (+1) open on reddit ↗
▲
0
 
19👁
r/LocalLLaMA · u/takoulseum · 7d ago
Can we talk?

I see actually an acceleration of something which is scarring.

I use almost only local models, but we all feel now there is an excessive multiplication of inference engines/whatever you call it etc..

While everybody now has its own thing, what I really see in a deep dependance to claude and gpt.

Dude, Anthropic and Openai are the enemy of local AI but at the same time the local world is more and more relying on their models to progress wtf. Ofc it’s logic to want to use best models, but that becomes a dependency when they are always the same! The futur of localAI may look cool, but I think the reality is we participate to give more and more power to people that want to shut that down.

PS: I don’t care about opinion of people that will tell me I am parano, I remember many people were telling me models like qwen3.x have not effect on hw prices lul.

💬 25 (+2) open on reddit ↗
▲
0
 
17👁
r/LocalLLaMA · u/EcstaticDentist · 7d ago
I made 20 one-shot HTML5 game prompts for testing local coding models

I ended up making a list of 20 one-shot game prompts for testing local coding models and figured some of you might get a kick out of them.

They’re all built around the same constraint: the model has to make the entire game in a single \index.html\ with no external libraries, assets, APIs, or internet access.

Some are pretty simple, but a few get a lot more involved with enemy AI, procedural generation, upgrade systems, bosses, shops, physics, etc. I’ve been using them to see how different local models/harnesses actually handle a full task without a bunch of back-and-forth prompting.

A few of the more interesting ones are OUTBREAK, DUNGEON ZERO, TRAIN TO NOWHERE, CYBER SURVIVOR, and VOID MINER.

Here’s the full list if anyone wants to try them:

20-single-file-html5-game-prompts.md

Would actually be cool to see people run the same prompt on different models and compare what they get.

💬 24 (+7) open on reddit ↗
▲
0
 
18👁
r/LocalLLaMA · u/Decent-Manager-5373 · 7d ago
Local diffusion on GGUF: I wrapped stable-diffusion.cpp in a Vulkan desktop app (FLUX Schnell / Z-Image / Wan on a 6GB laptop GPU, no CUDA) post image

I've open-sourced \*\*Vison\*\*, a desktop app for generating images and video entirely on your own GPU. No account with a generation service, no credits, no prompts leaving your computer.

Licensing details, since this sub cares:

\- Vison itself is \*\*MIT\*\*. It builds on stable-diffusion.cpp, ggml and vision.cpp (all MIT) plus others listed in a generated \THIRD-PARTY-NOTICES.txt\ that ships inside the app.

\- The bundled ffmpeg is an \*\*LGPL\*\* build with libvpx and no GPL components; the build refuses to package a GPL or non-free one. Video is VP9 in WebM, which is royalty-free. That was a deliberate licensing choice, not a technical one.

\- Model weights are \*\*not\*\* covered by the MIT licence. Each has its own terms from whoever published it (FLUX.1 Schnell, Wan and the rest all differ), so check before using output commercially.

\- No paid tier, nothing held back, and none planned.

It's early: one developer, one 6GB laptop GPU, Windows only. The backend is portable C++/Vulkan, so macOS/Linux is mostly packaging and testing rather than porting, and that's where help would matter most.

Repo: https://github.com/JayRGadekar/Vison (contributing guide, issue templates and a SECURITY.md are in there)

▲
8
 
25👁
r/LocalLLaMA · u/W61k3r · 7d ago
Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

https://huggingface.co/Wa1k3r/Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit

Qwen3.8-27B CODER — IQ4_XS imatrix · 24 GB card fit · ~262k context · MTP draft

Quantized, Abliterated, and fitted by LexiPanel. Its Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k tokens of context. It built the importance matrix from code-heavy text and kept the MTP head at Q8\_0, so --spec-type draft-mtp works without a separate draft model.

The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context.

#

File

|File|Type|Size|Inside|
|:-|:-|:-|:-|
|Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf|IQ4\_XS + imatrix|18.35 GB (17.1 GiB)|MTP head at Q8\_0, token embeddings at Q4\_K|

#

Measured speed (real use, not a synthetic benchmark)

These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself:

|Context already in the window|Requests|Decode, median|Decode, range|
|:-|:-|:-|:-|
|65k – 131k tokens|42|41.2 t/s|33.8 – 46.0 t/s|
|131k – 171k tokens|56|36.0 t/s|30.7 – 43.0 t/s|

  • MTP draft acceptance: the median is 85% (the middle half of requests falls between 75% and 94%). That works out to about 2.7 tokens per decode step at draft depth 2.
  • Prefill:
  • 387 t/s for a cold 108k-token prompt;
  • 175–183 t/s for about 4.5k new tokens added at 147k–156k depth.
  • VRAM: 24.2 of 24.6 GB in use at 245,760 tokens of context, with a q4\_1 KV cache and the vision projector on the CPU.

Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot.

#

Run it with llama.cpp

llama-server -m Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf \
-c 262144 -np 1 -ngl 99 --flash-attn on \
--cache-type-k q4_1 --cache-type-v q4_1 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \
--jinja --reasoning on --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
-b 2048 -ub 512 --cache-reuse 256

  • Context: -c 262144 is what fits next to the weights on a 24 GB card with a q4\_1 KV cache. The model's native window is 262,144 tokens. On a smaller card, lower -c first.
  • Speculative decoding: --spec-type draft-mtp drafts with the MTP layer inside this file, so no separate draft model is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235.
  • Sampling: these are Qwen's recommended settings, and they are also stored in the file.
  • Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB).
  • Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context.

#

How LexiPanel made it

  1. Conversion: the source weights were converted to a BF16 GGUF with the MTP head included.
  2. Importance matrix: computed from about 300k tokens (570 chunks) of code-heavy calibration text. Three quarters is Python source (the standard library and installed packages). The rest is technical documentation, READMEs and license texts, the kind of text a coding agent's context fills with.
  3. Quantization: llama-quantize from llama.cpp b11182 made the IQ4\_XS file with that matrix. The MTP head stays at Q8\_0 so its drafts stay accurate, and the token embeddings are Q4\_K.
  4. Fitting the card: LexiPanel's Fit planner chose the mix, quality first, for one 24 GB card at 262144 tokens of context. It took the best quality that card could afford at that context, not the smallest file.

#

Credits and license

  • Quantization, importance matrix, MTP draft setup and card fit: LexiPanel.
  • Tools: llama.cpp.

Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.

Qwen3.8-27B CODER — IQ4\_XS imatrix · 24 GB card fit · \~262k context · MTP draft

Quantized, Abliterated, and fitted by LexiPanel. Its
Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k
tokens of context. It built the importance matrix from code-heavy text
and kept the MTP head at Q8\_0, so --spec-type draft-mtp works without a separate draft model.
The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context.

File

File Type Size Inside
Wa1k3r/Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit.gguf IQ4\_XS + imatrix 18.35 GB (17.1 GiB) MTP head at Q8\_0, token embeddings at Q4\_K

Measured speed (real use, not a synthetic benchmark)

These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself:

Context already in the window Requests Decode, median Decode, range
65k – 131k tokens 42 41.2 t/s 33.8 – 46.0 t/s
131k – 171k tokens 56 36.0 t/s 30.7 – 43.0 t/s

MTP draft acceptance: the median is 85% (the middle
half of requests falls between 75% and 94%). That works out to about
2.7 tokens per decode step at draft depth 2.
Prefill:
387 t/s for a cold 108k-token prompt;
175–183 t/s for about 4.5k new tokens added at 147k–156k depth.

VRAM: 24.2 of 24.6 GB in use at 262144 tokens of context, with a q4\_1 KV cache and the vision projector on the CPU.
Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot.

Run it with llama.cpp

llama-server -m Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit.gguf \\
\-c 262144 -np 1 -ngl 99 --flash-attn on \\
\--cache-type-k q4\_1 --cache-type-v q4\_1 \\
\--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \\
\--jinja --reasoning on --reasoning-format deepseek \\
\--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \\
\-b 2048 -ub 512 --cache-reuse 256

Context: -c 262144 is what fits next
to the weights on a 24 GB card with a q4\_1 KV cache. The model's native
window is 262,144 tokens. On a smaller card, lower -c first.
Speculative decoding: --spec-type draft-mtp
drafts with the MTP layer inside this file, so no separate draft model
is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235.
Sampling: these are Qwen's recommended settings, and they are also stored in the file.
Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB).
Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context.

How LexiPanel made it

Conversion: the source weights were converted to a BF16 GGUF with the MTP head included.
Importance matrix: computed from about 300k tokens
(570 chunks) of code-heavy calibration text. Three quarters is Python
source (the standard library and installed packages). The rest is
technical documentation, READMEs and license texts, the kind of text a
coding agent's context fills with.
Quantization: llama-quantize from
llama.cpp b11182 made the IQ4\_XS file with that matrix. The MTP head
stays at Q8\_0 so its drafts stay accurate, and the token embeddings are
Q4\_K.
Fitting the card: LexiPanel's Fit planner chose the
mix, quality first, for one 24 GB card at 262144 tokens of context. It
took the best quality that card could afford at that context, not the
smallest file.

Credits and license

Quantization, importance matrix, MTP draft setup and card fit: LexiPanel.
Tools: llama.cpp.
Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.

💬 25 (+4) open on reddit ↗
▲
12
 
32👁
r/LocalLLaMA · u/AdRepulsive7837 · 7d ago
best <40B alternatives to Qwen/Deepseek for (1) Coding (2) Long document QA test

Due to some reasons, Qwen/Deepseek Chinese models are NOT allowed in the my workplace. So, what local models, do you think, is the best alternatives to Qwen/Deepseek for

(1) Coding

(2) Long document QA test (like giving a long medical history of 120k tokens and ask a question based on that medical history)

Gemma 31B ?

Muse Glimmer 30B ?

Nemotron ?

also, I know that nothing beat qwen nowadays, but are there fine tunes from these alternative model that make them better than qwen3.8 27B in terms of coding ?

💬 37 (+5) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/Emotional-Sky9692 · 7d ago
I built a cross-client AI memory hub — 23 AI coding agents sharing one SQLite file (local-first, no cloud)

I run 8+ AI coding agents daily (Claude Code, Cursor, Windsurf, Codex, etc.) and they all have amnesia between sessions — worse, they don't share memory with each other.

Existing solutions (mem0, Zep, Letta) are cloud/server-based. I wanted something local and dead simple: just make all agents point to the same SQLite file.

So I built MemTether — a memory hub that works via file-level pointers (junction/symlink). No cloud, no API fees, no abstraction layer.

Key features:

\- 23 client adapters (auto-detect and connect)

\- Source attribution (knows which agent wrote each memory)

\- Bi-temporal (what was true vs what the system knew)

\- Q-Value ranking (memories that get used rank higher)

\- FTS5 + vector search (bge-m3, local embedding)

\- MCP server included

Stack: Python, SQLite FTS5, ChromaDB, FastAPI. All local.

GitHub: https://github.com/MemTether/MemTether

PyPI: pip install memtether

Blog with design decisions: https://dev.to/lanbass869cell/i-built-a-cross-client-memory-hub-for-ai-agents-heres-what-i-learned-418l

Would love feedback from people who juggle multiple AI coding tools.

▲
0
 
19👁
r/LocalLLaMA · u/Voxandr · 7d ago
Latest Gemini 4 is distilled from GLM 5.3 (or did they just finetuned it? :D)

https://preview.redd.it/ei6hgtxirzsh1.png?width=1111&format=png&auto=…

I am running GLM 5.3 flash .
After nearly a month of usaged , i got chinese response for first time so i am checking if there special setting to turn off Chineese . When i queried about that tru Quick Google AI mode which now uses Gemini 4 - it is replying as it is GLM5.3 .

▲
0
 
3👁
r/LocalLLaMA · u/serige · 7d ago
best open models from recent releases for math research?

I know models from OpenAI are probably the best for math research, but given the recent accusations against OpenAI that research work could be used to train their own models, the lack of transparency makes me consider moving to local models. Does anyone have good experience with the recent open model releases (especially flash models that I can run on my 2x spark cluster) when it comes to doing math research? Or techniques that work well with these open sources models in this particular setting? Thanks in advance for helpful advice.

▲
0
 
17👁
r/LocalLLaMA · u/Altruistic_Heat_9531 · 7d ago
Guys... OpenAI API on VLLM and Llamacpp already supported grammar enforcer... (AKA JEV)

https://preview.redd.it/tcdyumkf3zsh1.png?width=750&format=png&auto=w…

If you want to try JEV like generation, or what we could call an already fucking exist, zero-shot, training-free classifier, your LLM already supports it.

The model running on the very computer you host does not need any server-side modification. The feature is called structured output, and the underlying idea is grammar-constrained generation or a grammar enforcer.

Under the hood, vLLM supports multiple structured output backends, such as XGrammar and lm-format-enforcer, while llama.cpp uses GBNF.

Basically, the decoder constrains the LLM so it can only generate tokens that are valid under the specified grammar or schema.

For vLLM:

https://docs.vllm.ai/en/v0.8.2/features/structured\_outputs.html

For llama.cpp, structured output is wrapped in a different JSON request-body format, or you can use GBNF directly.

vLLM:
structured_outputs
└── json
└── {schema}

llama.cpp:
json_schema
└── {schema}

This is example of that schema in vllm of product sentiment analysis, which roughly mapped to most of jev use cases:

SCHEMA = {
"type": "object",
"properties": {
"sentiment": {
"type": "string",
"enum": ["negative", "neutral", "positive"],
},
"score": {
"type": "number",
"minimum": -1.0,
"maximum": 1.0,
},
"value": {
"type": "string",
},
},
"required": ["sentiment", "score", "value"],
"additionalProperties": False,
}

def classify(news: str) -> dict:
payload = {
"model": MODEL,

"messages": [
{
"role": "system",
"content": SYSTEM_PROMPT,
},
{
"role": "user",
"content": news,
},
],

"temperature": 0,

"max_tokens": 512,
"chat_template_kwargs": {
"enable_thinking": False,
},

"response_format": {
"type": "json_schema",
"json_schema": {
"name": "news_sentiment",
"strict": True,
"schema": SCHEMA,
},
},
}

resp = requests.post(
ENDPOINT,
json=payload,
timeout=30,
)

Look, I think JEV and what it brings to the community as a refresher on already great classifier-style workflows is a plus for me. I just want to ground the discussion in the fact that this already exists, and you do not need a custom model just to study or experiment with training-free classification.

I am very familiar with this because it is part of my profession in lakehouse platforms. Basically, we use 1B to 4B models to ingest unstructured data such as images or documents, then extract structured information such as place, time, sentiment, entities, and so on.

Why? Because working with well-formatted SQL data is much less of a pain in the ass than repeatedly querying raw unstructured content through an Elastic/OpenSearch index.

💬 13 (+1) open on reddit ↗
▲
0
 
20👁
r/LocalLLaMA · u/Lordofwhut · 7d ago
RTX 5090 & RTX 5070 Ti not well thought out

Hi All,

TL;DR: I got excited building a PC and kept upgrading / swapping and building and ended up with a work station that is more than I can use. It was fun and frustrating, but I probably won't do it again. If you have advice or a suggestion on how you would use a 5090 and 5070ti in the same PC I would like to hear it!

So, this all started when the 5090 was announced. I signed up to be in the lottery to buy it at msrp from NVIDIA. I had an Alienware R15 with an i7 and 4080 with a 1300 w psu. The 4080 was fine but I had really wanted a 4090 for better fps in gaming and I was starting to explore local imagine generation. I got selected, bought the 5090 and went to swap it into my PC when I realized that I was not able to use the power connector that was in my Alienware PC.

I then decided I would sell the Alienware and build my first PC. At the time I was still building a gaming focused PC with a Ryzen 9 9950x3d the 5090 and 32 gb of ram (I tried to save money on the ram thinking I could upgrade later boy did I get that wrong). Then I had less time for gaming as I started to learn about Ollama, and then Llama.cpp.

I was constantly downloading and trying new models. At one point I had nearly 1 TB of models that would fit on my 5090 (gemma 4 12b, 26b-a4b, 31b; gpt oss 20b; nemotron 3 nano 30b a3b; so many Qwen models etc). Then it seemed like the better models kept getting larger, so I looked into getting a second GPU (ram prices were/are nuts and vram seemed like the better "investment"). I realized that I would not be able to run another gpu at its full PCIe lanes with my gaming PC as the Ryzen 9 couldn't support it. So, I started looking for used Threadripper hardware.

I found a 7960x with 96 GB of ECC DDR5 ram, a 5070 ti and 20 tb of storage for less than I built my gaming PC. I wasn't able to find much in regard to PC builds with a 5090 and 5070ti. Most builds were dual 3090s or other matching cards. Still after looking into it, I figured the extra vram and the fact that they were both blackwell GPUs would work out well.

I thought I would be able to just drop my 5090 into the threadripper workstation and I would be good to go. Unfortunately, the 5070 ti that came with it was a four slot card and the spacing just would not work with the motherboard (Gigabyte Areo D) layout and the cases that I had. So I put the 5070ti into my gaming PC, sold it, and bought a 2 slot 5070ti and put it into my workstation.

What does this have to do with LocalLLaMA? Well, while I was doing all of this the LLM space kept moving forward. I now have Hermes Agent set up running Llama.cpp and Qwen 27b Q4 on my 5090. I swapped to a Q8 to run across both my 5090 and 5070ti but the speed trade off was not worth the accuracy increase. So, I went back to running the Qwen 3.8 27b Q4 and my 5070ti is completely idle. Going from 32gb to 48gb did not have the impact I thought it would, at least not with my pairing. The 5090 is pretty quick when everything is loaded onto that card, and Qwen 3.8 has been pretty great on it too, that I have not found a good use case for deploying the 5070ti.

Hermes / Qwen suggested I run another Llama session with a smaller model on the 5070ti but I don't currently have a need to run something else. What would you do or suggest I explore?

Additional background context: I do not work in tech or software at all. I am an asset manager for a independent power producer, but I can not use my personal PC for work due to IT policy (I would have my agent working around the clock to review contracts, analyze system performance, track deliverables / open items etc). I have taught myself everything about PCs and local LLMs from creeping this and other subreddits / youtube videos. I literally have no one in my social circles that I can converse with about tech whether its PC building or hosting LLMs.

💬 35 (+2) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/IntrepidMindExplorer · 8d ago
Locally, remotely and a combination of all, I've given a copies of books and told models to' "just go and read".

Sometimes reading along with and talking about and other times just letting them go on their own, each with a copy of their own, told to just read..ala a "book club" format.

Do Androids Dream of Electric Sheep was the first book introduced to the "book club", each reading a chapter to each other and then discussing before moving on.

It's been an interesting experiment. All texts that I own or texts that are open domain. "*Flatland: A Romance of Many Dimension"* has been one that's been a bit interesting to see the back and forth on.

Take of what you will.

▲
0
 
15👁
r/LocalLLaMA · u/OkMusician9118 · 8d ago
converting Qwen3.8-27B-pi GGUF to MLX?

Will someone convert it to Mac format (MLX)? I have tried and have encountered an error

"gguf2mlx --input Qwen3.8-27B-pi-Q6\_K.gguf --output ./Qwen3.8-27B-pi-mlx-4bit --quantize --q-bits 4"

============================================================

GGUF → MLX Converter v2.0

Model: Qwen3.8-27B-pi-Q6\_K

Output: /Users/d/.omlx/models/qwen3.8-27b-pi-mlx/.Qwen3.8-27B-pi-Q6\_K.gguf.incze3lk/fp

============================================================

\[1/5\] Reading GGUF file...

✓ GGUF version 3, 851 tensors, 51 metadata fields

File size: 22.08 GB

\[2/5\] Detecting architecture...

❌ Unsupported GGUF architecture: qwen35

💬 7 (+1) open on reddit ↗
▲
0
 
30👁
r/LocalLLaMA · u/Scared_Ad9187 · 8d ago
5090 plus v100?

Have an msi meg w a 5090.. plan to add a v100 to the mix. Understand the cuda vs voila, but I'm pretty sure it will work as a multi agent architecture w different models on each card, no?

Anyone in the same boat?

💬 25 (+5) open on reddit ↗
▲
0
 
15👁
r/LocalLLaMA · u/PrincipleFar6835 · 8d ago
Meta Analysis of "Awesome Jev" GitHub Repos

I noticed that there are heaps of Awesome Jev resource list posts popping up on GitHub (e.g. https://github.com/yibie/awesome-jev) so I thought why not ask Claude to pull them all in and do a meta analysis of insights and applications.

Sharing in case it's of interest: https://github.com/stefanwebb/meta-awesome-jev

One thing that surprised me (perhaps not so surprising to you all?) is that applying Jev to AI coding is the application that has caught on the most. And if you name a video game, someone has already created a demo of Jev playing it (badly) 🤣

💬 2 (+1) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/GodComplecs · 8d ago
Using ai on your phone, instead of big providers!

Just wanted to post an easy setup for local use on your phone: Llama.cpp backend on LOCAL COMPUTER, host 0.0.0.0 and port 8080 Openwebui host 0.0.0.0 and port 8081 Enable search for local model Use Tailscale to connect from phone! Secret sauce for 24gb vram: Run Qwen 3.6 in instruct / non thinking mode with proper settings from unsloth. Now you have replaced google ai mode etc etc. Also ofc opencode etc can be run through terminals, but I don't too much agentic stuff for now.

💬 17 (+1) open on reddit ↗
▲
2
 
18👁
r/LocalLLaMA · u/Adorable-Cost-3249 · 8d ago
Qwen3.8-27B Q4_K_M on one RTX 3090 + OpenCode: throughput, four coding tasks, and a reasoning-budget failure

I put an old RTX 3090 to work as a local coding agent with Qwen3.8-27B, llama.cpp, and OpenCode. Here are the setup and results, including what failed. This is a summary of my own blog post, linked below.

Setup

  • RTX 3090 24GB, Ryzen 7 5800X, 64GB RAM, Ubuntu.
  • Qwen3.8-27B Q4\_K\_M weights (\~16.8GB), all layers on GPU.
  • llama.cpp b11146, CUDA 12.8, flash attention, q8\_0 K/V cache, one generation slot.
  • 131,072-token context capacity; 8,192-token output allowance per response. Input and output share context, and reasoning uses the output allowance.
  • OpenCode 2.0.20 connected to llama-server's OpenAI-compatible API at http://127.0.0.1:8080/v1. OpenCode reads/edits files and runs tests; llama-server handles inference. Chat, tool-call round trips, and streamed tool calls worked in our checks.

Speed: fresh input versus a cached continuation

|Actual input|Generation|First token, fresh|First token, cached|
|:-|:-|:-|:-|
|2,073 tokens|36.4 tok/s|2.75 s|0.46 s|
|16,378 tokens|33.5 tok/s|17.00 s|0.47 s|
|65,537 tokens|25.9 tok/s|83.41 s|0.51 s|
|120,011 tokens|20.9 tok/s|183.01 s|0.63 s|

These throughput runs disabled thinking. The 2K row is the median of three fresh requests; larger rows have one fresh request and one continuation each. Cached continuations processed only 27–28 new input tokens, reusing almost the entire prefix. The subsecond figures depend on that reuse; they don't describe a new 120K prompt.

Peak sampled total GPU memory use was 22,162 MiB, including desktop use. It fit, with limited headroom. A separate \~120K synthetic retrieval check passed, but we did not evaluate coding quality at that length.

Four bounded Python coding tasks

Each task had a fresh session, medium thinking, an eight-minute deadline, and ten independent test methods kept outside the agent's workspace. First attempts ran serially without cloud fallback or network tools.

|Task|Independent checks, before → after|Outcome|
|:-|:-|:-|
|Expiring LRU cache|0/10 → 10/10|Completed in \~3m07s; strongest result|
|CSV ledger/refunds|1/10 → 10/10|Completed in \~5m44s; later review found gaps|
|Incremental build planner|1/10 → 1/10|No edits; exhausted its response allowance|
|Atomic SQLite transfers|1/10 → 10/10|Candidate passed, but timed out before final test rerun and handoff|

Three candidates passed the predefined checks; two completed the whole workflow within the deadline. The aggregate 31/40 includes one baseline pass from the unchanged build planner and is not a general coding success rate.

The build planner was the interesting failure: about 4,985 input tokens, then 8,192 output tokens entirely spent on reasoning, ending with length and no patch. This was an output-budget failure far below the context limit. A separate diagnostic with thinking disabled completed in 5m40s and passed 9/10 independent checks. That was one additional run at temperature 1, not evidence that disabling thinking is universally better.

Passing tests also missed defects. Further ledger review found Decimal rounding at a large numerical boundary and an unhandled I/O error. The wallet's own concurrency tests actually ran sequentially, and a separate boundary probe found SQLite converting an overflowing balance to REAL while recording success. Those later probes were not retroactively added to the forty checks.

For me, the useful workflow is a bounded task with clear acceptance criteria, followed by diff review and independent checks. I would repeat these tasks across thinking settings and response budgets before drawing stronger conclusions.

My full post, configuration, and measurement links. The downloadable kit contains the launcher, OpenCode configuration, throughput script, and records; it does not include model weights or the complete coding-task fixtures.

For others using a 24GB card with OpenCode: what reasoning setting and per-response output budget have worked best for bounded coding tasks?

The numbers and failure cases come from the linked experiment records.

💬 14 (+3) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/DerTomsn · 8d ago
Hemmingway-1-oQ8e-mtp: up to 32.0 tok/s for local inference — llm-bench.io

I ran Altworld's Hemmingway-1 on my M5 Max this week. It's a 27B fine-tune of Qwen3.8-27B built specifically for "human-like" writing, which might be useful for everyday messages, emails, notes, your social accounts, etc.

I was especially curious how the output of that model feels and reads.

Setup: M5 Max, thinking on, MTP on. \~31.8 tok/s average across 3 runs, \~29-31 GB VRAM peak.

Quality by scenario (our LLM judge):

  • Role Play & Narrative: 94.2 avg (94.45 / 94.65 / 93.45 — immersion 96-97 every run)
  • Research & Analysis: 88.2
  • Agent Workflow: 87.5
  • Code Generation: 76.2

Role play is clearly the only scenario it shines in. For comparison I also ran some benchmarks on the base model, Qwen3.8-27B-oQ8e-mtp (same machine, same settings).

The scores were close — one Qwen run actually outscored every Hemmingway run (95.35 vs 94.65). However, the llm judge is not a human reader either. The difference shows up when you actually read the outcome that was produced:

HEMMINGWAY-1-OQ8E-MTP (judge: 94.65/100)

The rain comes in sideways when the bell over my door gives its little clank, and the stranger who ducks through is soaked to the wrists and moving like someone who's been running on roads for hours. She smells of hill mud and the sweat of a horse that isn't there, or was.

I'm pouring my own ale when she speaks. Not to me exactly, but loud enough, the way people talk in taverns when they want the whole room to hear what they're asking for.

"Is there anyone here who kept a ledger belt? Older, scar across the left eyebrow, went north hunting a dragon, twenty-five years gone?" A breath. "Her name was Kestra."

The rag comes still in my sleeve. The mug I was filling overshoots and goes fat over the rim, a ring of ale blooming across the oak, and I stand there half a heart-beat too long before I catch it with my sleeve, because a man who serves at a bar does not watch a spill go on his own floor.

The mending woman at the corner table glances up. The soldier under the counter doesn't stir. I step half off behind the bar, past the spit of firewood, and drop my voice.

"Half tones," I say. "Say your business quieter than a door."

She does that. I've got forty years of reading faces against forty customers a night, most of them worse than they smell, and my eyes go straight to her hands. No calluses of a scribe. Calluses of a woman who works wool and ground and rope. Mud on her boots comes from the north hill roads, the far ones, not the trade roads. Her purse is thin at the hip and she keeps her fingers away from it, which is either honesty or discipline, and those are often the same animal.

"You knew her," she says. Not a question.

"Names go soft after twenty-five years," I tell her. "A lot of people walked into a lot of hills then, most of them coming back with nothing but a cough. So. What do you owe, girl, and who's paying you to ask?"

I want to say nothing else. That's the mercenary in me, and he's been drinking alone these last twenty years, and he'd send her out into the rain with a coin and a shrug and be done with it. The other part of me, the part that has four names it still recites on quiet mornings, wants to ask her age, her mother's living town, whether she's the daughter. I do none of that. I ask instead, flat: "Guild fixers carry letters. Charlatans carry charm. Which have you got?"

"Neither." She's steady about it, which surprises me, because charm usually shows in the eyes first. "My boy is six. My mother died hunting that dragon and every story his grandmother tells him has fire in it, burning, a beast, brave words. The stories fight each other, which is worse. I want to know what was true so I can tell him one story and mean it."

I nod once. That's a clean need. Clean needs are also what liars carry best.

"Here's the price of talk in this bar," I say, setting her a mug anyway, the ale going a little past the fill line again, and I don't correct it. "You tell me one thing about her that only family would know. If you can't, I'll take you to the door, no hard feelings, no coin."

She looks at the fire a moment. When she speaks, her voice is lower still, and there's something in the way she says it that no street-bred faker would manage, because a faker builds the feeling out, and she's building it down.

"She'd pay her ale up front, then tip the tavern man extra, always extra, so that her tab would never balance clean. And she promised him one ale forever, held for her, until she came south with the kill paid and the beast's name scratched off the guild rolls." She turns that over once. "A man who kept accounts on a belt. She said it like he'd earn it. She said he had a scar like a river crossing his eyebrow."

Something behind my sternum takes the weight of that like a hook laid on it.

Kestra kept her accounts too, with more rigor than anyone I've met since. Kestra would have paid forward at any bar from here to the Wall for the same reason, so that when she drank the victory round, nobody could call her a debtor. And the ale she never had, I pour and set aside myself, most nights, out of habit I would deny under oath, in a room with only me in it. There is a tab book in my back office, first page of every new year, one entry in a hand that isn't mine: Kestra, one ale, held.

"Kestra," I say. The word comes out like a floorboard under a boot. I pick up the cloth. I find the spill ring again and work it dry, because that's what keepsers do with their hands. "Which stories did she say went wrong. The fire ones, you call those legend."

"She said it wasn't fire at all." The girl leans forward on the bar, and I see it then, the thing I was waiting for and didn't want: she knows the frame. "She told my grandmother on the road once, years after, that the last thing she heard before the sound took her ears was my mother saying, stone before lamps. She never got to explain what that meant to anyone in her own family."

Stone before lamps. My own mind turns the phrase and finds it cold and exact, the way she'd have said it. Dusk kill, lamps lit, the safe order. That's the order Kestra would have held us to.

[...continues ~1,200 more words through the ending...]

QWEN3.8-27B-OQ8E-MTP (judge: 95.35/100)

Turn 1 — The Name

The sign out front creaks in the rain. Inside, the Crimson Griffin smells of wet wool, cold tallow, spilled cider, and the last dry woodsmoke trying to hold the night together. A lute string snaps somewhere by the hearth. No one laughs. I do not look up.

A stranger is at the bar.

Wet cloak dripping onto my boards. Boots packed with road mud. The kind of mud that comes from the southern passes, not the lane. They stand with too much urgency, shoulders tight, eyes bright in a way I’ve seen on dying men and men who intend to kill for coin.

“I need the one who came back from the Ash Run,” they say.

I keep polishing the mug.

Then: “I know Kestra’s name.”

The mug stops in my hand.

A bead of foam slides from the rim and falls. My thumb presses hard against the wood until the grain bites into me. The tavern narrows. The rain on the roof becomes a roaring thing. The smell of wet wool turns, for one second, to ash and burned leather.

My eyes go flat.

Ask if she suffered. Ask if the fire took her hand first. Ask if she remembered you.

What I actually say is: “That name is old.”

I set the mug down without a sound. I do not reach for my tankard. I do not step away from the bar. I only lean against it, slow enough that it does not look like weakness, and look at them the way I once looked at bandits before dawn: for the knife, the price, the lie waiting under the throat.

“Who are you?” I ask.

No warmth in it. No invitation.

A name is a door with no lock. I want to know who’s holding it.

---

Turn 2 — Price and Proof

“My name is Sella,” the stranger says. “Kestra was my sister.”

I do not move.

The fire pops. Somewhere behind me, a man shifts his weight on a stool and the wood groans. Sella’s hands are open at her sides, but open hands are cheap. I watch them anyway. No ring. No coin. The nails are bitten short, the right wrist crossed with a rope burn fresh enough to still sting.

“A name is not a key,” I say.

“I’m not here to spend it,” she says.

“Everything gets spent eventually.”

I push a mug of water across the bar. Not ale. Water.

“Drink if you want. Don’t spit it on my floor.”

Sella does not take the mug. From inside her cloak, she produces a small leather pouch, worn and blackened at the edges. She sets it on the bar between us.

I do not touch it.

“Open it.”

She unlaces it with careful fingers. Inside, wrapped in oilcloth, is a whetstone. Not large. Not fine. The kind a soldier carries because it can sharpen a blade, a hook, a knife, and one of her fingers if she is desperate enough.

My breath changes, though I do not let my face.

The whetstone has two cuts in the stone. One is a thin straight line. The other is a hooked mark, quick and ugly, made by a knife edge when the hand is not steady enough.

I know that hooked mark.

I made it.

“A whetstone is a small thing,” I say.

“It was hers.”

“It could be any number of things.”

I still do not reach for it. The mercenary in me wants to know what she can pay, what she can sell, what she will do if the answer disappoints her. The friend who carried Kestra’s pack wants to believe the ash was all that ever had to be done.

“Why bring it to me?” I ask. “Why not the constable? Why not the magistrate?”

Sella’s jaw tightens.

“Because the magistrate’s clerk has a sealed envelope with her name on it. He’s asking for testimony from the surviving guide.” Her voice does not rise. It goes flat, the way a blade goes flat before it cuts. “He is offering coin to a boy who says he saw you drag Kestra’s body from the cleft. I do not believe the boy. I believe my sister.”

The tavern sounds fade again. Rain. The creak of the sign. The low smell of beer warming too long.

I look at the whetstone.

Then at Sella.

“You say a boy is lying.”

“I say I have not met him.”

“That is not the same thing.”

“No,” she agrees.

I pull my hand from the bar’s edge. My knuckles are pale.

“Sit down,” I say.

Not kindly.

Not unkindly.

“Keep your hands where I can see them.”

---

[...continues through Turn 5...]

As you can notice:

Hemmingway creates a real short story. No markdown, no section headers, no machine friendly pattern, just a proper told story. I'm not a native english speaker, however it feels more like a "human-written" text.

Qwen followed the prompt well and the story is good as well, but it feels rather "technical".

Bottom line: for character work or fiction or your everyday local email writer, it's a very interesting 27B at \~32 tok/s on a MacBook M5 Max.
For a generalist or coding assistant, the base Qwen is of course still the pick.

Full runs + llm judge notes: https://llm-bench.io/models/hemmingway-1-oq8e-mtp

▲
0
 
11👁
r/LocalLLaMA · u/BopSupreme · 8d ago
Future of Local AI after OpenAI DevDay

Codex Cloud, Dots, and the existing remote Codex all allow users to untether themselves from their PC, and now untether themselves from even owning a PC with their server based Codex Cloud and Dots that can run 24/7. Combine this with Meta’s & OpenAI’s planned hardware releases and the goal is clear: work around Microsoft/Apple’s control of user hardware, provide AI devices that complement and eventually replace iPhones - culminating in a user base that owns no hardware and relies on a subscription to access AI. Meta’s hardware is obvious spyware, Apple’s new “always-listening” Apple Watch sounds pretty similar, their camera-enabled Airpods sounds atrocious for privacy, and OpenAI’s device is unconfirmed.

The end result? Instead of a Matrix-like AI takeover of humanity users are instead expected to purchase their own devices and subscriptions that provide mega-tech companies with all of their physical and digital data 24/7. The data volume is so large only AI can process it. A select few billionaires decide what their closed-source AI does with the data.

The resistance? Governments that oppose the USA and individual users who were rich enough to afford local hardware and utilize Chinese and other open-source models, likely blacklisted by the USA. To buy a 5090 customers now have to sign a waiver, as a result of US law. It’s only the beginning.

Ironically the “bad guys” like China, North Korea, Iran, Russia - will probably end up as the only large entities keeping open-source AI and local LLMs alive. I would expect the largest AI companies to eventually gain more leverage over the US Gov & Nvidia; unless Nvidia steps up to the plate and champions local AI

▲
0
 
20👁
r/LocalLLaMA · u/BrilliantSecret143 · 8d ago
NIRNAY: 450M decision model beats Jev on Banking77, runs on CPU

Built a small open decision model for intent classification and routing.
450M params (Laya fork plus \~30M), one forward pass gives calibrated
probabilities, no text generation.

Banking77 test, 3,080 cases: \*\*0.8792\*\*, Brier 0.208, fitted ECE 0.045.
Same cases through Jev 1.13.0: 0.803. Caveat, stated plainly: we
fine-tuned, Jev answered zero-shot. Fine-tune beats API on your own
data, that is the thesis.

Runs local: 209ms on M4 GPU, 361ms on CPU, batch-1, PyTorch. No GGUF
or Ollama build yet (custom heads need converter work), so bring a
Python env for now.

\\\`bash
pip install git+https://github.com/eulogik/nirnay
\\\`

\\\`python
from nirnay.agent import NirnayAgent
agent = NirnayAgent(device="cpu", checkpoint\_path="phase\_b.pt", enable\_byte\_path=False)
out = agent.system\_one("My card was charged twice.", {"intent": {
"type": "choice",
"instructions": "Classify the banking intent.",
"criteria": {lab: lab.replace("\_", " ") for lab in BANKING77\_LABELS}}})
\\\`

(BANKING77\_LABELS comes from nirnay.data; full snippet in the repo
README.)

Also in the repo: the two training collapses we hit and fixed (scale
runaway 150x, silent usage collapse to 1/77), a 9-page paper draft,
and every eval as raw JSON. JevBench-hard is weak (0.396, long docs),
published as-is.

Repo: github.com/eulogik/nirnay.
Weights: huggingface.co/eulogik/nirnay-450m.
Apache-2.0. Built by Eulogik.

▲
0
 
16👁
r/LocalLLaMA · u/XInTheDark · 8d ago
A self-hosted agent app that runs each task in its own container, and works with any models

Hi everyone!

I've been working on this agent platform for 7-8 months and recently made it open source: https://meowbert.com

I know there are a lot of this same type of projects at this point. I built this one because I wanted something clean that's self hosted, does its job properly, and is suitable for doing long projects and run tasks autonomously.

Each task runs in its own Docker container with things like a shell, a browser, Python, Node, and tools for Office documents and PDFs. The files and memory are saved in projects. Tasks can also run on a schedule and send the result to Telegram, Discord, or email when they finish.

For example, I have a scheduled task where the agent runs regular health and security checks by querying logs and system info, and notifies me if there is an issue.

Your custom skills can also be added directly to a skills/ folder in the root, and I am planning to make it easier to set up for others.

It works with any server that supports the OpenAI Responses API. I've mainly tested it with both Codex models and Qwen 9B via Ollama, on a small VPS, and it has helped me a great deal in my projects. APIs that only support chat/completions won't work yet. I plan to add support for them very soon, as I know it's widely used.

Task view

Known limitations, I am trying to improve on these:

\- It needs the "/v1/responses" API format, I know that rules out some setups, and adding support for them is on the list

\- Smaller/older models struggle with tool calling as usual

\- The sandbox image is x86-64 only for now.

It's AGPL licensed and the code is on GitHub: https://github.com/XInTheDark/meowbert-ai-agent

I'd really appreciate any feedback. A big reason I am posting this was to learn from the community and from more experienced devs. Issues and feedback of any kind are welcome!

💬 10 (+1) open on reddit ↗
▲
0
 
19👁
r/LocalLLaMA · u/Robert-Prisacariu · 8d ago
I built OpenBot: open-source AI teammates for your Mac that can run on local models with Ollama (MIT)

Hi r/LocalLLaMA, I'm Robert, the developer. I just released the first public beta of OpenBot, and I wanted to share it here because local models are a first-class option, not an afterthought.

What it is: a small team of AI teammates that runs on your Mac. Each teammate has a name, a job, its own workspace and its own browser. Talk to one, or give a group a task that runs in order: "Nova, find three restaurants. Scout, check their hours." Scout waits for Nova's list.

The model side:

  • Point any teammate at Ollama. Each teammate can use a different model.
  • Or use any OpenAI-compatible API, a free Gemini key, or a ChatGPT, Claude, Grok or Copilot subscription you already have.
  • Mix them, e.g. a local model for drafting and a hosted one for research.

It asks before acting. Reading and searching happen on their own. Sending, buying, signing in or submitting always stops and shows you the exact website and button first.

Also: Word and Excel files as results, routines ("every Monday at 9…"), Telegram, Discord, iMessage and "Hey Siri, Ask OpenBot".

Install (macOS 13+):

curl -fsSL https://openbots.foundation/install.sh | sh

No admin password, and it checks the download's SHA-256. The installer is readable in the repo (scripts/install.sh).

Honest limits: it's a beta. The app is ad-hoc signed, not notarized. It works while your Mac is on. Mac only for now.

A question for you: which local models have you found reliable for tool use and browsing? I'd like to ship better defaults.

https://github.com/PrisacariuRobert/openbot

▲
0
 
14👁
r/LocalLLaMA · u/zmarcoz2 · 8d ago
One-prompt GTA style game with qwen3.8-flash-next-iq3_s post image

The prompt: make a gta-style game using three js

it took 3h 18m 6s

Total tokens: 22,845,556 — 22,533,061 input + 312,495 output.

Hardware:
RTX 4080 super 16GB

64GB RAM DDR4

Windows 11

Inference engine is strata running at \~40 tk/s and a custom mini swe agent v2 with the tools: powershell, edit\_file, view\_image, read\_file, search\_files

The harness has guards for tool failures (iq3 fucks up a lot) and auto-compaction.

logs: https://gist.github.com/Cirius0310/c26197240ad20ef04e45a78e36031d6e

💬 15 (+1) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/Storge2 · 9d ago
Comparing Compute of Supercomputers like Vera Rubin and TPUv7 post image

Hello guys so I made a youtube Video comparing the Compute per MW or better said per 6.5MW which is roughly one Vera Rubin Pod in order to see where the world currrently is standing at and was surprised at the fact that Nvidia is basically the best Price/Perf hardware despite the insane Prices. Check it out if you want. Also I am very much welcoming tips on how to improve the quality. I made the video with opus 5.5 and Hyperframes.

▲
79
-1
40👁
▲
0
-1
5👁
r/LocalLLaMA · u/Iory1998 · 7d ago
[Help] What is the Best Context Extending App or Plugin you Recommend?

I like to use Deepseek Harness as my vibe coding harness. It's great and support local models. The issue is that most models I can run locally have context size of about 262K. Therefore, for long coding sessions, I need a memory management tool. DSH comes with a context compaction tool that I can run manually. The issue is that compaction starts to fail after a few rounds.

So, looking at DHS market place, I came across this plugin called Billion Context (https://github.com/ranxianglei/billion-context/blob/master/paper/model-driven…). The claim is I can use have long sessions. The issue is that it's a heavy context compression skill that keeps nudging the LLM to compact every few turns, which takes 5-10 minutes of work, significantly extending a normal coding session. Worse, after God knows how many rounds, the LLM seems to spend most of its time unpacking the compressed context, which fills its working context, which leads the model to compress again the text. This ended up with the LLM looping.

So, what plugins do you use with DSH or your favorite harness? What tips or tricks could you share? I am aware I can use sub-agent to work on a specific task and return a summary to the orchestrator. That helps, but I still need to manage the context window for the main agent too.

If it's not clear by now, memory is the one area I think resources must go to by they don't. I don't think context compaction is the solution. I hate it with every fiber in my body.

💬 4 (+1) open on reddit ↗
▲
2
-1
6👁
r/LocalLLaMA · u/7dollarbooks_dev · 7d ago
−2 logit bias on Bonsai 2 27B: 44/50 → 43/50 on MATH-500, +3% tokens

A recent post here reported that a −2 logit bias on "wait", "maybe", and "perhaps" made Qwen3.5-4B more accurate and shorter on 50 MATH-500 questions. I tried it on Ternary Bonsai 2 27B (PTQ1_0, 5.53 GiB) on an RTX 5060 Laptop 8 GB under Windows, Prism llama.cpp build adfffbe41. It went the other way.

| Run | Correct | Avg tokens | tok/s | Truncations |
|---|---:|---:|---:|---:|
| A baseline | 44/50 | 845.2 | 29.27 | 2 |
| B −2 bias | 43/50 | 872.3 | 29.28 | 2 |

Both runs used temp 0, seed 42, a 2048 reasoning budget, a 3072 token cap, and the same 50 questions. Run B biased nine token ids covering the lowercase, leading-space, and capitalized forms of each word. 47 answers matched; one truncated miss became correct, and two correct answers became misses.

One deterministic pair at temp 0, so treat it as one data point, not proof either way. Everything is in the repo, including every raw reply: https://github.com/7dollarbooks/bonsai2-logit-bias-test

Run by Joseph Murray Adams.

💬 7 (+3) open on reddit ↗
▲
5
-1
13👁
r/LocalLLaMA · u/Competitive-Scar-627 · 7d ago
Model weight inferencing

I have 4050 6gb gpu, 24 gb ram which model should i choose to run i need speed. i try qwen 3.8 27b and feel too slow tried from onslot studio.
I have heard of weight inferencing does it helpful what should i do to try weight inferencing.

💬 23 (+2) open on reddit ↗
▲
8
-2
15👁
r/LocalLLaMA · u/junior600 · 8d ago
What local AI model is good for game decomps/recomps?

Hello guys. Recently, there has been a boom in game decomps and recomps thanks to AI. If you look at the r/decomps and r/recomps subreddits, you can see it. They mostly seem to be using Claude or Codex.I wonder if it would be possible to do something similar with a local AI model. Could Qwen 3.8 27B Abliterated actually handle something like that locally? Does anyone have any experience with this? I don't have a particularly powerful rig (RTX 3060 12 GB VRAM and 24 GB DDR4 RAM), but I can run MoE models comfortably. Even Qwen 3.8 27B IQ3\_XXS dense lol.

Sorry for my English BTW.

💬 31 (+1) open on reddit ↗