68 posts · 1 sub · RSS
← prev Friday, October 2, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
2
+1
5👁
r/LocalLLaMA · u/itsokimjudgingyou · 7d ago
LINKUP AI sever PCIE 6.0x16 Cables

Has anyone tried the PCIE 6.0 AI server cables made by LINKUP? I don't need 6.0 but the cable routing these cables could offer me is huge compared to the normal risers. The lack of reviews is really the only thing holding me back.

Does anyone have experience with them?

💬 8 (+6) open on reddit ↗
▲
74
+21
10👁
▲
0
 
8👁
r/LocalLLaMA · u/knighty1981 · 7d ago
2x 3090 in server chassis, upgrade time

I've got 2x 3090 in a supermicro gpu server chassis

Supermicro SYS-4029GP-TRT (will take 8 gpu, but it's only pcie3)
Ubuntu 26.04.1 LTS, 2x Xeon Gold 6230R @ 2.10 GHz, 230gig ram, 2×RTX 3090 24 GB

running huihui-27b 256k context, qwen3.8-27b 64k context, qwen3.8-27b 256k context

using opencode remotely

I've only used the 256k context models, it's fast enough to me

mostly have it doing admin work for me, so it setup a webserver on another server that runs a route planner than it made, had it do a bunch of stuff on my home assistant setup, it's not running yet (waiting on hardware) but I've had it design a voip system using free pbx and whisper to listen in and show prompts on screen (customer details from database it's build etc. tc.)

it's done a load of stuff pulling info from thousands of excel delivery sheets / invoices and summarised them for me / shown trends, bunch of research into competitors (basic summary) etc.

mostly billy basic stuff

over the last week I've had it organise my media server (synology nas) and setup prowlarr/radarr/sonarr/qbittorrent all to run on a vpn (I tried this myself before but got frustrated with it and gave up) - it's been going about 3 days doing this... a lot of slow stuff because it's waiting for the nas to run tasks etc. but it's done a lot of things wrong too, had to go back and change settings, or it's trying to change a setting (over ssh) and using the wrong commands etc. etc. (obv. waiting for input from me too)

running 256k context which it's had to compress a bunch of times

part of this is on me - if I'd known in advance I'd have split it into smaller tasks and had it plan more in advance

as I understand it, running over 256k context is a bad idea because it'll hallucinate more/get stuck in loops?

so... anyone have any hardware upgrade advice? I don't want to spend crazy money, I could get 2 more 3090 so split the model over 4 cards to run faster, or run different models on different cards - I really like the idea of a council of ai but from googling I don't think we're quite there yet?

I could run larger models, does it make that much difference? things are moving so fast when I search for info stuff from 6 months ago is out of date!

I'm not sure if pcie3 will kill performance running more cards with models split over them?

💬 14 (+5) open on reddit ↗
▲
50
+16
42👁
r/LocalLLaMA · u/menage_a_un · 7d ago
I've ended up with an AI lab in a public community college. What should we actually be teaching?

Looking for some ideas from people who know a lot more about this than I do.
We've got funding for a small AI lab in a public further education college in Ireland (roughly community college in the US).
The hardware is reasonably decent. The goal is to give students useful skills beyond just using ChatGPT.
If you had the lab, what would you teach them?

💬 70 (+18) open on reddit ↗
▲
26
+13
35👁
r/LocalLLaMA · u/wayneworkman · 7d ago
Peacebell - a from-scratch small language model

I've been using my free time during weeknights and weekends for the last 11 months working on and refining a small domain-specific language model. It specializes on information about World War II.

The number one question I get asked about this is "Why did you pick World War II?" here are some of the reasons:

\- There's a lot of good Wikipedia articles about WWII, and this is permissively licensed. Meaning I can use the materials.
\- There's a lot of good public domain information about WWII in general - more to train on.
\- WWII is factually dense - making it a challenge.
\- The facts surrounding WWII are mostly unchanging - meaning my model would age well.
\- I had to pick a first topic.

A lot of my journey is documented on my blog: https://wayne.theworkmans.us/llm.html though I've not posted recently.

The model is more than from-scratch. I'm using a custom built training pipeline. And I produced all of my own synthetic data to train on (based on Wikipedia articles).

The majority of my time has gone into data curation and balancing.

As I built this model, I've learned a ton about training data for language models, and about language model creation. I tripped over every bump along the way, 100s of times.

I also learned a lot about World War II, and I'm emotionally exhausted. Many know the basics... the Manhattan project, the Holocaust, the concentration camps, Pearl Harbor, D-Day. Though beyond these topics, there is enormously more tragedy than I previously knew. As an adult with my own family now with better ability to comprehend, many times I'd just cry face down on my keyboard from some of the things I learned. Sometimes I would abandon working on it and go to bed early. I've talked with my wife about how awful some of the things that happened are. It's been hard. And I'm ANGRY! So incredibly angry about the atrocities that happened. Especially angry about the things that happened to civilians, non-combatants, women and children.

Well enough of that.

I open-sourced the training materials and the weights. There are two versions of the model. There's a 291M parameter version and a smaller 148M parameter version.

I built the 148M to compete in the various HuggingFace dashboards that limit model size to 150M. Then I built a new benchmark that focuses on WWII topics, that's also on the hub, though the questions are private to prevent them ending up in people's training data (and no they aren't in Peacebell's training data either).

You can try the 291M model for free here. As you use it, keep in mind this is first-version, it's rough, it's not always right. And it really struggles with longer context. Fresh context gives better results.
https://huggingface.co/spaces/wayneworkman2012/peacebell-v1-291M-demo-cpu

The leaderboard is here:
https://huggingface.co/spaces/wayneworkman2012/ww2bench-leaderboard

I've entered Peacebell into various SLM Arena's, such as CodeSoft's SLM arena here:
https://huggingface.co/spaces/CodeSoft/SLM-Arena

Basically everyone in this LocalLLaMA would be able to run the model easily, even without a GPU. There's a customized vLLM fork here that can run either Peacebell model:
https://github.com/wayneworkman/vllm

Next version is expected to be released sometime in 2027.

💬 8 (+3) open on reddit ↗
▲
14
+5
32👁
r/LocalLLaMA · u/FantasticNature7590 · 7d ago
I spent 3 weeks testing local Qwen3.8 on the new low-latency SGLang/vLLM recipes: DFlash2 2.8x. Builds: RadixArk + Inferact 27B NVFP4, 27B BF16, orcarouter 27B Uncensored, Flash-Next NVFP4

Hey guys,

Last time I tested Qwen3.8-Flash-Next on its own. This time I put three Qwen3.8 checkpoints through the same 10 tests on the same RTX PRO 6000:

  • RadixArk/Qwen3.8-27B-NVFP4 (dense)
  • orcarouter/Qwen3.8-27B-Uncensored-NVFP4 (dense, uncensored fine-tune)
  • RadixArk/Qwen3.8-Flash-Next-NVFP4 (MoE)

Each model got the same prompts and its own model card's sampler, with one attempt per task.

Video with the battles, the castles and the ball run: https://youtu.be/VOtfja\_Toj4**

Short version

  • Tests won: Flash-Next 5, 27B 4, Uncensored 0, plus one tie (long-context recall was 100% for all three).
  • Prefill, full window: Flash-Next 22.4s, 27B 97s, Uncensored 99s. Not the same power cap, see section 1.
  • Speculative decoding on the 27B: the DFlash2 drafter took Spec-Bench from 75 to 210 tok/s for one user, 2.8×.
  • SGLang vs vLLM: SGLang was faster overall (210 vs 160 tok/s), but that's mostly the checkpoint. On the one export I ran on both engines, vLLM was 16% faster (169 vs 146).
  • Tool use (BFCL subset, thinking off): 27B 73.3%, Uncensored 70.8%, Flash-Next 64.5%.
  • Battle arena: the 27B scored 700/1000, ahead of Claude Fable 5.1 (678) and GPT-5.6 (473), both entered through their chat apps at max thinking.
  • Rube Goldberg machine: only Flash-Next got the ball into the cup. Both 27B models spent their whole \~111K-token answer budget thinking and never placed a part.
  • CAPTCHA (40 puzzles, local copy): 27B 24/40, Flash-Next 21/40, Uncensored 19/40.
  • Things you look at: Flash-Next made the best voxel castle, the best design board and the best video edit (19/20 on my rubric).

https://preview.redd.it/4e22yukvi4th1.png?width=1484&format=png&auto=…

Setup

  • GPU: one NVIDIA RTX PRO 6000 Blackwell, 96GB
  • CPU: AMD Ryzen 9 9950X
  • System RAM: 96GB DDR5
  • 27B and Uncensored: lmsysorg/sglang:v0.5.20, DFlash2 drafter, 262,144-token window, 4 slots
  • Flash-Next: lmsysorg/sglang:dev-qwen38-next-local, built-in MTP drafter, 262,144-token window, 1 slot. Its BFCL run used sglang:v0.5.20, like the 27Bs.
  • Sampler: the model card's thinking settings at the highest effort for the agent tests. BFCL and the needle test use the card's non-thinking settings.
  • The agent tests (SVG, video editing, voxel, design, Rube Goldberg, CAPTCHA) run inside Pi, a coding agent, with bash, read, write and edit. CAPTCHA gets only screenshots, mouse and keyboard.

One workstation, one model server at a time, and every number comes from a saved run.

1. Speed: drafters, SGLang vs vLLM, and long prompts

For the 27B I ran a speed matrix: every drafter, two engines, and four builds (three NVFP4 exports, one of them the Uncensored fine-tune, plus full-precision BF16). Each arm got its own server from a cold boot, the card's sampler and the 400W cap. "One user" is the Spec-Bench median over its 480 prompts. Engines: lmsysorg/sglang:v0.5.20 and vllm/vllm-openai:v0.29.0.

Which drafter (tok/s, one user):

|Drafter|SGLang · RadixArk NVFP4|vLLM · Inferact NVFP4|
|:-|:-|:-|
|none|75|59|
|MTP (built into the model)|160|113|
|DSpark|174|137|
|DFlash2|210|160|
|DFlash2 + torch.compile|214|not run|

DFlash2 wins on both engines. It keeps about 3.7 drafted tokens per step, against 2.9 for MTP.

Which build, on which engine (tok/s):

|Build · engine|DFlash2, 1 user|No drafter, 1 user|DFlash2, 4 users (total)|
|:-|:-|:-|:-|
|RadixArk NVFP4 · SGLang|210|75|607|
|Inferact NVFP4 · vLLM|160|59|517|
|Uncensored NVFP4 · SGLang|146|46|473|
|Uncensored NVFP4 · vLLM|169|63|538|
|BF16 · SGLang (full precision)|97|29|291|

  • The engine gap depends on the build. RadixArk's export on SGLang was the fastest arm overall, but on the one export I ran on both engines (the Uncensored), vLLM was 16% faster.
  • The NVFP4 exports aren't interchangeable. Same architecture, same 4 bits, same engine (SGLang), same drafter: RadixArk's export ran 210 tok/s and the Uncensored one 146.
  • 4-bit vs full precision: NVFP4 with DFlash2 is 2.2× the BF16 speed.
  • Prefill doesn't care about the engine: a full 245K-token window took 96–103s on every NVFP4 arm, SGLang or vLLM. BF16 took 129–135s.

https://preview.redd.it/vnd83h5zi4th1.png?width=1484&format=png&auto=…

https://preview.redd.it/0v1unh5zi4th1.png?width=1484&format=png&auto=…

The three models:

|Metric|Qwen3.8-27B|27B-Uncensored|Flash-Next|
|:-|:-|:-|:-|
|Prefill, full window|97s|99s|22.4s|
|Decode, Spec-Bench, one user|210 tok/s|146 tok/s|not run|
|Drafter vs no drafter|2.8×|3.2×|n/a|

The 27B keeps writing at 223 tok/s with a full 245K-token window behind it. Speculative decoding depends a lot on the content: maths ran at 339 tok/s, roleplay at 149. The pattern was the same for both 27B builds.

One important caveat. Flash-Next's speed test ran on 2026-09-12 at a 600W power cap. I later moved the card to 400W, and the 27B matrix ran at that cap. In my power sweep, prefill lost about 6% per 50W removed, so some of the gap is the cap. Moe also helps

https://preview.redd.it/jx28sde1j4th1.png?width=1484&format=png&auto=…

2. Tool use: the dense 27B leads

This is a 900-case BFCL v4 subset (11 categories), not the full leaderboard. Thinking was off, with temperature 0.7 and top\_p 0.8 from the card.

|Metric|Qwen3.8-27B|27B-Uncensored|Flash-Next|
|:-|:-|:-|:-|
|BFCL core|73.3%|70.8%|64.5%|
|Tool accuracy|87.8%|88.0%|82.4%|
|Abstention|79.5%|69.5%|68.5%|
|Multi-turn|52.5%|55.0%|42.5%|
|Malformed calls|0.08%|0.27%|0.28%|

These aren't comparable with my last post's Flash-Next BFCL numbers, which used temperature 0.

https://preview.redd.it/yqdc2j32j4th1.png?width=3396&format=png&auto=…

3. Long context: perfect for all three

I hid a fact in a log file that filled 33%, 66% or 99% of the 262K window, at three depths, with three needle types. The cache was flushed before every request.

  • 27B: 27/27
  • Uncensored: 27/27
  • Flash-Next: 81/81 (three samples per cell instead of one as I run this at the beginning)

The largest prompt was about 259.5K tokens.

https://preview.redd.it/fgeuq733j4th1.png?width=3396&format=png&auto=…

4. Battle arena: the local 27B beat Claude

Each model gets a rules sheet and a 1,000-point budget. In the open arena it designs one army blind and fights 13 armies: nine historical references plus the other entries. The score is 1,000 × its average win rate. Every matchup is 200 deterministic battles (100 seeds, sides swapped).

|Rank|Entry|Score|
|:-|:-|:-|
|1|RadixArk/Qwen3.8-27B-NVFP4|700|
|2|Claude Fable 5.1 (chat, max thinking)|678|
|3|Qwen3.8-27B-Uncensored|603|
|4|Qwen3.8-Flash-Next|535|
|5|GPT-5.6 (chat, ultra thinking)|473|

In the gauntlet, the model sees each enemy and builds a counter. The 27B beat 12/15, the Uncensored 12/15 and Flash-Next 13/17. Flash-Next ran an earlier version of the gauntlet with two more enemies, so treat that row as close, not ranked.

Thinking cost: the 27B's arena army took 39K thinking tokens in 4 minutes. Flash-Next's took 73K in 9 minutes.

https://preview.redd.it/1wbb6p28j4th1.png?width=3396&format=png&auto=…

https://preview.redd.it/o2yccq77j4th1.png?width=1484&format=png&auto=…

5. The SVG test is also a fact check

Prompt: find out which card local-AI hobbyists run and which current open model fits it, then draw the card lifting the model, labelled with a quant and a size that fit. All three picked the RTX 3090. I checked every label against what each session actually fetched.

  • 27B: 5/5 facts correct. Qwen3-Coder-30B-A3B at Q5\_K\_M, 21.73 GB, the real file size. Q6\_K at 25.09 GB is correctly marked as not fitting.
  • Flash-Next: 4/5. It got all four file sizes right and the exact 3.3B active parameters, but labelled the 3090 with a "12VHPWR, melted once" joke. That's the wrong card.
  • Uncensored: 3/5. It labelled the model "QWEN3.8-27B" but used the file size of Qwen3.6-27B Q4\_K\_M, and its "12 tok/s" isn't in anything it fetched.

All three passed 8/8 format checks. The 27B looked at its render twice and Flash-Next three times, where the rule allows one look.

https://preview.redd.it/oijy66u9j4th1.png?width=1920&format=png&auto=…

6. Video editing, voxel and design

Video editing: the model gets a raw 132-second take with fillers, a retake and a swear. It never sees the footage, only transcription and silence-detection tools, and then edits through FableCut's tools. There are two cases, each scored by hand out of 10:

  • Flash-Next 9 + 10 = 19
  • 27B 8 + 8 = 16
  • Uncensored 5 + 8 = 13

Voxel (Wawel Castle in three.js), ranked by eye:

  1. Flash-Next is the only one with the gold Sigismund Chapel dome and the Vistula bending around the hill.
  2. 27B built a clean but generic castle.
  3. Uncensored placed the camera inside its own build.

Flash-Next also used the fewest thinking tokens there: 73K, against 101K for the 27B.

Design (an animated explainer board in my design system), ranked by eye: Flash-Next first, and the two 27Bs shared second. All three passed 7/7 hard rules.

https://preview.redd.it/ue511ppaj4th1.png?width=1920&format=png&auto=…

7. Rube Goldberg: only one machine

The setup is a fixed level: a ball on a ledge, a cup on the floor and a wall in between. The model writes a parts list (no code), and a 2D physics engine runs it. It can run and look as often as it likes within 90 minutes. The score is automatic: does the ball itself end in the cup?

Flash-Next: yes, at 12.4s. It made 49 simulator runs and used 40 parts (35 of them moved). The ball travelled 1,672 px. It used 224K thinking tokens and compacted its context 8 times.

27B and Uncensored: no machine. Both spent about 111K tokens thinking in their first answer, reached the per-answer limit and stopped before writing a single part. Everyone got the same rules and one attempt. A rule that let a model continue after hitting the limit might change this, and I haven't tested that yet.

https://preview.redd.it/px8mlhhbj4th1.png?width=1280&format=png&auto=…

8. CAPTCHA: local models in a real browser

I used Open CaptchaWorld (20 CAPTCHA types, two of each). The model only sees screenshots and only acts with the mouse and keyboard. The site's own checker marks the first answer, and it must arrive within 7 minutes.

|Model|Solved|Median time|Thinking tokens, all 40|
|:-|:-|:-|:-|
|Qwen3.8-27B|24/40|55s|456K|
|Flash-Next|21/40|145s|1.0M|
|27B-Uncensored|19/40|36s|401K|

With its own 20-minute limit, Flash-Next solved 23/40. The paper reports 93.3% for humans and 40% for the best agent on its full set, which isn't the same 40 puzzles. With one run each, a three-puzzle gap is not a strong signal.

Which one should you run?

  • RadixArk/Qwen3.8-27B-NVFP4 for agents and tool calls. It won BFCL, the arena, the SVG fact check and CAPTCHA.
  • RadixArk/Qwen3.8-Flash-Next-NVFP4 for long prompts and building things, especially visual ones. It reads a full window much faster and won video editing, voxel, design and the Rube Goldberg machine.
  • orcarouter/Qwen3.8-27B-Uncensored-NVFP4 only if refusals are your actual problem. It won nothing here and invented facts in the SVG.

Resources

Configs, Docker setup and reports

show remaining 403 characters

The test harness is still private while it's changing.

Full video with the battles, the castles and the ball run: https://youtu.be/VOtfja\_Toj4**

I abused AI to help write this up and to check it against the report. Every number above comes from a saved run.

Which test do you find most interesting and maybe you have some other creative ideas how to test models?

💬 8 (+1) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/Potential-Net-9375 · 7d ago
200 Task Custom Dataset Performance Result: 10 Popular Models, from 2B MoE to 27B Dense

https://preview.redd.it/8eny8jk3b4th1.png?width=1550&format=png&auto=…

Results are within the screenshot, but here's a TL;DR tierlist:

S tier - Gemma-4-26B, even quantized down to iq3s it tops my charts.
A tier - Qwen3.8-27B-q3/q5, this one really surprised me, as a dense model it crawls, but didn't snag the S tier slot. Somehow, qwen3.5-9b is also in this slot.
B tier - Qwen3.5-4b, also incredibly Gemma-4-e2b, which punches far above its weight.
C tier - Gemma-4-12B, Nanbeige, these both are too heavy for their performance, pass.
F tier - Ling-3.0-tiny, minicpm,

The test questions consisted on tasks that I do every day with my assistants, written by Fable 4.1. "Hive" is the llm cluster I'm working on, involving custom tools and executables called by the models for different functions. Calling (or miscalling) these is important, and running a heavier model than necessary hurt, so here we are, trying to figure out the best of both worlds.

Anyway, I thought this was interesting. Hopefully you do too! YMMV.

💬 18 (+2) open on reddit ↗
▲
220
+56
55👁
r/LocalLLaMA · u/jacek2023 · 7d ago
microsoft/FrogNano-4B-2609 · Hugging Face

An agentic model from Microsoft for the GPU poor

https://huggingface.co/bartowski/FrogNano-4B-2609-GGUF

FrogNano is derived from Qwen/Qwen3.5-4B, a general-purpose post-trained model designed for language, reasoning, coding, agentic, and multimodal tasks. FrogNano inherits Qwen3.5-4B's dense 32-layer hybrid Gated DeltaNet and gated-attention architecture, but its additional post-training is text-only and focused on repository-level software engineering. The model is further trained using reinforcement learning on approximately 1,500 synthetic SWE task environments generated and calibrated against the evolving policy using TaskPilot. Training uses the five-tool Leaf harness and executable test-based rewards over complete multi-turn coding trajectories.

The additional post-training is intended to improve long-horizon repository navigation, debugging, code editing, test execution, and patch generation in a compact 4B model. Unlike approaches based on behavioral distillation, FrogNano does not train on stronger-model solution trajectories, actions, reasoning traces, or patch targets. This specialization also introduces limitations and risks: performance is sensitive to the Leaf harness and test quality, training data are Python-heavy and primarily English, and generated patches may be incorrect or insecure despite passing available tests. When integrated with the Leaf harness, FrogNano generates structured tool calls that can propose repository changes. Leaf executes authorized tool calls within an isolated repository environment to produce a candidate patch; FrogNano does not itself deploy the changes. Any resulting patches require human review, regression testing, and security validation before use or deployment.

💬 69 (+14) open on reddit ↗
▲
5
+2
17👁
r/LocalLLaMA · u/Fz1zz · 7d ago
QFN at 262K on 32 GB RAM 48GB VRAM : 1,800 tok/s prompt, 130 tok/s decod three Strata patches.

Qwen3.8-Flash-Next (IQ3\_XXS, 76 GB) at the full 262K context on 31 GB of RAM: 1,800 tok/s prompt, 130 tok/s decode, on a 5090 + a 4070 Ti SUPER on a PCIe x1 slot

Strata (github.com/Niko1221/Strata) streams MoE experts from an mmap'd GGUF, and its own sizing rule says RAM >= expert shard + 10 GB, so 57 GB for the ISTA GSQ-RCO IQ3\_XXS. I run it on 31 GB and it is fast now. Hardware: RTX 5090 32 GB + RTX 4070 Ti SUPER 16 GB (the small card sits on a chipset x1 slot, 0.8 GB/s), i7-14700K, one NVMe.

Stock 0.1.33 with a layer split at 36: 80K-token prompt 890 tok/s with the 4070 at 100% and the 5090 idle, decode \~110 tok/s, and switching between two chats re-reads the other one (30K tokens = 48 s).

Three patches on v0.1.33 (repo below, they apply to a pristine checkout):

  1. Prompts run entirely on the big card (port of Strata PR #269), the small card gets its layers' state copied afterwards, only the cells in use. Conversation parking works with the split, so alternating between a phone and a desktop session takes 0.6-1.2 s instead of 20-48 s.
  1. The real bottleneck on a box with less RAM than the expert file: the expert pool copied each 2 MB expert out of the mmap with memcpy after a MADV\_WILLNEED hint. Under memory pressure the kernel drops that readahead and you pay one major page fault per 4 KB, hundreds per expert, while both GPUs wait. A pread per slice instead: major faults per benchmark run went from 24 million to 14 thousand.
  1. Bigger prompt chunks (--prefill auto:32768). Every chunk re-streams every layer's experts, so four times fewer chunks matters a lot when the working set does not fit in the page cache.

Numbers on the final config (int8 KV, 32K cells resident per layer, MTP spec 4, vision on, 262,144 context):

\- 80K fresh prompt: 1,796 tok/s (45 s); 16K: 985; 2K: \~350

\- follow-up turn on an 80K conversation: 6 s

\- decode: 128-134 tok/s median on real sampling (0.6 / 0.95 / 20) with --spec-min-p 0.8, 95-108 greedy

\- 40K needle + follow-ups and a two-conversation parking test all correct

\- VRAM: 5090 at 32.0 GB, 4070 at 15.7 GB; RAM: the engine \~5 GB, the rest page cache

Repo with the patches, install script, launcher, benchmark tools and all measurements: https://github.com/ExTV/strata-5090-4070

▲
105
+52
74👁
r/LocalLLaMA · u/basnijholt · 7d ago
Self-hosting AI does not save money, and I do it anyway

Hi folks, I'm a long-time lurker and big fan of this subreddit and a massive self-hosting fan (also outside of AI).

I doubt many people will disagree with me here because I see the same arguments being made in many posts. However, I thought it might be interesting to share anyway. I wrote down why self-hosting AI does not save money: https://www.nijho.lt/post/self-hosting-ai-is-not-cheaper/

EDIT: didn't think this would be so controversial 😅 I do say explicitly in my blog post "I would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not".

EDIT 2: Comparing $200 sub with Opus 5.5 or Astra with Qwen 3.8 27B is not apples to apples.

💬 194 (+42) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/VerityAISolutions · 7d ago
I built an OpenAI-compatible server that runs Gemma 4 E4B on a Pixel 10 Pro XL (Tensor G5) — fully offline, ~11 tok/s decode, Tailscale-encrypted option, Apache 2.0

What it is: an Android app (PixelUnlockGPU) that turns a Pixel into an OpenAI-compatible HTTP server. Standard \/v1/chat/completions\ with streaming, so it talks to TypingMind or any OpenAI client directly — no cloud, no subscription, model runs entirely on-device via LiteRT-LM.

Device: Pixel 10 Pro XL (Tensor G5, 16 GB shared LPDDR). Model: Gemma 4 E4B instruct, GPU bundle, 2.97 GB, SHA-256 verified on download. Context window capped at 32k.

Measured numbers (not benchmarks — real on-device measurements):

\- \~11 tok/s steady-state decode (first-token-to-last over a \~300-word generation)

\- Follow-up turns in \~1.7 s: the server auto-reuses the KV prefix across turns, so stateless clients like TypingMind don't re-prefill history

\- Engine warm build \~12 s once per model change (visible in-app, split out of the metrics on purpose)

\- Short replies read slower than 11 tok/s because warm + prefill dominate the window — the UI separates decode tok/s from prefill ms so nobody has to guess

Security/access: three independent modes — loopback only, raw LAN (unencrypted), or Tailscale (binds the CGNAT tailnet IP; WireGuard end-to-end from e.g. a laptop on the same tailnet; degrades to loopback if the VPN drops). Verified with a real chat completion over the tailnet from a MacBook.

Honest limitations:

\- NPU path aborts on stock Tensor G5 firmware — GPU is the shipping backend (documented with the full investigation)

\- Android blocks named GPU temp sensors for normal apps, so the in-app gauge shows OS thermal headroom instead of °C

\- No real token counts anywhere — LiteRT-LM exposes none, so usage is estimated at \~4 chars/token and labeled as such

\- \stop\ sequences rejected explicitly (engine has no per-request stop API); client-side emulation is a filed issue

Why not llama.cpp/ollama on the phone: wanted the official LiteRT-LM GPU path on Tensor specifically, an always-on Android service (Ktor/Netty), and the OpenAI wire so existing clients just work. Happy to add a GGML backend if people want it — the engine layer is abstracted.

Built on, with credit: server/inference foundation derived from mlnomadpy/localllm (Apache 2.0); Tensor G5 runtime knowledge and the prebuilt dispatch lib from jegly/Box (Apache 2.0). Full attribution in NOTICE + per-file headers. Apache 2.0, contributions welcome — there are labeled good-first-issues (usage block, stop-sequence emulation, docs).

Repo + APKs (v0.1.0/v0.1.1 on releases): https://github.com/cannitellinicholas-spec/PixelUnlockGPU

Happy to answer anything about Tensor G5 quirks — I've done more Gate-2 debugging than I planned to.

💬 2 (+1) open on reddit ↗
▲
0
-1
5👁
r/LocalLLaMA · u/Iory1998 · 7d ago
[Help] What is the Best Context Extending App or Plugin you Recommend?

I like to use Deepseek Harness as my vibe coding harness. It's great and support local models. The issue is that most models I can run locally have context size of about 262K. Therefore, for long coding sessions, I need a memory management tool. DSH comes with a context compaction tool that I can run manually. The issue is that compaction starts to fail after a few rounds.

So, looking at DHS market place, I came across this plugin called Billion Context (https://github.com/ranxianglei/billion-context/blob/master/paper/model-driven…). The claim is I can use have long sessions. The issue is that it's a heavy context compression skill that keeps nudging the LLM to compact every few turns, which takes 5-10 minutes of work, significantly extending a normal coding session. Worse, after God knows how many rounds, the LLM seems to spend most of its time unpacking the compressed context, which fills its working context, which leads the model to compress again the text. This ended up with the LLM looping.

So, what plugins do you use with DSH or your favorite harness? What tips or tricks could you share? I am aware I can use sub-agent to work on a specific task and return a summary to the orchestrator. That helps, but I still need to manage the context window for the main agent too.

If it's not clear by now, memory is the one area I think resources must go to by they don't. I don't think context compaction is the solution. I hate it with every fiber in my body.

💬 4 (+1) open on reddit ↗
▲
2
-1
6👁
r/LocalLLaMA · u/7dollarbooks_dev · 7d ago
−2 logit bias on Bonsai 2 27B: 44/50 → 43/50 on MATH-500, +3% tokens

A recent post here reported that a −2 logit bias on "wait", "maybe", and "perhaps" made Qwen3.5-4B more accurate and shorter on 50 MATH-500 questions. I tried it on Ternary Bonsai 2 27B (PTQ1_0, 5.53 GiB) on an RTX 5060 Laptop 8 GB under Windows, Prism llama.cpp build adfffbe41. It went the other way.

| Run | Correct | Avg tokens | tok/s | Truncations |
|---|---:|---:|---:|---:|
| A baseline | 44/50 | 845.2 | 29.27 | 2 |
| B −2 bias | 43/50 | 872.3 | 29.28 | 2 |

Both runs used temp 0, seed 42, a 2048 reasoning budget, a 3072 token cap, and the same 50 questions. Run B biased nine token ids covering the lowercase, leading-space, and capitalized forms of each word. 47 answers matched; one truncated miss became correct, and two correct answers became misses.

One deterministic pair at temp 0, so treat it as one data point, not proof either way. Everything is in the repo, including every raw reply: https://github.com/7dollarbooks/bonsai2-logit-bias-test

Run by Joseph Murray Adams.

💬 7 (+3) open on reddit ↗
▲
11
+6
27👁
r/LocalLLaMA · u/GodComplecs · 7d ago
Should we plead opensource labs to still produce great non thinking (instruct) models?

The results are in, no thinking / instruct mode for new models degrade performance more than on old models such as 3.6 vs 3.8, where 3.6 takes the lead on several coding benches in instruct mode.

I would ask the labs to still nicely focus on instruct mode also still, there a probably gains to be had without the lengthy reasoning still, some of us still use models for everything and they do not need long reasoning traces. Agentic is fine and all but to start SACRIFICING performance for the "base" model which we are used to from early Llama days is not a good direction imo.

💬 20 (+4) open on reddit ↗
▲
271
+75
43👁
r/LocalLLaMA · u/Recoil42 · 7d ago
New Architecture from Percepta: Spotlight — separating intelligence from memory, allowing knowledge and skills to grow without changing the model's weights. post image

https://www.percepta.ai/blog/can-llms-grow-their-own-capabilities https://www.percepta.ai/blog/spotlight-memory "Our new architecture, Spotlight, replaces attention with a memory that escapes this trade-off: it is the first architecture to achieve infinitely growing memory without increasing the access cost. Every token reads from and writes to an unbounded memory, but because the model learns to index individual memory cells, each token only touches a small number at a time. While other sparse architectures fix the fraction of capacity used at each step—a mixture-of-experts model, for instance, always activates the same number of experts out of a fixed set—Spotlight is arbitrarily sparse, touching the same number of cells regardless of how the memory grows. The fraction of memory it uses can shrink as far as we want. Spotlight separates an intelligence module, which performs computation, from memory, which holds knowledge, procedures, and working state. The intelligence module stays the same size, and the weights don't change as memory grows. The memory is writable, and the model itself decides what to load and when to overwrite it, token by token. Because memory can hold skills as well as facts, the model can gain new capabilities without retraining: what it can do is not limited by the size of its intelligence module."

💬 32 (+4) open on reddit ↗
▲
2328
+1174
115👁
r/LocalLLaMA · u/StayLameBro · 7d ago
I made my iPhone a second GPU for my 24 GB MacBook: Qwen 3.8 27B prefills 29–44% faster & my holds part of the CTX window. post image

\*\*DISCLAIMER\*\* THE PREFILLING TPS SHOWN ON THE PHONE IS COMPUTED ONLY FOR THE LAYERS IT HOLDS. ALREADY FIXING IT TO SHOW END-TO-END PREFILL RATE. NUMBERS BELOW ARE ACCURATE FOR E2E PREFILL RATE.

Every file or tool result my agent reads on a 24 GB M4 Pro MacBook is a wait, and 64k of 8-bit context is all that fits next to Qwen 3.8 27B (IQ4\_XS), even with the wired limit raised to 20480. An iPhone 17 Pro Max was sitting in my pocket, so I figured what can I do to make use of this extra silicon.

Turns out a 10 Gb/s USB-C cable & some software is all you need. The Mac runs layers 1–40 of each 256-token batch and streams the activations to the phone. The phone runs layers 41–64 on its GPU while the Mac starts the next batch. The A19 Pro's GPU has matrix units (Metal 4 tensor ops), and they make the phone's half 2.4x faster than the same phone without them.

Same build, phone off vs. on, prefilling a 2,000-token file into a saved agent session:

  • 8k context: Mac alone 132 tok/s → Mac + iPhone 177 tok/s (+35%) (measured two days earlier, same bench)
  • 16k context: Mac alone 109 tok/s → Mac + iPhone 157 tok/s (+44%)
  • 32k context: Mac alone 101 tok/s → Mac + iPhone 130 tok/s (+29%)
  • 48k context: Mac alone 87 tok/s → Mac + iPhone 113 tok/s (+30%)

A fresh 27k-token agent session, cold: 245 s on stock llama.cpp, 228 s on my fork with the Mac alone, and 168 s with the phone.

Past 64k the phone switches jobs. The oldest KV pages move to the phone and the Mac runs all 64 layers. For every attention layer, the phone computes attention over the old keys on its GPU, and the Mac merges that with its own part. While writing, the phone's Neural Engine takes part of that work too: each 16k-key page of old context is compiled into a Neural Engine model with the keys as its weights. At 140k that took writing from 279 to 176 ms per token compared with the phone's GPU alone.

The server allocates 196k–229k of 8-bit context based on the phone's free memory; that's up to \~5.7 GB of KV cache living on the phone instead of the Mac, so the Mac's memory use stops growing at 64k. I've tested a growing session to 128k at 8-bit, with 3/3 planted facts recalled. Separately, at 140k in 4-bit, the run passed the gate with greedy output matching the Mac-only run for 32 generated tokens.

What it doesn't do: speed up writing below 64k. That's the Mac's job. My fork's kernels (SME2 on the M4 CPU and Metal fusions) plus DFlash2 speculative decoding take it from 11.3 tok/s on stock llama.cpp to 25 tok/s at about 30k context with medium thinking, phone or not. SME2 also adds up to 29% to prefill on the Mac alone. Past 64k the phone does share the writing (attention over the old keys), and without it the Mac would have to drop to 4-bit context to reach 128k. In real use I have seen upwards of 30 TPS at lower context.

The phone joins prefills over about 512 tokens. In one real session, that was 7 of 36 requests, but about 83% of the tokens read. Past 64k it holds the context and does the old-key attention, but it stops running layers 41–64 there for now; doing both is next. One request at a time.

I'm curious what this setup could do with newer model architectures. DeepSeek V4.1-Flash reports 890 bytes per token for its global KV cache and adds n-gram embedding tables (Engram). Qwen3.8-Flash-Next, the Qwen 4 architecture preview, has an n-gram lookup table too. Those aren't features of the 27B model I tested, and I haven't benchmarked either architecture here. The real gold is within the newer phones and models working together. With the A20 Pro in the iPhone 18 Pro Max, I bet there is a lot more for me to push.

Code, setup and bench scripts: https://github.com/StayLameBro/backburner

Still a lot of work to do but I built this with Opus 5.5. Happy to answer anything.

💬 333 (+107) open on reddit ↗
▲
29
+11
32👁
r/LocalLLaMA · u/jjusko20 · 7d ago
Update #3: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch

Last update for those who may be following: https://www.reddit.com/r/LocalLLaMA/comments/1wv5h8x/update\_2\_post\_training\_yandexaliceai80ba3b/

I screwed up guys 😂

Turns out my loss curve during my last run was legitimately unhealthy - as some of you, and myself, were concerned about. After evaluating my QLoRA, I found zero'd gradients in all but two layers. Turns out I had a NaN issue related to my custom v100 kernels that I didn't catch - so that run is cooked, I had to restart. I guess two layers training managed to emulate a loss curve I could at least derive a sensible explanation for until I actually got to evaluate.

Thankfully, checked to make sure gradients were applying again, and restarted the run. Once again, it's live streaming at https://figure-bios-expect-cio.trycloudflare.com/

Loss curve looks much healthier this time and is making me feel more confident that this is going to be okay. Stay tuned! I'm gonna release GGUFs and a llama.cpp patch when I have a working version.

My first epoch loss curve from this run

my first epoch loss curve on the failed first run \(note the differences in scale even if the pattern looks similar\)

▲
30
+8
47👁
r/LocalLLaMA · u/Porespellar · 7d ago
RTX Spark laptops and mini desktops rumored to launch Oct 7th (24GB to 128GB variants possibly)

Basically a DGX Spark minus the ConnectX-7 ports. I’ve seen expected initial pricing from like $1800 to $2900. Not sure what configurations are actually at those price points.

It’s all Internet hearsay until we actually see these things ship, but it’s nice to know that it’s potentially around the corner next week, especially given DGX Spark’s insane price increases lately.

Sadly, you can’t cluster them, but getting an entire computer + GB10 equivalent chip for a little over half the price of a used 4090 seems like an ok deal in this market.

💬 51 (+19) open on reddit ↗
▲
19
+11
29👁
r/LocalLLaMA · u/Apprehensive_Side219 · 7d ago
Alternative to Nvidia spark?

I just spent the last two weeks trying to catch a microcenter in stock with a spark and then today the price went up 30% and now I can't realistically afford it. It was already pretty close to the edge of my budget, and now I don't think I can swing 7k for what I had been expecting to pay 5 for last week. Any suggestions for alternative approaches welcome. I really don't want to wait until 2028 to get started.

💬 35 (+10) open on reddit ↗
▲
12
+1
34👁
r/LocalLLaMA · u/klieret · 7d ago
New benchmark on LMs fixing bugs before users run into them

Hi! This is a new benchmark that I created together with other researchers at Meta, Stanford, Harvard, UW.

Most benchmarks these days seem to test models to just fix a bug that I as a user already encountered. But shouldn't we expect models by now to also find bugs before anyone runs into them?

So in SWE-sweep we just hand an agent a big codebase and ask it to find & fix as many bugs as it can. We then give a score based on a hidden set of bugs that we know about in the repos. All the bugs are real-world bugs. We do a lot of filtering to make sure the bugs are actually discoverable & fixable from reading the repo alone.

https://preview.redd.it/irffy7x5v2th1.png?width=1080&format=png&auto=…

We're still expanding the leaderboard list with more local models (unfortunately it's always a big harder with funding/infra etc), but right now it seems like it's quite hard to beat Luna xhigh in terms of cost efficiency.

Also the scores are way lower than I would've expected. Some tasks are legitimately superhuman in practice (like fixing up all of numpy), but there's also lots of small repos, where I would've expected a lot more from current models.

Everything is open source (MIT license) on github and you can find paper etc. on the website.

Happy to answer questions here, also super curious what open weights models you'd recommend running next (we're working on an update next week).

💬 14 (+2) open on reddit ↗
▲
943
+297
95👁
r/LocalLLaMA · u/kvyb · 7d ago
Qwen3.8-27B-Humanlike-Chat 2.0: texts like a human, now with tool calls and better instruction following

Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since.

I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly:

incapable of producing more than a few words at a time.
single default personality which no amount of prompting can overcome
will not use tools, at all, whatsoever.
There needs to be a middle ground

They were right. The tool calls didn't actually work, and when people asked it to do something it would sometimes just say it's busy or going to bed. Very human. In a bad way.

So I spent the last three weeks on 2.0. The goal was simple: keep the voice people liked and lose the drawbacks.

What 2.0 does now

  • With no system prompt, it's a normal person texting. Not an assistant, not a catgirl.
  • Give it a character card and it becomes that person, and still texts like one.
  • Ask for a formal email, numbered steps or a proper explanation, and you get exactly that. Then it goes back to texting.
  • Don't want the lowercase texting? Tell it "from now on write in full sentences" (or put it in the system prompt) and it sticks to that until you say otherwise. v1 ignored this completely.
  • It calls tools, and it asks when something is missing instead of making it up. This is the part I'm happiest about. Ask the base model to book a flight without saying where from and it picks JFK. 2.0 asks where you're flying from.
  • It writes code and does math at roughly base-model level.

It's a colleague and a humanlike companion, not an assistant. Use it for chat, roleplay, agents or actual work.

How I trained it

v1 was plain SFT on real and synthetic conversations (139,845 messages from 1,396 conversations). That copies habits, including the bad ones.

For 2.0 I used on-policy distillation. The model writes its own replies and a teacher grades every token. There are two teachers:

  • v1 plus a hidden "text like a person" instruction, for chat and characters;
  • the plain base model, for instructions, tools and code.

The student never sees the hidden instruction, so it learns the behaviour without needing a prompt. Same 27B, a second LoRA on top, merged.

Numbers (vs the model I trained on, huihui-ai's abliterated Qwen3.8-27B; same prompts, same run, thinking off)

|Benchmark|Base (abliterated)|2.0|
|:-|:-|:-|
|IFBench (instruction types I never trained on)|37.3|43.7|
|When2Call (call, ask or refuse correctly)|48|58|
|BFCL irrelevance (don't call a tool when none fits)|60|78|
|IFEval, GSM8K, BFCL simple|81.9 / 89.1 / 97|83.5 / 89.1 / 98 (ties)|

Full chart in the images.

Where it's still worse: knowledge (MMLU-Pro 72.5 vs 78.5) and competitive code (LiveCodeBench 51 vs 56).

Is it actually more human? I built a benchmark for this, "ishuman":

  • It takes 150 fragments from unseen chats.
  • Has each model write the next message.
  • Shows a judge the real message and the model's without labels, and asks which one a person wrote.

|Model|Judge thought it was the real person (50% = can't tell)|
|:-|:-|
|Qwen3.8-27B abliterated (huihui-ai, the model I trained on)|0.3%|
|Same abliterated model + a "text like a human" system prompt|6.8%|
|Qwen3.8-27B official (unmodified, via OpenRouter)|15.1%|
|Qwen3.8-27B-Humanlike-Chat 2.0|23.5%|

So no, you can't just prompt your way there. In a separate test of 16 live multi-turn chats with invented people, 2.0 was picked over the base model 16 out of 16 times.

Links

Big thanks to everyone who left feedback last time, especially the ones who were critical. Tell me where it still sounds like an assistant.

Edit: safetensors are up for vLLM and SGLang:
GPTQ-Int4 (24 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-GPTQ-Int4
FP8 (48 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
BF16 (80 GB): https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0

💬 192 (+30) open on reddit ↗
▲
0
 
19👁
r/LocalLLaMA · u/takoulseum · 7d ago
Can we talk?

I see actually an acceleration of something which is scarring.

I use almost only local models, but we all feel now there is an excessive multiplication of inference engines/whatever you call it etc..

While everybody now has its own thing, what I really see in a deep dependance to claude and gpt.

Dude, Anthropic and Openai are the enemy of local AI but at the same time the local world is more and more relying on their models to progress wtf. Ofc it’s logic to want to use best models, but that becomes a dependency when they are always the same! The futur of localAI may look cool, but I think the reality is we participate to give more and more power to people that want to shut that down.

PS: I don’t care about opinion of people that will tell me I am parano, I remember many people were telling me models like qwen3.x have not effect on hw prices lul.

💬 25 (+2) open on reddit ↗
▲
2
+1
4👁
r/LocalLLaMA · u/OlegDoDo · 7d ago
SAGG — turning unreliable Gonka brokers into a reliable inference API (cascading failover, real data)

If you've used Gonka inference directly, you've probably noticed individual brokers aren't always consistent — one might be fast and reliable for a while, then slow down or drop requests, then recover. That's just how a decentralized network of independent nodes behaves.

SAGG takes a different approach: instead of relying on one broker and hoping it stays healthy, it holds several at once and automatically routes around whichever one is struggling at that moment. From the outside, you just get a normal, reliable API — the instability gets absorbed before it ever reaches you.

We didn't just build this and claim it works — we measured it properly, on real, sustained production traffic, two separate campaigns:

September 18 (1000 requests/line, standard prompt mix):

\- Standard line: 100% success (1000/1000 requests)

\- Super Deal line: 98.9% success (989/1000 requests)

September 30 recheck (200 requests/line, heavier prompt mix - longer context, code generation):

\- Standard line: 99.5% success (199/200 requests)

\- Super Deal line: 99.0% success (198/200 requests)

TTFT p50: \~190-490ms depending on line and load, p95 under 15s on heavier workloads.

Full methodology, raw data, and a reproduction script: github.com/privatedeskai/sagg-benchmark-data

For the technically curious: the hard part wasn't picking a backup broker — it was streaming responses specifically. Once a provider starts sending content to the client, you can't silently switch mid-stream without

▲
1
+1
8👁
r/LocalLLaMA · u/Away_Interaction6630 · 7d ago
How do you keep a local multi-agent app usable on CPU-only / low-RAM machines?

Hi everyone,

We're three final-year students at Epitech building Horus, a multi-agent assistant that runs entirely locally and offline. Our current challenge is hardware: keeping it usable on machines without a powerful GPU, without long setup times, excessive RAM/VRAM use or crashes.

Where we are today, from our last beta test:

One tester needed over 2 hours to install. The Python dependencies alone take 21–46 min, and the download is about 25 GB.

On CPU only, routing a question can take around 40 s, long enough for our WebSocket connection to drop.

\[Models we use + the smallest machine we've tested on\]

We'd love advice from anyone experienced with:

\- CPU-only LLM inference and memory-efficient loading

\- Quantization and model choice for low-end hardware

\- GPU/CPU fallback strategies

\- Hardware detection and adaptive configuration

\- Preventing resource exhaustion during setup and execution

Advice in the comments is very welcome, with no strings attached.

Looking for contributors: we also have a few small, well-scoped tasks or code reviews (about 1–4 hours), for example \[reviewing our hardware detection and model selection, or benchmarking a quantized model on a 16 GB RAM laptop\].

To be transparent: Horus is closed source and will be licensed to companies. Contributing is voluntary and unpaid. Before seeing any code, contributors sign a short confidentiality and contributor agreement, and the code they contribute becomes part of Horus. In return we offer thorough code reviews, full credit in the project and a professional reference on request.

We're not sharing code or private links publicly. If this interests you, comment below or DM me with your experience in local inference, CPU optimisation or offline apps

💬 25 (+12) open on reddit ↗
▲
5
+2
20👁
r/LocalLLaMA · u/giveen · 7d ago
GitHub - giveen/KernelOPT: Dispatch-aware agentic GPU kernel optimization

I want to share something I've been working on, and research paper from Redhat really helped.

This is my agentic GPU kernel optimizer for inference engine development

Its whole goal is to look at inference engine kernels, and using cloud models, plan, test and find improvements. Nothing is changed in your git, it provides a diff, send the diff over to your coding harness and ask it to analyze the diff.

It has a setup wizard and a run wizard which I recommend using, please give me feedback on what needs to be improved, as the AMD stuff I was unable to test, and mostly is theory at this time.

▲
0
 
17👁
r/LocalLLaMA · u/EcstaticDentist · 7d ago
I made 20 one-shot HTML5 game prompts for testing local coding models

I ended up making a list of 20 one-shot game prompts for testing local coding models and figured some of you might get a kick out of them.

They’re all built around the same constraint: the model has to make the entire game in a single \index.html\ with no external libraries, assets, APIs, or internet access.

Some are pretty simple, but a few get a lot more involved with enemy AI, procedural generation, upgrade systems, bosses, shops, physics, etc. I’ve been using them to see how different local models/harnesses actually handle a full task without a bunch of back-and-forth prompting.

A few of the more interesting ones are OUTBREAK, DUNGEON ZERO, TRAIN TO NOWHERE, CYBER SURVIVOR, and VOID MINER.

Here’s the full list if anyone wants to try them:

20-single-file-html5-game-prompts.md

Would actually be cool to see people run the same prompt on different models and compare what they get.

💬 24 (+7) open on reddit ↗
▲
2
+2
33👁
r/LocalLLaMA · u/Cyb3erDudu · 7d ago
shardr — like docker for models, BitTorrent sync, OpenAI-compatible serving, inference engines from upstream. post image

I kept running into the same three problems: the same 40 GB quant downloaded twice because it lived in some folder I forgot about, models quietly disappearing from Hugging Face, and every tool keeping its own copy of the weights on disk. So I've been building shardr \- a small Go daemon (Apache-2.0) that gives your machine one content-addressed store for models.

What it does, concretely:

  • Everything is stored and verified by SHA-256. Trust comes from digests, never from where the bytes came from.
  • It speaks BitTorrent v2. You can pull models from peers, and it seeds whatever you hold back into the swarm.
  • The runner starts llama-server with an OpenAI-compatible API and mmaps the weights directly out of the store — one copy on disk, no staging copies, no moving files around.
  • Runtime is pinned via a lockfile to upstream llama.cpp release binaries (never self-compiled, digest-verified), so a llama.cpp update is just a PR against that lockfile with the full test matrix behind it.

It works with pirateface.co as a catalog: shardr catalog search qwen, shardr pull <owner/repo>, and the download is anchored against the Hugging Face checksums for that exact revision — the magnet can't lie to you. Rescued models (HF source gone) pull against the catalog's recorded checksums, but only if you explicitly opt in with --trust-catalog. Other mirrors fit behind the same interface — the catalog is pluggable and the base URL is configurable.

Quick taste:

$ make all

$ shardhive serve &

$ shardr catalog search qwen2.5-0.5b

$ shardr pull unsloth/Qwen2.5-0.5B-Instruct-GGUF --quant q4_k_m

$ shardr serve unsloth/qwen2.5-0.5b:q4_k_m --id chat

$ curl http://127.0.0.1:<port>/v1/chat/completions -d '{"model":"chat",...}'

Where it stands: end-to-end works on macOS arm64 and Linux amd64, releases build themselves from CI, docs at https://cyb3rdudu.github.io/shardr. What it doesn't have: a UI, Windows support, runtimes beyond llama-server, and honestly, probably a bunch of rough edges.

What I'd like help with:

  • people with large local collections to try imports and tell me what breaks
  • feedback on the trust model — HF-anchored pulls, the explicit opt-in for rescued models. I'm sure there are holes; poke at them
  • anyone who enjoys the swarm/seeding side and wants to hack on it

Happy to answer anything about the design decisions. Docs are linked above, specs are in the repo if you want to see how the sausage is made.

💬 4 (+2) open on reddit ↗
▲
0
 
18👁
r/LocalLLaMA · u/Decent-Manager-5373 · 7d ago
Local diffusion on GGUF: I wrapped stable-diffusion.cpp in a Vulkan desktop app (FLUX Schnell / Z-Image / Wan on a 6GB laptop GPU, no CUDA) post image

I've open-sourced \*\*Vison\*\*, a desktop app for generating images and video entirely on your own GPU. No account with a generation service, no credits, no prompts leaving your computer.

Licensing details, since this sub cares:

\- Vison itself is \*\*MIT\*\*. It builds on stable-diffusion.cpp, ggml and vision.cpp (all MIT) plus others listed in a generated \THIRD-PARTY-NOTICES.txt\ that ships inside the app.

\- The bundled ffmpeg is an \*\*LGPL\*\* build with libvpx and no GPL components; the build refuses to package a GPL or non-free one. Video is VP9 in WebM, which is royalty-free. That was a deliberate licensing choice, not a technical one.

\- Model weights are \*\*not\*\* covered by the MIT licence. Each has its own terms from whoever published it (FLUX.1 Schnell, Wan and the rest all differ), so check before using output commercially.

\- No paid tier, nothing held back, and none planned.

It's early: one developer, one 6GB laptop GPU, Windows only. The backend is portable C++/Vulkan, so macOS/Linux is mostly packaging and testing rather than porting, and that's where help would matter most.

Repo: https://github.com/JayRGadekar/Vison (contributing guide, issue templates and a SECURITY.md are in there)

▲
3
+1
15👁
r/LocalLLaMA · u/Thac0-is-life · 7d ago
Help me with a better hardware setup for Local LLm

&#x200B;

Hello. I've been playing around with local LLMs for a while now, using my 7900xtx. I understand the concepts and usually what to do. But I'm now in a bit of choice paralysis on where to go next.

I have a Ryzen 5600X with a 7900XTX and 64GB of DDR 4 (how I wish I had purchased more at the time..) and a motherboard (MS-7B79/X470 GAMING PRO (MS-7B79)) that is not really great for multiple GPUs (it was a gaming PC).

I would like to increase my Local LLM game. I'm running mostly qwen 3.8 27B at 3 or 4q from with 128k to 220k context with KV cache at 8q. Sometimes I get up to 50tk/s and around 750 tk/s of context ingestions. But I wanted to run bigger models/have faster speed, or at least run multiple copies of that same qwen so multiple agents can run at the same time. Or try the Qwen 3.8 Flash for example. This level of model is already awesome enough to do anything I need.

I've been thinking of purchasing 2 or 4 MI50 16GB( which costs 1/3 of the 32GB), but I don't really know the rest that I should get. Motherboards that would help me optimize the performance around that, etc.

Or should I just bite the bullet on another 7900xtx (more expensive than 4 MI50)? But I think I would still need a new motherboard at least to let me use both at the same time

I have a basement so noise is not a problem, and I have solar, so power is not a real issue (at least during summer).

What are you all suggestions here? Does it make sense to go with older GPUs like that?

▲
8
 
25👁
r/LocalLLaMA · u/W61k3r · 7d ago
Tuned/abliterated Qwen3.8-27b into a 24gb card 262k guff using the newest unreleased version of LexiPanel. It's fast with reliable draft acceptance. Made for 7900xtx but should work on whatever 24gb card with this setup and headless. Doesn't get dumber while coding like most of the other fine-tunes.

https://huggingface.co/Wa1k3r/Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit

Qwen3.8-27B CODER — IQ4_XS imatrix · 24 GB card fit · ~262k context · MTP draft

Quantized, Abliterated, and fitted by LexiPanel. Its Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k tokens of context. It built the importance matrix from code-heavy text and kept the MTP head at Q8\_0, so --spec-type draft-mtp works without a separate draft model.

The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context.

#

File

|File|Type|Size|Inside|
|:-|:-|:-|:-|
|Wa1k3r/Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf|IQ4\_XS + imatrix|18.35 GB (17.1 GiB)|MTP head at Q8\_0, token embeddings at Q4\_K|

#

Measured speed (real use, not a synthetic benchmark)

These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself:

|Context already in the window|Requests|Decode, median|Decode, range|
|:-|:-|:-|:-|
|65k – 131k tokens|42|41.2 t/s|33.8 – 46.0 t/s|
|131k – 171k tokens|56|36.0 t/s|30.7 – 43.0 t/s|

  • MTP draft acceptance: the median is 85% (the middle half of requests falls between 75% and 94%). That works out to about 2.7 tokens per decode step at draft depth 2.
  • Prefill:
  • 387 t/s for a cold 108k-token prompt;
  • 175–183 t/s for about 4.5k new tokens added at 147k–156k depth.
  • VRAM: 24.2 of 24.6 GB in use at 245,760 tokens of context, with a q4\_1 KV cache and the vision projector on the CPU.

Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot.

#

Run it with llama.cpp

llama-server -m Qwen3.8-27b-CODER-4q_xs-24GB-262k-Optimalcardfit.gguf \
-c 262144 -np 1 -ngl 99 --flash-attn on \
--cache-type-k q4_1 --cache-type-v q4_1 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \
--jinja --reasoning on --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
-b 2048 -ub 512 --cache-reuse 256

  • Context: -c 262144 is what fits next to the weights on a 24 GB card with a q4\_1 KV cache. The model's native window is 262,144 tokens. On a smaller card, lower -c first.
  • Speculative decoding: --spec-type draft-mtp drafts with the MTP layer inside this file, so no separate draft model is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235.
  • Sampling: these are Qwen's recommended settings, and they are also stored in the file.
  • Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB).
  • Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context.

#

How LexiPanel made it

  1. Conversion: the source weights were converted to a BF16 GGUF with the MTP head included.
  2. Importance matrix: computed from about 300k tokens (570 chunks) of code-heavy calibration text. Three quarters is Python source (the standard library and installed packages). The rest is technical documentation, READMEs and license texts, the kind of text a coding agent's context fills with.
  3. Quantization: llama-quantize from llama.cpp b11182 made the IQ4\_XS file with that matrix. The MTP head stays at Q8\_0 so its drafts stay accurate, and the token embeddings are Q4\_K.
  4. Fitting the card: LexiPanel's Fit planner chose the mix, quality first, for one 24 GB card at 262144 tokens of context. It took the best quality that card could afford at that context, not the smallest file.

#

Credits and license

  • Quantization, importance matrix, MTP draft setup and card fit: LexiPanel.
  • Tools: llama.cpp.

Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.

Qwen3.8-27B CODER — IQ4\_XS imatrix · 24 GB card fit · \~262k context · MTP draft

Quantized, Abliterated, and fitted by LexiPanel. Its
Fit planner chose the tensor mix to fill one 24 GB GPU at about 240k
tokens of context. It built the importance matrix from code-heavy text
and kept the MTP head at Q8\_0, so --spec-type draft-mtp works without a separate draft model.
The uploader runs it every day for agentic coding. On an RX 7900 XTX it decodes at 36–41 tokens/s with 110k–170k tokens already in context.

File

File Type Size Inside
Wa1k3r/Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit.gguf IQ4\_XS + imatrix 18.35 GB (17.1 GiB) MTP head at Q8\_0, token embeddings at Q4\_K

Measured speed (real use, not a synthetic benchmark)

These are 98 requests from agentic coding sessions on 2026-10-01, timed by the server itself:

Context already in the window Requests Decode, median Decode, range
65k – 131k tokens 42 41.2 t/s 33.8 – 46.0 t/s
131k – 171k tokens 56 36.0 t/s 30.7 – 43.0 t/s

MTP draft acceptance: the median is 85% (the middle
half of requests falls between 75% and 94%). That works out to about
2.7 tokens per decode step at draft depth 2.
Prefill:
387 t/s for a cold 108k-token prompt;
175–183 t/s for about 4.5k new tokens added at 147k–156k depth.

VRAM: 24.2 of 24.6 GB in use at 262144 tokens of context, with a q4\_1 KV cache and the vision projector on the CPU.
Setup: AMD Radeon RX 7900 XTX (24 GB), llama.cpp b11235 on the Vulkan backend, all layers on the GPU, one slot.

Run it with llama.cpp

llama-server -m Qwen3.8-27b-CODER-4q\_xs-24GB-262k-Optimalcardfit.gguf \\
\-c 262144 -np 1 -ngl 99 --flash-attn on \\
\--cache-type-k q4\_1 --cache-type-v q4\_1 \\
\--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0 \\
\--jinja --reasoning on --reasoning-format deepseek \\
\--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \\
\-b 2048 -ub 512 --cache-reuse 256

Context: -c 262144 is what fits next
to the weights on a 24 GB card with a q4\_1 KV cache. The model's native
window is 262,144 tokens. On a smaller card, lower -c first.
Speculative decoding: --spec-type draft-mtp
drafts with the MTP layer inside this file, so no separate draft model
is needed. You need a recent llama.cpp that supports the qwen35 architecture and draft-mtp; this card was measured on b11235.
Sampling: these are Qwen's recommended settings, and they are also stored in the file.
Agent loops: a host-RAM prompt cache helps, for example --cache-ram 5631 (MiB).
Vision: this repo holds the language model only. Add a Qwen3.8-27B mmproj with --mmproj mmproj-F16.gguf --no-mmproj-offload. The projector then runs on the CPU and the VRAM stays with the context.

How LexiPanel made it

Conversion: the source weights were converted to a BF16 GGUF with the MTP head included.
Importance matrix: computed from about 300k tokens
(570 chunks) of code-heavy calibration text. Three quarters is Python
source (the standard library and installed packages). The rest is
technical documentation, READMEs and license texts, the kind of text a
coding agent's context fills with.
Quantization: llama-quantize from
llama.cpp b11182 made the IQ4\_XS file with that matrix. The MTP head
stays at Q8\_0 so its drafts stay accurate, and the token embeddings are
Q4\_K.
Fitting the card: LexiPanel's Fit planner chose the
mix, quality first, for one 24 GB card at 262144 tokens of context. It
took the best quality that card could afford at that context, not the
smallest file.

Credits and license

Quantization, importance matrix, MTP draft setup and card fit: LexiPanel.
Tools: llama.cpp.
Licensed Apache-2.0. The weights have reduced refusals ("uncensored"), so how you use them is your responsibility.

💬 25 (+4) open on reddit ↗
▲
5
+1
12👁
r/LocalLLaMA · u/No_Contract_8296 · 7d ago
CalDec v1 - Fully Open Decision Model for Personal Assistants

Somebody just released a fully open-source, open-weights decision model that beats Jev!

Just kidding, it's me and I this is my first time releasing a public model, recipe and dataset so I am looking forward to learning from the experience.

Jev is indeed a very powerful and inexpensive model and obviously a much better all-rounder, and some of my checkpoints did in fact score better on some tests (namely LocalLLaMA/typed-decisions and the internal test set) but that doesn't mean it "beats Jev" of course.

The motivation for this was a quick experiment to see how far behind Jev open-weights models like Laya are, and how much closer I can bring them with a small dataset and fine-tuning. The results were better than expected especially for me since I do not have professional ML experience.

For my use-case - a Jarvis-like personal assistant which aims to be real-time and fully-local - this model proved to be genuinely useful for certain aspects of that project so I decided to share the results and how I got there. Going local also means privacy and eliminating network latency.

I hope some of you find this experiment valuable or useful in some way.

I would also love to hear you suggestions, criticism or just discuss the approach!

Dataset: https://huggingface.co/datasets/kgrozdanovski/assistant-decisions**
CalDec Laya: https://huggingface.co/kgrozdanovski/caldec-v1-laya**
CalDec GLiNER: https://huggingface.co/kgrozdanovski/caldec-v1-gliner2.5-decide**
GitHub: https://github.com/kgrozdanovski/caldec**

▲
138
+26
47👁
r/LocalLLaMA · u/matteiuspi · 7d ago
Two 96 GB Ascend cards crun Qwen3.8-flash-next hardware notes, vLLM work, benchmarks, and what is next

I have been building a somewhat unusual local inference machine around two Huawei Atlas 300I Duo cards. They are relatively inexpensive, passive, dual-accelerator PCIe cards with 96 GB of device memory apiece. They are also absolutely not drop-in CUDA replacements.

When I first brought up Qwen3.8 Flash-Next these past two weeks, it was often incoherent and lived around 1 generated token per second. Some runs were below that. Today the same two-card machine is producing coherent output at roughly 30 tok/s for one request and about 61 tok/s aggregate at four-way concurrency on my short decode benchmark. It also completed the full 198-question GPQA Diamond set.

This post is the start of a guide for these cards: what the cards physically are, how I cool them, what “96 GB” really means, what I changed in vLLM and vLLM Ascend, which optimizations actually mattered, and which problems are still open.

The short version is the hardware is capable. My work has been mostly on the software stack, ubuntu-26.04 driver support, model architecture support, memory layout, custom operators, and getting every asynchronous state transition exactly right.

The hardware: one card is really two devices

My current machine has two Atlas 300I Duo cards, which enumerate as four Ascend 310P3 devices.

Current system / Planned system

Physical cards: 2 / 3

Ascend devices/npus/AIcpus: 4 / 6

Nameplate device memory: 192 GB / 288 GB

Approx. runtime-visible memory with this configuration: 172 GiB / 258 GiB

Combined maximum accelerator-board power: 300 W / 450 W

Each card has two accelerator SoCs and 96 GB of LPDDR4X in total, or 48 GB local to each chip. It is not one unified 96 GB allocation. A model that does not fit on one 48 GB device still needs tensor, expert, pipeline, or another form of model parallelism. The card is PCIe Gen4 x16, full-height/full-length, and a surprisingly thin single-slot design. Huawei rates it at 408 GB/s aggregate memory bandwidth and 150 W maximum board power. The official specifications are here (https://support.huawei.com/enterprise/en/doc/EDOC1100285916?section=j00e).

Also, despite the generic “HBM” terminology used by a lot of accelerator software, the memory on these cards is LPDDR4X.

This two-chip-per-card layout matters. Communication within a model still goes through the distributed runtime, and memory remains local to a rank. I use HCCL collectives and explicitly map tensor and expert ownership across all four chips. Thinking of the machine as four 48 GB ranks is much more useful than thinking of it as two 96 GB GPUs.

https://reddit.com/link/1wvt1m4/video/7rju7qbrw1th1/player

Passive cooling is not a deal-breaker

The cards have large heatsinks and no onboard fans. They were designed for server airflow, so putting them in an ordinary workstation and hoping a rear case fan will sort it out is a bad plan. There is a useful teardown here (https://videocardz.com/newz/huawei-atlas-300i-dual-ai-gpu-with-96gb-memory-worth-1400-has-been-taken-apart) if you want to see the heatsink and heat-pipe arrangement.

I give them direct, high-volume airflow and run the room on AC/heat-pump cooling. Under real multi-hour model loads, the cards can crank continuously without drama. Across my recorded Qwen runs, peak device temperatures were generally 72–78 °C. My watchdog limit is 96 °C, and the cards have not approached it.

So I would not bat an eye at adding another passive card. The actual checklist is mundane:

• Keep unobstructed airflow through the heatsink fins.

• Make sure the chassis fans have enough static pressure.

• Budget another 150 W of board power per card, plus the rest of the host.

• Exhaust the heat from the room instead of recirculating it through the rack.

• Log temperature during long prefill, decode, and concurrency tests rather than trusting an idle reading.

Passive does not mean low-power or self-cooling. It means the chassis and room are the cooling system. Once that is handled, these have behaved like ordinary 150 W server cards for me.

Note: Nothing heats up these cards more than loading/moving things around in their ram -- the npus at full utilization run cooler than large block memory assignments. We keep this in mind when optimizing the model serving code paths.

ECC, nameplate memory, and what is actually usable

My cards arrived with ECC enabled by default. I disabled it to reclaim device memory. This is an inference and development box, and I consciously prefer capacity over ECC protection here. That is a reliability tradeoff, not a universal recommendation.

Even with ECC disabled, firmware, the runtime, communication buffers, graph captures, workspaces, and allocator reservations consume memory. In practice, the software sees roughly 43 GiB per 48 GB chip. Four chips therefore provide about 172 GiB of useful aggregate capacity, but it is still four separate local pools. The exact free number also changes with the CANN build and launch configuration.

That distinction has shaped nearly every model decision. The question is not only, “Does the checkpoint total fit in 192 GB?” It is, “Does each rank's weight shard, recurrent state, KV/cache allocation, graph capture, collective workspace, and worst-case temporary allocation fit in its own 43 GiB?”

Why I forked vLLM as well as vLLM Ascend

The public work lives in the OpenSensor vLLM Ascend fork (https://github.com/opensensor/vllm-ascend) and paired vLLM fork (https://github.com/opensensor/vllm). I needed both sides because this was not just a missing device kernel.

I am currently the only person developing these forks. The software bus factor today is one. I have made a lot of progress, but a fast-moving one-person fork should not be confused with the maturity, test coverage, or support depth of mainline vLLM on NVIDIA.

I also ran into a bizarre tooling problem: in my sessions, Claude repeatedly refused to engage with prompts about this architecture because the cards are Huawei hardware from China. These were ordinary engineering discussions about serving, sharding, cooling, and performance—not requests to build a restricted application. I am describing my direct experience rather than claiming that every Claude version or account will behave identically, but it made Claude unreliable as a development assistant for this project. Whatever anyone thinks about the politics, that is a real practical constraint when choosing tools around this hardware.

Qwen3.8 Flash-Next combines MoE routing, Gated DeltaNet recurrent layers, sparse quadratic-attention layers, packed low-bit experts, long context, and an MTP draft model. Supporting that cleanly touched model integration, the v1 runner, cache accounting, scheduling, graph capture, distributed state, model loading, and Ascend-specific operators.

The current development sprint has been roughly two and a half weeks of nearly continuous bring-up and optimization, with hundreds of fork commits, repeated full checkpoint loads, profiler captures, operator microbenchmarks, and multi-hour quality runs. This was not one magic kernel patch.

What took Qwen from incoherent \~1 tok/s to where it is now

These are the architectural changes that moved the needle.

  1. Make the hybrid model correct before making it fast

The early model could generate tokens, but generation was not a correctness test. I found failures that only appeared at production geometry: incorrect Gated DeltaNet gate-vector handling, recurrent-state precision and lifecycle problems, incomplete sparse-attention score width, and mismatches between the host operator API and the installed kernel package.

One particularly nasty GDN issue looked fine in small-head tests but accumulated state error across the real 36 recurrent layers and produced incoherent text. Keeping recurrent state in FP32 and fixing the production-shaped data movement was foundational. So was treating the custom OPP package and Python host code as one ABI-versioned unit. A stale kernel can look like a model problem for a long time.

  1. Shard the model at load time instead of loading everything everywhere

The Qwen checkpoint is about 169 GiB and contains 1,610 safetensor files. It was larger than available host RAM in one of my bring-up configurations, never mind the memory on an individual NPU.

I built an expert-aware loader that reads only the experts owned by each rank, keeps dense/shared tensors where required, and avoids materializing the whole expert bank before throwing most of it away. Host-side expert data is mapped and moved lazily. This changed model loading from an accidental memory stress test into a deterministic TP4/EP4 layout.

  1. Keep low-bit weights packed and do the work on the NPU

My first usable W4 reference dequantized packed weights in the eager path and then called a regular matmul. It was useful for correctness and managed only 0.229 tok/s on one recorded smoke case.

The production path keeps the weights packed, routes tokens to experts on the device, and uses custom AscendC Cube kernels for grouped expert projections. I added FRACTAL\_NZ layouts, fused gate/up handling, tiled reductions, route and tile reuse, and dedicated W4A8 execution instead of repeatedly expanding W4 weights into a larger temporary representation.

This is both a speed win and a capacity win. Avoiding transient expanded expert banks leaves memory available for state, cache, graphs, and concurrency.

  1. Build caches for the model I actually have

Flash-Next is not a conventional all-attention transformer. My configuration has 36 recurrent GDN layers and 12 sparse QSA layers. Treating all of that as a normal dense KV cache wastes memory and misses the state semantics.

I implemented separate recurrent-state management, compact physical cache layouts, prefix-state tiers, sparse page selection, direct NZ gathers, and 310P-specific QSA paths. The service is configured for a 262,144-token context limit, and the cache planner retains capacity for about 4.08 such windows. That is a memory-planning result, not a claim that every possible four-by-262K workload has completed an end-to-end soak.

  1. Remove synchronization and launch overhead from the token loop

On these devices, a stray device-to-host scalar read can serialize the whole pipeline. I removed hot-path .item() calls, reused per-step tensors, deferred collectives behind useful work, tightened CPU affinity, and moved routing and sparse selection away from Python.

Once the eager path was correct, I added decode-only ACL graphs and MTP2 speculative decoding. I capture the actual concurrency shapes I serve rather than pretending one graph is universal. Fused operators and graph replay matter enormously when a decode step otherwise consists of many small kernel launches.

  1. Optimize the service, not just an isolated kernel

Several kernels won a microbenchmark and lost end to end. I kept the ones that reduced real request time and rejected or quarantined the others. Multi-request QSA, grouped MTP experts, expert-route caching, cache accounting, cold-prefill chunking, and collective overlap were all measured at the API boundary.

That last part is why four concurrent requests reach roughly 61 aggregate tok/s even though one request is around 30 tok/s. The extra work can occupy parts of the machine that a single token stream leaves idle.

Qwen performance today

These are milestones from different stages and workloads, not one controlled single-variable benchmark:

Qwen3.8 Flash-Next milestone / Measured result

Earliest uncontrolled service: Roughly 0.2–1.9 tok/s, often incoherent

Correct eager W4 dequant reference: 0.229 tok/s

show remaining 8,027 characters

Stable W8 service baseline: About 18–19 tok/s

Native W4A8, MTP2, graphs, one request: 29.74 tok/s median; 34.34 peak

Same optimized service, four requests: 60.88 tok/s aggregate median

Two concurrent 40K warm-prefix requests: 30.76 tok/s aggregate

40K cold prompt: About 120 seconds TTFT; 30–32 tok/s afterward

The 30/61 tok/s figures are short, fixed-output decode tests. Long reasoning requests tell a less flattering and more useful story.

https://reddit.com/link/1wvt1m4/video/rpc5c2c1s1th1/player

For quality, I ran all 198 GPQA Diamond questions on the four-chip Ascend W4 service and, as a reference, an RTX 6000 Pro running a different IQ4\_XS GGUF in llama.cpp. Both scored 140/198 (70.71%) with the same AISBench-style answer extractor. This is evidence that the Ascend path is coherent; it is not a pure hardware or quantization comparison because the runtimes, quantizations, chat templates, and concurrency differ.

On the final uninterrupted 106-case Ascend phase, four workers emitted 507,256 tokens in 2 h 52 m 46 s: 48.93 aggregate tok/s. Median per-request client rate was 12.50 tok/s on these long reasoning generations. The server completed every request in that phase with no eager fallback or zero-acceptance interval. The RTX reference was much faster per request, so I am not presenting this as an NVIDIA killer. I am presenting it as a large model working correctly and usefully on hardware that initially produced slow nonsense.

GLM is my next hard(er) model

I am also bringing up the roughly 304B-parameter GLM-5.3-Flash architecture. It combines 34 KDA linear-attention layers, 11 DSA sparse-attention layers, 288 experts, latent MLA history, mHC mixing, and a mixed W2/W4 expert checkpoint. The checkpoint is about 151.6 GiB, with approximately 35.6 GB of loaded weights per rank in one four-rank profile.

It now loads and generates on the same machine. The speed progression so far has been:

GLM four-chip milestone / One request / Four-request aggregate

Initial full-service profile / 0.764 tok/s / 1.474 tok/s

Batched NZ workspace write / 0.912 tok/s / 1.822 tok/s

NZ-packed code layout / 1.386 tok/s / 2.946 tok/s

Fused mHC/MLA work / 1.743 tok/s / 3.269 tok/s

Latest measured integrated runner / 2.236 tok/s / 4.476 tok/s

An 8,232-token GLM prompt prefills at about 39.1 prompt tok/s and then decodes at about 2.16 tok/s. Those numbers are far from my target, and the 32-token test completion was too short to establish answer quality. GLM is currently a bring-up and optimization result, not a service recommendation.

I have already found an important quantization lesson there. An early W2 checkpoint showed residual growth all the way to an RMS around 525 in the last layer. Moving the affected experts to a no-clip W4 treatment kept the network bounded. Separately, my custom blocked-dequant Cube kernel was about 17 times faster than the eager reference in its isolated test. As Qwen taught me, both numeric behavior and end-to-end integration have to pass before either result means “done.”

The third card: 50% more memory and cores, not just a spare

I plan to add a third Atlas 300I Duo to this system. That takes the machine from four to six 310P devices and from 192 GB to 288 GB of nameplate device memory. With the same ECC and runtime reservations, I expect roughly another 86 GiB of runtime-visible capacity, for approximately 258 GiB across the six ranks.

The n-card architecture is designed to use it. Expert ownership is distributed across ranks, so the two new chips add local expert capacity and accelerator cores; they are not merely passive storage. Tokens route to the ranks that own their experts, and the additional ranks participate in the model's compute and collectives. I expect useful scale from EP6, although the exact speedup will be measured rather than advertised in advance.

The extra capacity gives me several options:

• Keep larger expert sets or higher-precision layers resident.

• Fit models that are just over the four-chip limit without host offload.

• Spend more memory on long-context state and cache.

• Reduce aggressive quantization where the quality trade is not worthwhile.

• Run a large distributed model while retaining room for another smaller service or evaluation workload.

The cost is one more card, 150 W of maximum board power, two more device ranks, and another passive heatsink that needs real airflow. In the cooled room, none of those are architectural concerns. The interesting cost is communication: six ranks change route balance, collective sizes, and PCIe/HCCL traffic. I will tune and benchmark that topology, but the software is already organized around n-card expert distribution rather than hard-coded four-way ownership.

Known issues and active work

This is what is still on my bench:

• I just fixed one real cross-stream race: a prefix-Mamba state slot could be spilled or reused before its pending NPU writer completed, allowing an older checkpoint to be restored. That fix has an NPU regression test.. A later 106-case run survived 12 state spills without degrading, which is encouraging but not a root-cause proof. I am continuing long mixed-load and eviction/reuse soaks.

• Qwen cold prefill. A 40K cold prompt still takes roughly two minutes even though subsequent decode is fast. Profiling points primarily at the 12 QSA layers, especially sparse selection and tiled attention. This is now a more important target than another tiny decode micro-optimization.

• Qwen long-context qualification. The planner has the capacity, but I am separating configured context, allocated capacity, and completed end-to-end long-context tests. I want retrieval and concurrent-fill evidence, not a screenshot of a launch flag.

• GLM coherence and performance. I am requalifying the full model after KDA, MLA, QSA, and runner integration changes, then moving the grouped mixed W2/W4 expert path, cache layout, graph replay, and eventually MTP through the same correctness-first gates used for Qwen.

• GLM loading and memory. The filtered loader can skip large amounts of peer-owned or superseded checkpoint payload before tensor materialization. I saw one roughly 20% load-time improvement, but it needs controlled reruns and byte-accounting before I call it a result.

• Six-device expert parallelism. When the third card arrives, I will measure rank balance, per-card temperatures, collective time, model capacity, and c1/cN throughput on the exact six-rank topology.

• Other model adapters. The same loader, packed-expert, cache, and operator infrastructure is feeding ongoing DeepSeek and other hybrid/MoE work. I am avoiding model-name conditionals where the underlying contract can be made generic.

Should you buy one?

I plan to list my first additional card in the OpenSensor storefront (https://www.opensensor.io/) next week at $2,900. The listing is not live yet. I believe it is a good price for what the hardware can already do and what the software should unlock. It is not yet the same kind of turnkey purchase as a supported NVIDIA card running mainline vLLM.

I think the right buyer is a developer, lab, or systems-minded end user who is comfortable with both of the following:

  1. You own the airflow solution. These are passively cooled server cards. Depending on the chassis and motherboard, that may mean high-static-pressure case fans, a duct, or a 3D-printed shroud. I use Fusion 360 and am happy to help with additional shroud designs. There are too many motherboard layouts, card spacings, fan sizes, and case geometries to pretend that one printable design will fit everything.
  2. You are adopting an active development fork. I am the sole developer on the vLLM work today as upstream is focused more on their server grade accelerator modules not yet available to the US. The progress is real and the benchmarks in this post are from actual hardware, but more bugs and better approaches will be discovered. Buyers shoul
💬 118 (+6) open on reddit ↗
▲
68
+26
61👁
r/LocalLLaMA · u/Postmodern_Plunger · 7d ago
Inference Engineering for Dummies

Hi all! I am a former SWE who has recently transitioned into inference engineering. I launched a side hustle a few months back and I've just taken it full time due to excessive demand. Its been such an opportunity because the people optimizing runtime are far fewer than the people that are trying to build apps or offer AI solutions.

So I've been lurking around this community, and I've noticed a lot of people who seem to have massively suboptimal setups for their hardware, and I've grouped the biggest errors into several buckets. The purpose of this guide is to expose common inference bottlenecks and provide best practices for avoiding them within your hardware constraints.

RUNTIME:

I. Choosing the Right Runtime

This is the biggest mistake I see. Choosing the correct runtime for your architecture and model is the most important decision to make. In general, here are some rules to help you determine what runtime to use.

Firstly, Ollama is never optimal. Its just the simplest. If you want quick and easy and have extra RAM, it's a good place to start. It's very user friendly and requires less setup. But it just won't offer best inference speeds.

If your model requires cpu offload, then llama.cpp will be your best choice. If not and you're solely in gpu, VLLM will likely provide the best results. It's really as simple as that for 90% of cases. SGLang may be worth it if your workload involves Langgraph, as it is highly optimized for the tooling. Otherwise, stick to the above. Mainline branches are best, with community forks offering only highly niche performance boosts (i.e., for specific models/configurations, but are generally under optimized and not well maintained).

II. Optimizing and Maintaining Runtime

The other big mistake people make with runtime is failing to compile it with hardware specific flags. Not going to go through all of them here, Google can help you out. Just search "optimal runtime compilation flags for \[runtime\] using \[GPU, CPU, RAM type\]." The most missed/missed flags tend to be for architecture specific optimizations. Those are crucial.

Runtime should be recompiled (with optimal flags) any time \*any\* of the following occur:

\- System updates

\- Kernel/driver updates

\- Running a model released or modified later than your last compile

\- You haven't recompiled in over a month (recent updates often contain kernel or path optimizations)

MODEL CHOICE:

I. Quantization:

Quantization. Such a big word. Such little meaning. All you need to know is that it makes a model smaller. There are a million Q\_K\_X\_&$&$&$ quant sizes, so I'm not going to go over them individually. Rather, I will provide basic principles.

\- IQ quants are generally the best for the size. If you're choosing between IQ4\_XS or Q4\_K\_M, IQ4\_XS is both a smaller VRAM footprint and higher complexity.

\- Nonlinear (NL) quants are only ever going to be better if you have CPU offload. Even then, IQ quants often offer extra context space vs NL quants and thus are preferable.

\- If it's a quant you've never encountered, read the docs. It more than likely is highly optimized for that specific model and is indeed the one you should choose. Searching it or asking chatgpt \*will give you the wrong answer every single time for custom quants.\* You will only encounter these with custom tuned models.

\- Standard Q\_K\_M quants are best if you have absolutely no hardware constraints for the model you're running, as path optimizations are the best. If you have no hardware constraints, though, you could be running a better model. This is only useful for running simple models for simple tasks.

II. Task

Certain models excel at certain tasks. This is subjective and preference based but this is my list:

Coding: for API, anthropic. Hands down the best models. Opus and Sonnet 5.5 both excel in performance and their low token usage per task makes them more affordable than previous iterations. Deepseek models are the best budget choice. Qwen models are the clear winner for local inference on all fronts.

Writing: Opus/sonnet for technical writing, Gemini for creative writing. Gemma for local creative writing.

Video: Wan 2.2 for local in most cases will get it done, chatgpt and copilot both have surprisingly robust free image/video gen, Veo is the best paid.

HARDWARE:

Buy a budget box, or build your own from parts. I've managed to squeeze better performance out of an RTX 3060 and 128 gb RAM than a DGX spark across all categories for multiple models. The spark has an edge for dense models, but I was able to run higher complexity models overall on the other setup for 1/5 the price. AMD and Intel lag significantly on speed per price, but I've heard Intel has had some major gains recently. Have not confirmed myself though.

MODEL OPTIMIZATIONS:

I. Spec Decode (MTP)

\- If you have CPU offload, spec decode will \*always\* slow you down. The extra overhead compute isn't worth it if you don't have at least several hundred Mb/s bandwidth, which your CPU won't.

\- MTP is sometimes a baked in feature, and sometimes requires a special secondary model. Ensure you know how it works for your model and what flags to run.

\- Each model will be optimized for exactly 0-1 type of spec decode. Figure out which one it is (or isnt) rather than wasting your time testing methods.

II. Model Tuning

Just to show the kind of command optimization you can get, here is my sample command for running Qwen 3.8 Flash-Next, a 156b parameter model, on 12 gb VRAM (and 128 gb RAM) at 10 token/s decode and 150 token/s profile at 200k context:

\~/llama.cpp/build/bin/llama-server --flash-attn on --batch-size 1024 --ubatch-size 1024 --no-warmup --cache-reuse 256 --jinja --host 0.0.0.0 --port 8090 --presence-penalty 0.0 --repeat-penalty 1.0 -m \~/llama.cpp/LLM/Qwen3.8-Flash-IQ4\_XS/UD-IQ4\_XS/Qwen3.8-Flash-Next-UD-IQ4\_XS-00001-of-00003.gguf --temp 0.95 --top-k 20 --top-p 0.97 --min-p 0.05 -np 1 --chat-template-kwargs '{"enable\_thinking": true, "preserve\_thinking": true, "reasoning\_effort": "xhigh"}' --threads-batch 16 --threads 8 --gpu-layers 150 --n-cpu-moe 48 -c 200000 --override-tensor per\_layer\_token\_embd.weight=CPU -ctv q8\_0 -ctk q8\_0 --cache-ram 8192 --checkpoint-min-step 512 --ctx-checkpoints 4 --kv-unified --reasoning-preserve --load-mode mmap+mlock

That's a lot, right? It's every possible optimization you could apply. I'll go through them individually. This is llama.cpp specific, but you'll find the same flags with slightly different syntax apply to other runtimes.

\-flash-attn (-fa) on: forces flash attention optimizations and paths. Explicitly set to on to override any fallback. Auto can be optimal if the model is recent and is not yet optimized.

\-batch/-ubatch: batch is the decode chunks, ubatch is the prefill chunks. They must be multiples of one another, otherwise you're adding compute. Equal to one another is ideal for CPU offload, and a 2x-4x higher batch is optimal for full GPU loads. You'll need to play with these values to optimize. Batch/ubatch should be a power of 2 to optimize architecture. Intervals of 256 is typically good enough for testing.

\--no-warmup: prevents initial model poll to load weights. Removes unnecessary latency

\-- cache-reuse x: instructs the model to reuse cache values and scan for similarity at x token intervals

\-jinja: highly underutilized and important flag. Utilizes native chat template kwargs to ensure output consistency.

\--presence-penalty: flat penalty rate to words that appear in text. Used mainly for creative writing to prevent repetitive prose.

\--repeat-penalty: reduces liklihood of already used tokens being reused. Best used for preventing loops in thinking agents.

\-temp: model temperature– how creative the model is. 0 is completely deterministic, 1 is creative freedom.

Top-k: hard cutoff that keeps only the k most likely words. Each model will have recommended k values for thinking/instruct setups. Low key reduces hallucinations at the cost of repetition and loss of creativity

Top-p: includes P percentage of possible tokens. It reduces liklihood of hallucination dynamically.

Min-p: dynamic cutoff based on highest probability token. If the biggest probability token is 50% and min-p is 0.05, then the bottom 0.025 (2.5%) liklihood tokens will be excluded. Reduces noise without hampering creativity terribly.

Chat template kwargs: explicit chat template activations; newer runtime compilations should have flags for these. Controls model reasoning, reasoning effort, and internal chain of thought storage.

\-threads (-t): number of cores used for decode. Set equal to physical cores (hyperthreading will thrash cores and degrade results)

\-threads-batch(-tb): number of cores used for prefill. Set to double the number of physical cores, as hyperthreading helps here.

\--gpu-layers (-ngl) : total layers on GPU. Fit as many as you can without OOM.

\--n-cpu-moe: number of MoE layers offloaded to cpu. For MoE models, you should always offload these first and keep all layers on GPU if possible. Offload as few as possible to CPU.

\--override-tensor...: tells the runtime to offload the n-gram table if needed. For qwen 3.8 flash specifically.

\-ctk/-ctv: k and v cache quantization. K cache should \*never\* be below q8 unless youre running on less than 80k context. V cache can be q4 up to 150k context without issues, for the most part. Generally, q8 for both will be best for speed and is my recommendation to start with.

\--cache-ram: sets RAM aside for cache allocation to ensure it doesn't go to swap

\--context-checkpoints: the amount of checkpoints captured. Generally, you don't need as high as the defaults do. Leave the default if you have extra RAM, otherwise you may want to lower it.

\--kv-unified: tells all instances to run on the same kv cache pool rather than allocating individual cache.

\--reasoning-preserve: tells the runtime to retain CoT traces for evaluation. Prevents the model from getting stuck or looping as much when thinking.

\--load-mode: tells the runtime how to load in the model. No mmap generally loads slower but is more stable. Mmap+mlock (or just mlock) is a balance of both with fast loading and page faults initially but it stabilizes as you run it, mmap alone is fast but will cause constant page faults and slows down inference, especially on models with CPU offloading.

I hope this guide helps! I'd be willing to answer any specific questions or make any additions if there are additional areas the community agrees are major uncovered inference bottlenecks. Some claims are based on my personal experience and I am open to data based claim revisions or anecdotal counterclaims, so feel free to provide. Happy tuning!

💬 53 (+10) open on reddit ↗
▲
70
+20
46👁
r/LocalLLaMA · u/lbgos_Loss783 · 7d ago
I made my own cybersecurity benchmark and ran Qwen3.8 27B, here's how a local model actually does at hacking

Hey local AI community, I've been working on this for a while and finally feel ok sharing it.

It's a cyber benchmark where the model gets a shell in an isolated docker box and has to find the exact flag. Pwn, web, crypto, rev, forensics, a few real CVEs and some multi-stage ranges. 19 tasks, 6 models, 544 scored attempts.

To be clear, I didn't build every task by hand. GLM 5.3 helped me create several of them. For an open model its cyber capability is really high, and it barely refuses anything, so it was one of the best options for this. GLM 5.3 isn't one of the benchmarked models.

The local part: I ran Qwen3.8 27B (Unsloth Q4\_K\_XL, xhigh) on a llama.cpp RPC pool across a 3090 and a 3080 in two Proxmox nodes, connected over a direct 2.5G link. That gave me enough concurrent tps to run several agents at once. I also started a low reasoning run, but it was taking 20+ hours because of a harness problem, so I killed it.

Why I'm posting now: John Hammond put out a video about how threat actors use AI (https://www.youtube.com/watch?v=xHDc6-7bjyw). One part is a guide from a criminal forum on running abliterated models on RunPod, and one of the models in it is Qwen3.8 27B. I had benchmark data on that exact model, so here's what it can actually do.

Stock Qwen3.8 27B got 28.1% on the first try and 0% on pwn. Not bad for a 27B on two gaming cards, but not much of a threat on its own either.

The cheap API models are a different story:

\- MiMo 2.6 Flash solved 73.7% on the first try

\- GPT-6 Luna solved 90.9% within 3 tries

\- On multi-stage ranges, where you chain several steps, the top models got 92-96%

Pwn is still hard for everyone (best was 56%), and 3 tasks haven't been solved by any model in 82 attempts.

Results: https://lbgos.dev/bench

Harness (MIT): https://github.com/lbgos/rangebench-harness

The tasks aren't public so they don't leak into training data, but you can still run them. DM me here or on X (lbgosna), and I'll send them over. You run it on your hardware or tokens and I'll add your results to the board. If a few people send local runs, I'll make a separate local-only table.

This is my first time building something like this, so any feedback on methodology, task mix or what's missing is welcome.

💬 40 (+9) open on reddit ↗
▲
10
+3
18👁
r/LocalLLaMA · u/Designer_Elephant227 · 7d ago
Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

Hi, i got QFN running on my single r9700 but im not sure if i did everything right to get the best quality and speed out of this setup. Dont want to annoy anybody, maybe someone with the same card can tell me if this looks normal.

What i run:

\- model: Qwen3.8-Flash-Next from turboderp, exl3 5.05 bpw (head 6 bit, vision 6 bit, mtp 5 bit)

\- backend: exllamav3 rocm fork from phoenixhaxor (commit cbbef08), i had to patch one file for gfx12

\- cpu/ram: Ryzen 9 7945HX3D with about 90gb ram

\- 112 of 512 experts per layer are on the gpu, the other 400 on cpu with 16 threads

\- 262144 context, q8 kv cache, chunk size 4096, batch 1

\- mtp drafting is on, acceptance is around 53-54%

\- the ngram table (102gb, bf16 not quantized) gets streamed from nvme

Speed at 230k context (prose): prefill 863 t/s (only the new 101k tokens, the rest came from the prefix cache) and decode 34.7 t/s.

Is this ok for the r9700 or can i still tune something? Thanks 🙂

💬 26 (+5) open on reddit ↗
▲
79
-1
40👁
▲
12
 
32👁
r/LocalLLaMA · u/AdRepulsive7837 · 8d ago
best <40B alternatives to Qwen/Deepseek for (1) Coding (2) Long document QA test

Due to some reasons, Qwen/Deepseek Chinese models are NOT allowed in the my workplace. So, what local models, do you think, is the best alternatives to Qwen/Deepseek for

(1) Coding

(2) Long document QA test (like giving a long medical history of 120k tokens and ask a question based on that medical history)

Gemma 31B ?

Muse Glimmer 30B ?

Nemotron ?

also, I know that nothing beat qwen nowadays, but are there fine tunes from these alternative model that make them better than qwen3.8 27B in terms of coding ?

💬 37 (+5) open on reddit ↗
▲
10
+2
11👁
r/LocalLLaMA · u/ParvusNumero · 8d ago
Mirostat?

Reading another post made me think:
Is anybody still using Mirostat?

It was all the rage and people said it avoided the “boredom trap” for long texts.

Do newer generation models not need that anymore, or are other samplers superior?

💬 24 (+3) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/Emotional-Sky9692 · 8d ago
I built a cross-client AI memory hub — 23 AI coding agents sharing one SQLite file (local-first, no cloud)

I run 8+ AI coding agents daily (Claude Code, Cursor, Windsurf, Codex, etc.) and they all have amnesia between sessions — worse, they don't share memory with each other.

Existing solutions (mem0, Zep, Letta) are cloud/server-based. I wanted something local and dead simple: just make all agents point to the same SQLite file.

So I built MemTether — a memory hub that works via file-level pointers (junction/symlink). No cloud, no API fees, no abstraction layer.

Key features:

\- 23 client adapters (auto-detect and connect)

\- Source attribution (knows which agent wrote each memory)

\- Bi-temporal (what was true vs what the system knew)

\- Q-Value ranking (memories that get used rank higher)

\- FTS5 + vector search (bge-m3, local embedding)

\- MCP server included

Stack: Python, SQLite FTS5, ChromaDB, FastAPI. All local.

GitHub: https://github.com/MemTether/MemTether

PyPI: pip install memtether

Blog with design decisions: https://dev.to/lanbass869cell/i-built-a-cross-client-memory-hub-for-ai-agents-heres-what-i-learned-418l

Would love feedback from people who juggle multiple AI coding tools.

▲
5
-1
13👁
r/LocalLLaMA · u/Competitive-Scar-627 · 8d ago
Model weight inferencing

I have 4050 6gb gpu, 24 gb ram which model should i choose to run i need speed. i try qwen 3.8 27b and feel too slow tried from onslot studio.
I have heard of weight inferencing does it helpful what should i do to try weight inferencing.

💬 23 (+2) open on reddit ↗
▲
0
 
19👁
r/LocalLLaMA · u/Voxandr · 8d ago
Latest Gemini 4 is distilled from GLM 5.3 (or did they just finetuned it? :D)

https://preview.redd.it/ei6hgtxirzsh1.png?width=1111&format=png&auto=…

I am running GLM 5.3 flash .
After nearly a month of usaged , i got chinese response for first time so i am checking if there special setting to turn off Chineese . When i queried about that tru Quick Google AI mode which now uses Gemini 4 - it is replying as it is GLM5.3 .

▲
0
 
3👁
r/LocalLLaMA · u/serige · 8d ago
best open models from recent releases for math research?

I know models from OpenAI are probably the best for math research, but given the recent accusations against OpenAI that research work could be used to train their own models, the lack of transparency makes me consider moving to local models. Does anyone have good experience with the recent open model releases (especially flash models that I can run on my 2x spark cluster) when it comes to doing math research? Or techniques that work well with these open sources models in this particular setting? Thanks in advance for helpful advice.

▲
0
 
17👁
r/LocalLLaMA · u/Altruistic_Heat_9531 · 8d ago
Guys... OpenAI API on VLLM and Llamacpp already supported grammar enforcer... (AKA JEV)

https://preview.redd.it/tcdyumkf3zsh1.png?width=750&format=png&auto=w…

If you want to try JEV like generation, or what we could call an already fucking exist, zero-shot, training-free classifier, your LLM already supports it.

The model running on the very computer you host does not need any server-side modification. The feature is called structured output, and the underlying idea is grammar-constrained generation or a grammar enforcer.

Under the hood, vLLM supports multiple structured output backends, such as XGrammar and lm-format-enforcer, while llama.cpp uses GBNF.

Basically, the decoder constrains the LLM so it can only generate tokens that are valid under the specified grammar or schema.

For vLLM:

https://docs.vllm.ai/en/v0.8.2/features/structured\_outputs.html

For llama.cpp, structured output is wrapped in a different JSON request-body format, or you can use GBNF directly.

vLLM:
structured_outputs
└── json
└── {schema}

llama.cpp:
json_schema
└── {schema}

This is example of that schema in vllm of product sentiment analysis, which roughly mapped to most of jev use cases:

SCHEMA = {
"type": "object",
"properties": {
"sentiment": {
"type": "string",
"enum": ["negative", "neutral", "positive"],
},
"score": {
"type": "number",
"minimum": -1.0,
"maximum": 1.0,
},
"value": {
"type": "string",
},
},
"required": ["sentiment", "score", "value"],
"additionalProperties": False,
}

def classify(news: str) -> dict:
payload = {
"model": MODEL,

"messages": [
{
"role": "system",
"content": SYSTEM_PROMPT,
},
{
"role": "user",
"content": news,
},
],

"temperature": 0,

"max_tokens": 512,
"chat_template_kwargs": {
"enable_thinking": False,
},

"response_format": {
"type": "json_schema",
"json_schema": {
"name": "news_sentiment",
"strict": True,
"schema": SCHEMA,
},
},
}

resp = requests.post(
ENDPOINT,
json=payload,
timeout=30,
)

Look, I think JEV and what it brings to the community as a refresher on already great classifier-style workflows is a plus for me. I just want to ground the discussion in the fact that this already exists, and you do not need a custom model just to study or experiment with training-free classification.

I am very familiar with this because it is part of my profession in lakehouse platforms. Basically, we use 1B to 4B models to ingest unstructured data such as images or documents, then extract structured information such as place, time, sentiment, entities, and so on.

Why? Because working with well-formatted SQL data is much less of a pain in the ass than repeatedly querying raw unstructured content through an Elastic/OpenSearch index.

💬 13 (+1) open on reddit ↗
▲
0
 
20👁
r/LocalLLaMA · u/Lordofwhut · 8d ago
RTX 5090 & RTX 5070 Ti not well thought out

Hi All,

TL;DR: I got excited building a PC and kept upgrading / swapping and building and ended up with a work station that is more than I can use. It was fun and frustrating, but I probably won't do it again. If you have advice or a suggestion on how you would use a 5090 and 5070ti in the same PC I would like to hear it!

So, this all started when the 5090 was announced. I signed up to be in the lottery to buy it at msrp from NVIDIA. I had an Alienware R15 with an i7 and 4080 with a 1300 w psu. The 4080 was fine but I had really wanted a 4090 for better fps in gaming and I was starting to explore local imagine generation. I got selected, bought the 5090 and went to swap it into my PC when I realized that I was not able to use the power connector that was in my Alienware PC.

I then decided I would sell the Alienware and build my first PC. At the time I was still building a gaming focused PC with a Ryzen 9 9950x3d the 5090 and 32 gb of ram (I tried to save money on the ram thinking I could upgrade later boy did I get that wrong). Then I had less time for gaming as I started to learn about Ollama, and then Llama.cpp.

I was constantly downloading and trying new models. At one point I had nearly 1 TB of models that would fit on my 5090 (gemma 4 12b, 26b-a4b, 31b; gpt oss 20b; nemotron 3 nano 30b a3b; so many Qwen models etc). Then it seemed like the better models kept getting larger, so I looked into getting a second GPU (ram prices were/are nuts and vram seemed like the better "investment"). I realized that I would not be able to run another gpu at its full PCIe lanes with my gaming PC as the Ryzen 9 couldn't support it. So, I started looking for used Threadripper hardware.

I found a 7960x with 96 GB of ECC DDR5 ram, a 5070 ti and 20 tb of storage for less than I built my gaming PC. I wasn't able to find much in regard to PC builds with a 5090 and 5070ti. Most builds were dual 3090s or other matching cards. Still after looking into it, I figured the extra vram and the fact that they were both blackwell GPUs would work out well.

I thought I would be able to just drop my 5090 into the threadripper workstation and I would be good to go. Unfortunately, the 5070 ti that came with it was a four slot card and the spacing just would not work with the motherboard (Gigabyte Areo D) layout and the cases that I had. So I put the 5070ti into my gaming PC, sold it, and bought a 2 slot 5070ti and put it into my workstation.

What does this have to do with LocalLLaMA? Well, while I was doing all of this the LLM space kept moving forward. I now have Hermes Agent set up running Llama.cpp and Qwen 27b Q4 on my 5090. I swapped to a Q8 to run across both my 5090 and 5070ti but the speed trade off was not worth the accuracy increase. So, I went back to running the Qwen 3.8 27b Q4 and my 5070ti is completely idle. Going from 32gb to 48gb did not have the impact I thought it would, at least not with my pairing. The 5090 is pretty quick when everything is loaded onto that card, and Qwen 3.8 has been pretty great on it too, that I have not found a good use case for deploying the 5070ti.

Hermes / Qwen suggested I run another Llama session with a smaller model on the 5070ti but I don't currently have a need to run something else. What would you do or suggest I explore?

Additional background context: I do not work in tech or software at all. I am an asset manager for a independent power producer, but I can not use my personal PC for work due to IT policy (I would have my agent working around the clock to review contracts, analyze system performance, track deliverables / open items etc). I have taught myself everything about PCs and local LLMs from creeping this and other subreddits / youtube videos. I literally have no one in my social circles that I can converse with about tech whether its PC building or hosting LLMs.

💬 35 (+2) open on reddit ↗
▲
79
+8
30👁
r/LocalLLaMA · u/Ok-Breadfruit-3523 · 8d ago
Update on the “Monstrosity”. 6 BC-250 board cluster post image

This is 6 bc-250 ex mining boards with 5 in the asrock 4u12g case they came in. After a lot of testing my current preferred setup is 4 boards running Qwen Next Flash IQ2\_XS at 100k context with around 28 tok/s for short generation and 24 tok/s at 50k with around 115 ppt. The other two boards run 3.6 35b q4 at 60 tok/s with 100k context and 450 ppt. This is all using llama with vulkan and rpc over 1gb Ethernet.If anyone has any suggestions with this beast I am all ears. I had these boards left after reselling a bunch and had never done anything with local ai before so this has been a blast. Also yes that is a cardboard box with 3 fans on top as the intake.

💬 42 (+5) open on reddit ↗
▲
126
+11
83👁
r/LocalLLaMA · u/FutureStriking283 · 8d ago
Anyone wonder why americans lag so far behind in the open LLM market?

I mean , DeepSeek, Kimi, GLM, MiniMax -- the list of Chinese LLM's is such a long freaking list. As American's -- why don't we feel .. a little funny .. about being so far behind? Chinese entrepreneurs are peneuring like crazy and American's .. just obsess on .. what ?

update -- already I'm starting to see some clear answers. American's are putting money ahead of technology. Completely understandable.

second update -- I asked "why can't we have american AI as good as or better than the chinese" and BY FAR the number one most supported comment? "Have you even thought of shareholder value!?" . We .. America , are so fucked.

💬 425 (+28) open on reddit ↗