109 posts · 1 sub · RSS
← prev Sep 30, 2026 → Oct 1, 2026 next →
2026-09-30 → 2026-10-01 hourdayweekmonthyearall
allr/LocalLLaMA
▲
731
+57
70👁
r/LocalLLaMA · u/xenovatech · 10d ago
We just open-sourced the world's fastest WebGPU kernels for local AI on Hugging Face post image

The collection includes kernels for more than 200 common ML operations, all of which can run entirely locally in your browser on WebGPU. We're also working to upstream these optimizations to Transformers.js, ONNX Runtime Web, LiteRT.js, and more!

Kernels: https://huggingface.co/kernels?platform=webgpu
Blog: https://huggingface.co/blog/webgpu-kernels

💬 45 (+2) open on reddit ↗
▲
458
+24
68👁
r/LocalLLaMA · u/Rombodawg · 9d ago
Least to most expensive (Somewhat modern) GPU's with 32gb of vram (Under $1600) Based on ebay listings post image

I was researching prices on ebay and fed claude a bunch of images of listings. I had it make a chart and thought it would be useful to share.

💬 260 (+24) open on reddit ↗
▲
410
+28
45👁
▲
353
+57
63👁
r/LocalLLaMA · u/-p-e-w- · 9d ago
Heretic is on PewDiePie!

So I haven’t played a computer game in 20 years, and I know nothing about Minecraft, and I definitely prefer classical literature over YouTube culture, but even I have heard about the individual called PewDiePie, for two reasons:

  1. His monicker starts with the initials of my own name
  1. I remember a recurring Internet meme a few years ago where he was competing for the most subscribers with an Indian film music channel

I had never watched a single one of his videos, however.

Well, until today, when people started spamming me with messages informing me that Mr. Kjellberg aka PewDiePie has tried out Heretic and made a video where he talks about it:

https://m.youtube.com/watch?v=ODDJXGY_1kQ

(Heretic mentioned around 9:00)

Obviously I’m thrilled that a less technical audience is being exposed to my work, and the more people understand what is possible the better. I expect to be receiving a couple hundred more mails in the coming days asking how to run Heretic on ChatGPT (you can’t), or accusing me of working for the CIA (I don’t), but other than that, the more the merrier I guess 😏

Heretic 2.0 coming soon…

💬 91 (+4) open on reddit ↗
▲
270
+20
26👁
r/LocalLLaMA · u/Spiritual_Impress_30 · 10d ago
Thank You, Mradermacher.

best iq quants in the biz, got me gemma 4 26b to run 75tok/s tg and 1500 pp on 2x 4060 8gb using lmstudio serving to hermes, much work has been done.

▲
262
+24
53👁
r/LocalLLaMA · u/jacobpederson · 9d ago
Why am I like this? (Full Chat and Image generation on a 286 Tandy 1000 TL/3) post image

40 year tech gap? No problem! The Tandy runs DeskMind, a native DOS program. It talks over WiFi (a PicoMEM 2 card with mTCP) to a small Python server on my PC. That server drives Qwen3.8-27B (NInfer on a 5090) and Krea 2 (ComfyUI on a 4090). The 286 never sees JSON, base64 or a PNG. It gets plain text lines and pictures that are ready to copy into video memory.

Drawing from chat without tool calling. The system prompt tells Qwen to wrap a picture request in \<draw>...</draw>\. The server catches the tag mid-stream, runs Krea 2, dithers the result, and streams a \picture ready\ line. "Draw me a 286 AI logo" takes about 9 s from Enter to a thumbnail in the chat.

Qwen Vision sees what the Tandy sees. When you ask about a picture, Qwen gets the original and the 16-colour dithered version, so "why does the sky look striped?" has context. The latest picture stays in context for follow-ups.

The model knows where it lives. The system prompt knows it's talking IN a 286 with 80 columns and 16 colors. It keeps answers short and plain ASCII, and when asked about games it suggests Wolfenstein 3D or Commander Keen.

Streaming cleanup for a 1990 screen. Reasoning is stripped, Markdown is removed on the fly, Unicode becomes code page 437 (bullets turn into the CP437 block character), and tiny tokens are merged into \~48-character lines so the 286 isn't redrawing for every token.

Per-request reasoning effort: low for chat (replies start in \~2 s), medium for rewriting image prompts.

Prompt "enhancement" tuned for dithering: bold shapes, strong contrast, simple backgrounds. The rewrite shows up in an edit box on the Tandy before drawing, and the rules themselves can be edited from the Tandy.

\- \*\*A Dither Lab\*\* in the server GUI: Floyd-Steinberg, Atkinson, Bayer, Yliluoma and more, previewed at the Tandy's real aspect ratio.

Numbers: Krea 2 at 1024x768 in \~10 s 8 steps, the target is 640x200). About 1 s to send a 64,000-byte picture over the PicoMEM WiFi (56-79 KB/s).

Code (GPLv3): https://github.com/RowanUnderwood/DeskMind

added image gallery https://imgur.com/a/wysqoM1

💬 129 (+12) open on reddit ↗
▲
244
+11
31👁
r/LocalLLaMA · u/LambdaHominem · 9d ago
AI CEO Interviews (2026) post image
💬 34 (+1) open on reddit ↗
▲
230
+26
46👁
r/LocalLLaMA · u/Significant-Price695 · 10d ago
Oído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)

I'm part of the Lokutor team that built this.

Model: NVIDIA Conformer-CTC Small (13M params, int8). It runs on an ESP32-S3 with 8 MB PSRAM, no GPU or NPU. LibriSpeech WER is 3.7 / 8.2, versus 6.3 / 15.9 for Whisper tiny.en on a laptop. Under real noise (DEMAND: car, kitchen, cafeteria) plus babble and reverb, mean WER is 8.4 vs 12.1 for Whisper tiny.en. You can try the exact chip arithmetic on your laptop mic with live_demo.py. https://github.com/lokutor-ai/oido

💬 40 (+1) open on reddit ↗
▲
218
+8
47👁
r/LocalLLaMA · u/jacek2023 · 10d ago
add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp

now you can use GLM-5.3-Flash on your home computer

💬 89 (+10) open on reddit ↗
▲
207
+3
26👁
r/LocalLLaMA · u/jacek2023 · 9d ago
Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp

now you can use MTP with Qwen Flash Next, time to switch from Qwen 3.8 27B?

(merged after 17h of development)

quants: https://huggingface.co/ggml-org/Qwen3.8-Flash-Next-GGUF

link to the previous discussion (I deleted the old post to avoid duplicates): https://www.reddit.com/r/LocalLLaMA/comments/1wur4lt/qwen\_flash\_next\_mtp\_work\_restarted/

▲
184
+7
31👁
r/LocalLLaMA · u/Educational_Sun_8813 · 10d ago
Preorder for new AMD Ryzen™ AI Max 400 Series 192GB from framework just started

Framework Desktop
Framework Desktop DIY Edition (AMD Ryzen™ AI Max 400 Series) 192GB

💬 129 (-1) open on reddit ↗
▲
181
+8
46👁
r/LocalLLaMA · u/Creative-Type9411 · 9d ago
Finally got my 4th card in (64gb total) post image

I was waiting on the blowers for the T4s and posted this before it was finished, other than some braided cable sleeves for the fan wires its pretty much good, I was going to upgrade the CPU, but I'm getting great speeds comparatively to a CPU in my old box that had way more cores, so I don't think it's going to make a difference.

Fractal Design Torrent Mid-Tower Case w/Tinted Glass
SuperMicro X11SPA-T Motherboard
Xeon W3225
768GB DDR4 ECC 2666
4xTesla T4 16GB GPU
4x1tb Samsung 870 EVO SATA SSD Raid

Ubuntu 26.04/llama.cpp/openwebui+custom powershell harness

now i want more cards 👀

💬 108 (+5) open on reddit ↗
▲
179
+16
54👁
r/LocalLLaMA · u/ea_man · 9d ago
Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now post image

pi-llama-skip-reasoning is an extension for the Pi.dev harness that forces a local llama.cpp model to stop reasoning and answer / act immediately.

When you are deep into the ctx session and ask 27B a simple question about a fact or need a direct action, the model may still feel the urge to indulge in copious deliberation in the reasoning trace. This extension allows the user to force the model to snap out of the reasoning stage and provide the answer immediately.

Disclaimer: don't skip the reasoning for important problem-solving, that would hurt quality.

This uses the same mechanism the llama.cpp web interface uses to skip reasoning, so it's native to llama.cpp, this extension is meant for Pi.dev yet the same mechanism could work for other harnesses.

Usage: /skip-reasoning command or shortcut Alt+T ,
Install: pi install npm:pi-llama-skip-reasoning

\- https://pi.dev/packages/pi-llama-skip-reasoning

💬 82 (+4) open on reddit ↗
▲
163
+6
31👁
r/LocalLLaMA · u/WebAssemblyMan · 10d ago
DeepSeek now trained on Ascend 950 post image

26 months ago Liang Wenfeng said:
"Someone must step onto the frontier."

Now they are training their models on Ascend 950

▲
161
+3
50👁
r/LocalLLaMA · u/MLDataScientist · 10d ago
Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine

I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling and the developer fixed them.

Inference engine (only runs on Nvidia for now; AMD support is experimental): https://github.com/Niko1221/Strata

Here are some metrics with screenshots. My laptop has 5070ti 12GB VRAM, 64GB ddr5 RAM, Intel 275HX CPU, gen4 SSD.

Aquarium test \(unsloth studio connected via local API\)

The model I used was https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/tree/main/IQ3\_XXS which has a good quality for its size. Above, the model generated the aquarium test. At 43k context depth, it was running at 51 t/s. Stock llama.cpp reached only 23t/s with the same quant.

32k context read at 1500t\/s \(unsloth studio via local API\)

This quant could only reach 100t/s PP with stock llama.cpp using the same quant. Strata was reading 32k context text at 1500t/s. This is way above my expectation. This quant can load with up to 200k context at 8bit. However, I was only using 131k context.

Memory utilization

As you can see it is utilizing 11GB VRAM and 56GB RAM (includes system/OS programs).

This engine is specifically built for one model only and only select ggufs (ISTA-DASLab) work with it. You can use IQ3\_S from ISTA-DASLab which they claim recovers full model's performance on coding benchmarks. I tested IQ3\_XXS for some time and I would say it is an excellent model.

I never thought 12GB VRAM would be enough to run frontier models from 6 months ago locally on a laptop. What a time to be alive!

💬 143 (+3) open on reddit ↗
▲
157
+1
32👁
r/LocalLLaMA · u/soyalemujica · 9d ago
If one hour of AI is costing me 0.12€ is paying for frontier a cheaper option?

Running Qwen flash next of even Qwen 27b dense, I can do any,burning sticking to flash due to its speed, and the kwh cost is at 0.25€ where I live in, ranging from 0.11€ to 0.35€, so I used chatgpt to help me calculate the total kwh consumption on my 7900xtx plus 9800x3D, and well that is the result.

Judging by this, if deepseek flash is indeed then faster to use per 1m token, does it mean that frontier is cheaper for me or am I calculating something wrong ?

💬 175 (-1) open on reddit ↗
▲
122
+12
71👁
r/LocalLLaMA · u/aya-ifm · 9d ago
AMA about K2 Horizon, Meet our team from IFM

Hi r/LocalLLaMA

We’re researchers at the Institute of Foundation Models (IFM), an AI research lab dedicated to open and independent development of frontier-class foundation models.

We recently released K2 Horizon a connected fleet of six fully open models with size ranging from 0.9B to 375B. In addition to weights, we also open-sourced training data and recipes, training code, intermediate checkpoints, fine-grained training logs and evals.

Ask us anything about pre-training and data mixes, post-training, small models on-device, MoVA and sparse attention, deployment, what’s out now, and what’s coming next.

Participating in the AMA:

  • Hector Liu u/hunterhector
  • Alexander Moreno u/IFMAlex
  • Mikhail Yurochkin u/my-moonfolk
  • Rupesh Srivastava u/k2pt
  • Junlin Chen u/Junlin_Chen110
  • Haonan Li u/East-Career9147

We'll be live Mon, Oct 5, 8–10 PM PT. Questions are open now, so drop yours anytime!

Join IFM on Discord: https://ifm.ai/discord

https://preview.redd.it/awu0r6fbowsh1.png?width=3240&format=png&auto=…

💬 135 (+108) open on reddit ↗
▲
119
+1
33👁
r/LocalLLaMA · u/Usual_Maximum7673 · 9d ago
Jeff-Qwen3.5-0.8B v1.2 + 9 LoRA adapters: put it in front of Qwen3.8-27B for 38× faster decisions and +8.7 points accuracy, for under 2 GB extra memory

A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO 6000, but as a general zero-shot classifier it trailed the big models.

Then it occurred to me that most decisions an agent makes in front of a local model aren't open-ended. They fall into a handful of recurring kinds: is this a prompt injection, which tool to call, how urgent is this ticket, is this answer grounded in the sources. So I trained 9 LoRA adapters, one per job, and you pick the ones you need. The server loads the base once plus whichever adapters you choose (about 40 MB each), and every request either names an adapter or goes to plain Jeff.

That means you keep both: the base model stays untouched, so you still get Jeff's general zero-shot ability for anything new, and the adapters give you near-perfect accuracy in the domains you care about. Each adapter was also trained with 10% of the base model's own training data mixed in, to help it keep its general skills.

Everything is on jeffhub.ai: the adapters, the results, the docs. Code on GitHub, models on Hugging Face, and you can try all nine adapters in your browser.

The headline: I let Jeff + adapters answer first and pass only the queries it's unsure about to Qwen3.8-27B. Same test rows both ways, on an M4 Max:

|Measure|Qwen3.8-27B alone|Jeff + adapters, 27B only when unsure|
|:-|:-|:-|
|Accuracy (mean of 8 adapters\*)|86.6%|95.3%|
|Time per decision (mean)|8.1 s|0.25 s (38× faster)|
|Wrong answers|13.4%|4.7%|
|Memory|28.6 GB|under 2 GB for Jeff, even with all 9 adapters loaded (+6.9%)|

On the five decisions an inbox agent makes for every message (guard, triage, support intent, tool choice, grounding) alone: 87.7% → 95.7%, 39× faster. Jeff wins outright on 8 of the nine adapters and ties on grounding (96.3% vs 96.7%, at 20× the speed). On their full held-out test sets, six of the nine adapters score 97–98%. On a GPU, a decision takes about 30 ms, whether you load one adapter or all nine.

\*Emotion is left out of the averages: picking the single strongest of 27 emotions (or neutral) in short Reddit comments is hard even for people, and the human labels often disagree. Jeff + adapter scores 60.6% there against the 27B's 35.6%, at 42× the speed. Including it, the average across all nine adapters is 91.4% for Jeff + adapters against 80.9% for the 27B, so leaving it out makes the gain shown above smaller, not larger.

Caveats, up front:

  • the 27B ran in 8-bit with step-by-step reasoning off (with reasoning on, the speedup would be even more dramatic);
  • each task used a fixed sample of 300 held-out rows (500 for emotion and legal-clauses);
  • each adapter's "pass it on" threshold was chosen on separate calibration rows, before the test rows were scored.

Data: 4 adapters are trained on public data sets. 5 are mostly synthetic. Every generated row records which model wrote it, and the cards give the counts. Every data set went through a shortcut check and an independent review before training, and a lot of first drafts failed: things like the answer being given away by length.

What's open: weights (Apache 2.0), code (MIT), and each adapter's test and calibration sets, so you can check every number. The training data isn't published.

This is a community preview: I'd love feedback.

Next: over the next \~36 hours I'll train v1.3, a long-term-support base. The fixed parts of a prompt come first, so servers can prepare them once and reuse them, which means faster decisions. I'll then retrain all nine adapters on it and keep the request format stable, so others can build and submit their own adapters. The adapter kit, with the data checks I used, is in the repo.

I've got access to more hardware now, so if there's a decision you'd like an adapter for, tell me and I'll train it.

The goal: when the next generation of local models lands (like everyone, I'm watching for Qwen 4), anyone running one locally should also have a tiny, fast, well-calibrated decision layer in front of it.

💬 33 (+1) open on reddit ↗
▲
114
+4
44👁
r/LocalLLaMA · u/rmonsurate · 9d ago
Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)

We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next.

Victoria (coding and agents)

  • Qwen3.8-Flash-Next cut down by 44% using a paper / technique called REAP: 512 down to 288 per layer.
  • Retrained at 4-bit (NVFP4) afterwards, so it's trained for the format it ships in rather than just quantized after the fact.
  • Terminal-Bench 2.1: 70.0%, averaged over 3 runs with an 8h per-task timeout. Our previous NVFP4 build scored 62.5%.
  • HumanEval: 159/164.
  • 48.0 GiB of weights, including the draft head. The 95.4 GiB n-gram table is separate and not counted in that number.
  • 280 tok/s single stream on one B300 with the draft head, versus 135 without it.
  • GGUF Q4_K_M is 49.17 GiB. It scored 75.3% on Terminal-Bench (a single run, so treat it as noisy) and 93.2% on HumanEval (averaged over 5 runs).
  • Uses 35% fewer output tokens than our previous build.

Maple (Canadian questions)

Most models answer questions about taxes, benefits and regulations as if you live in the US. Maple is fine-tuned to default to Canada. On 600 held-out questions, with search:

  • Cites an official Canadian source: 6.0% before fine-tuning, 62.9% after.
  • Fully correct answers: 6.6% before, 21.8% after.
  • "No answer" responses: 47.2% before, 23.7% after.
  • It pushes Canada onto people who said they live somewhere else less often: 2.9% before, 1.0% after.

Coding holds up: 157/164 on HumanEval. Grading was done by an AI judge panel; human review hasn't happened yet.

Links:
https://huggingface.co/rmonsurate/Victoria
https://huggingface.co/rmonsurate/Maple

Happy to answer questions about running them.

Edit: llama.cpp users. The GGUF carries our draft head, and mainline llama.cpp doesn't know about it yet, so it fails with "expected 1256, got 1224". Your download is fine. For now, build from our fork: github.com/rmonsurate/llama.cpp, branch qwen4exp-mtp. Prebuilt binaries are on the way. Thanks to the reader who caught this.

💬 42 (+3) open on reddit ↗
▲
84
+5
41👁
r/LocalLLaMA · u/WebAssemblyMan · 9d ago
DeepSeek harness 0.2 - Optional Bundle Architecture, Windows Sandbox improvements, Async Question Mode, Desktop release, Web Search without key post image

Optional Bundle architecture: Schedule (session-local delayed / timed / interval reminders) was removed from the default set and made an explicit Optional Bundle. This cleanly separates “installed” from “enabled” and is the first systematic use of the Profile + Bundle model for official features.

• Windows Sandbox improvements: A new permission-diagnosis skill can detect common Access Denied causes and perform backed-up, recoverable permission fixes after user authorization, giving the Agent a reliable recovery path instead of blind retries.

• Async Question Mode (experimental): “Ask the user” is no longer a hard synchronous block. After a timeout the Agent can keep working while the user answers later, introducing asynchrony between interaction and execution.

• Model-layer polish: DeepSeek-account sessions can use Web Search without an extra API key; third-party model catalog updated (some old IDs removed); long model lists now support fuzzy search and keyboard navigation.

• Desktop release: Official Windows and macOS clients are out (Linux unsupported). Account login is supported, suggesting paid plans may be coming soon.

• Overall theme: Version 0.2 strengthens the Agent Runtime’s composability, recoverability, permission boundaries, and execution-state semantics — the practical foundations needed to move from a toy toward production use.

💬 17 (+3) open on reddit ↗
▲
83
+8
35👁
r/LocalLLaMA · u/Skyline34rGt · 10d ago
BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B

"AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.

AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.

  • Architecture: Dense Qwen3.8-compatible multimodal model
  • Parameters: 27B
  • Context length: 262,144 tokens

[](https://huggingface.co/BAAI/AREX-2#key-features)Key features

  • Long-horizon self-improvement: turns extra test-time rounds into useful solution refinement.
  • Feedback-driven reflection: reads scores, logs, errors, and timings to decide what to change next.
  • Cross-domain performance: training on coding and machine-learning tasks also improves the model's deep-research performance.
  • Long-horizon reasoning: sustains productive iteration as the task budget grows."

Gguf's - https://huggingface.co/mradermacher/AREX-2-GGUF

💬 29 (+2) open on reddit ↗
▲
58
+5
22👁
r/LocalLLaMA · u/jjusko20 · 9d ago
Update #2: Post training yandex/AliceAI-80B-A3B [instruct!] from scratch post image

Last update: https://www.reddit.com/r/LocalLLaMA/comments/1wu9ksu/update\_yandexaliceai\_80ba3b\_fine\_tune\_progress/ \- basically, an instruct fine tune on the base model using a synthetic distilled data set. I've been posting regular updates so I imagine at least a few people have seen this.

Live stream: https://figure-bios-expect-cio.trycloudflare.com/

UPDATE: Finished train. hopefully some examples soon.

The initial train is finally almost done, after about 48 hours of humming. While the loss curve looks a little crazy, I've done some analysis (and some chatting with the LLMs) to understand that my average loss each epoch has been steadily decreasing (few reasons the loss curve looks wacky, vocabulary size, low to high token counts in epochs, etc) - but I'm pretty happy with what I'm seeing so far.

I'm post training the attention and the shared expert, and leaving the base experts frozen - this is a behavioral and logic fine tune that preserves the original yandex training data.

I plan on, within the next few days, releasing a few gguf quants of this, along with a llama.cpp patch for running it locally. I'm not sure how well the initial fine tune is going to work out - loss looks good but I'll have to do some evaluating. Either way, I plan on continuing training with reinforcement learning and an extended SFT set, as I have room and a ton of capacity left in my QLoRA adapter. I'll release this version as a public checkpoint anyways though (kinda like how deepseek did it) so people can play around with it and hopefully get excited for new checkpoints.

Cheers! Stay tuned, this is a pretty fun model size to play with, I'm excited to release the instruct version. I'll open source whatever you guys want out of this - I already open sourced the distillation engine (see SFTMill, it's been posted in here in the last few days) - but I also have a custom kernel for training this for V100s and a few other patches I can share (this training has been plugging away on 3, 32gb v100s - man it took a while to get that to work). Mandatory plug for my own goals: if you're hiring remote or in NYC for a dev or ml engineer, hit me up!

God I hope it writes the adapter when this is done I didn't audit that code well enough.

💬 18 (+1) open on reddit ↗
▲
57
+3
54👁
r/LocalLLaMA · u/Effective-Ad2060 · 9d ago
We benchmarked 18 RAG pipelines against an agent loop on Google's FRAMES. The best pipeline hit 78.9%. The agent loop hit 92.7%.

Hybrid search, reranking, query decomposition, and query expansion are often treated as must-haves for good RAG. We wanted to see how much each actually helped, so we tested them. Same model, same embeddings, same documents, across all 824 multi-hop questions in FRAMES.

We built 18 pipeline variants. The best one scored 78.9%. Our agent loop (with retrieval tools) that could read the results and search again scored 92.7%—roughly the same as giving the model the right articles upfront.

The reranker results might surprise you. A small reranker dropped our best pipeline’s accuracy by 9 percentage points, while a larger one barely helped. I’d already suspected reranking wouldn’t help much here, but wanted to test that assumption.

Another thing we noticed: models sometimes fill in gaps from memory, even when you explicitly tell them to stick to the retrieved documents. Those answers can still be full of citations. We ended up checking every correct answer against what the system had actually read.

Here’s the write-up if you’re interested:
Agentic RAG vs. traditional RAG on FRAMES

Full disclosure: I work on PipesHub, which is open source. The benchmark code and runbook are in the repo: https://github.com/pipeshub-ai/pipeshub-ai/tree/frames

Quick note on what the numbers mean: they're end-to-end answer accuracy, not retrieval scores. Every answer was graded by an LLM judge (Claude Sonnet 5) using the FRAMES paper's own grading prompt, and independently by a second judge (Gemini Flash 3.8). The two agreed on almost every answer (Cohen's κ 0.93–0.98). We also checked each correct answer against the text the system was actually shown, so answers that came from the model's memory don't count as retrieval wins.

💬 54 (+4) open on reddit ↗
▲
54
+14
37👁
r/LocalLLaMA · u/Designer_Cost8989 · 9d ago
Index-Translate: 150 text languages, plus document translation, multilingual subtitles and dubbing

Quick update: we’ve opened a free public API for Index-Translate-35B-A3B! It’s OpenAI-compatible, and you can get started with our Python script—no extra dependencies needed.

I'm part of the BiliBili Index LLM team. We're sharing Index-Translate and its companion models for translating text, documents, and videos.

Index-Translate supports 150 text languages, with 2B, 9B, and 35B-A3B (preview) options. You can specify terminology, writing style, and output format—for example, keeping product names consistent, translating in a casual tone, or preserving JSON and placeholders during localization.

There are also models for more specific workflows:

  • Index-NativeLong: translate whole documents, using their context to help keep names and terminology consistent across passages.
  • Index-Homura: set a syllable budget for translated lines, useful for fitting a dubbing script.
  • Index-Echo: generate multilingual subtitles or translate speech into speech, using the source speaker's voice as a reference.

The attached video shows English → Japanese dubbing, followed by an English clip with subtitles in six languages. The 150-language coverage applies to the text models; Echo supports a smaller set of language pairs.

https://reddit.com/link/1wugf2t/video/9zajzj72zpsh1/player

Code and released weights are Apache-2.0.

Try the demo · GitHub · Models

What would you try it on—video subtitles, game localization, or documents? We'd especially appreciate examples where it gets your language pair wrong.

💬 22 (+3) open on reddit ↗
▲
54
+3
22👁
r/LocalLLaMA · u/jjusko20 · 10d ago
Update: Yandex/AliceAI 80B-A3B fine tune progress

loss curve \(taken from the last micro of every step, to explain the variation\)

some help from gemini 3.8 flash high

About 40% of the way done with the initial fine tune. The loss is so spiky because I accidentally used the last loss of each micro, rather than the average of each step

The training live stream is at: https://figure-bios-expect-cio.trycloudflare.com/ \- and it allows you to inspect any and all of the training data I'm using, if you're interested - I can also provide those roughly 3.5k examples as a dataset. It was generated from sftmill

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wslgkw/watch\_me\_posttrain\_aliceaifoundation80ba3b\_from/

▲
43
+10
36👁
r/LocalLLaMA · u/SnooPredictions515 · 8d ago
Running 95.5 GiB Qwen3.8-Flash-Next at 41–52 tok/s on a 64GB Mac (1.76x faster than llama.cpp): Slipstream release, 130k context scaling, + Swift variant

I've been working on getting the 95.5 GiB Qwen3.8-Flash-Next model to run fast on a single 64GB Mac. In my earlier post, I shared a custom expert-streaming fork of llama.cpp . It worked, but decode capped out around \~23–27 tok/s and slowed down as context grew.

Today I'm releasing Slipstream: a compiled C++ Metal inference engine with native SSD expert streaming and speculative drafting for Apple Silicon.

The main result: If you already downloaded my original V3 model (34k+ downloads), you don't need to re-download anything. You can run that exact checkpoint on Slipstream for a 1.76x speedup: 41–52 tok/s (up from 23.1 tok/s in llama.cpp) on the same 64GB Mac.

Even better: decode speed doesn't collapse at long context. Across 3,086 live requests in real coding sessions, it stays flat at 33–44 tok/s all the way out to 130,000 tokens.

Previous posts for context:

Open source resources:

1. How to run your existing V3 model on Slipstream

If you have the model from the last post (~/models/qwen38-flash-next-v3), you can point Slipstream directly at it.

Step 1: Clone & build (under 1 minute)

git clone https://github.com/npanj/slipstream.git
cd slipstream
make -j4

Step 2: Download the model (if you don't already have it)

Downloads the 3 GGUF shards + MTP draft head (~95.5 GiB total) huggingface-cli download nitinpanj/qwen38-flash-next-v3 \ --local-dir ~/models/qwen38-flash-next-v3

Step 3: Raise wired GPU memory limit & serve

Raise wired GPU memory limit once per boot (required on 64 GB Macs): sudo sysctl iogpu.wired_limit_mb=59392 # Serve your existing model: ./slipstream serve --model ~/models/qwen38-flash-next-v3 --port 8090

First Run Note: On first launch, Slipstream detects the multi-shard GGUF files and prepares optimized streaming package files into <model>/prepared/ (\~5–7 minutes). Subsequent launches load in \~10–15 seconds.

The server exposes a standard OpenAI-compatible API (http://127.0.0.1:8090/v1/chat/completions) ready for curl, Oh My Pi (omp), Claude Code, or OpenCode.

2. Speed: llama.cpp Fork vs. Slipstream (Same V3 Checkpoint)

Here is a direct head-to-head comparison running the exact same 95.5 GiB model files across 6 reasoning and coding tasks on the same M5 Pro (64 GB unified memory, temperature 0.0):

|Domain / Task|Prompt Task|llama.cpp Fork|Slipstream|Speedup|llama.cpp TTFT|Slipstream TTFT|
|:-|:-|:-|:-|:-|:-|:-|
|Math Reasoning|GSM8K (eggs problem)|24.0 tok/s|43.6 tok/s|1.82x|4,024 ms|2,337 ms|
|Math Derivation|MATH-500 series ($p - q$)|24.3 tok/s|43.1 tok/s|1.77x|1,655 ms|1,587 ms|
|Constraint Logic|3-chair deduction|25.4 tok/s|46.0 tok/s|1.81x|1,469 ms|1,042 ms|
|Python Coding|merge_intervals ($O(N log N)$)|19.7 tok/s|35.0 tok/s|1.77x|1,507 ms|1,070 ms|
|Systems Coding|Rust CSV parser|22.7 tok/s|37.5 tok/s|1.65x|1,257 ms|859 ms|
|Tech Writing|Multi-head attention|22.5 tok/s|39.4 tok/s|1.75x|1,267 ms|843 ms|
|AVERAGE|Across all 6 tasks|23.1 tok/s|40.8 tok/s|1.76x|1,863 ms|1,290 ms|

https://preview.redd.it/vh1danbnzwsh1.png?width=3000&format=png&auto=…

What made Slipstream faster:

  1. Asynchronous layer-ahead prefetch (fcntl(F_RDADVISE)): In llama.cpp, synchronous page reads for missed expert matrices stalled the GPU on NVMe latency (\~475 ms per chunk). In Slipstream, non-blocking read-ahead hints stream upcoming expert layers from SSD into RAM while the GPU is still executing the previous layer, cutting prefill staging latency by 28%.
  2. Hybrid MTP + Prompt Lookup speculation: During tool calls and code generation, Prompt Lookup Decoding (PLD) matches prompt anchors in under 50 ns with 0 allocations, preventing draft rejections. This lifted tool-calling decode from 5.6 tok/s to over 45 tok/s.
  3. Metal GPU-mapped n-gram tables: llama.cpp faulted on the 26.8 GiB n-gram table during prefill. Slipstream maps and gathers n-gram embeddings directly in Metal kernels.

3. Context Scaling: Real Telemetry up to 130,000 Tokens

On standard Transformers, decode slows down sharply as context grows because the KV cache swells and memory bandwidth saturates.

Qwen3.8-Flash-Next avoids that through its hybrid architecture:

  • 48 recurrent linear DeltaNet layers (fixed $128 \\times 128$ hidden state, $O(1)$ memory growth with context).
  • Only 16 full-attention layers.

Here is actual telemetry collected across 3,086 live requests during real agent coding sessions on my M5 Pro (64 GB):

|Context Range (Tokens)|Live Runs|Average Decode|Median (p50)|Peak Decode|Average TTFT|Notes|
|:-|:-|:-|:-|:-|:-|:-|
|< 1,000|314|41.5 tok/s|41.9 tok/s|59.8 tok/s|2.16 s|Short baseline|
|1k – 4,000|21|41.0 tok/s|42.5 tok/s|64.5 tok/s|5.26 s|Small documents|
|4k – 8,000|58|43.6 tok/s|43.2 tok/s|67.2 tok/s|7.36 s|Code review turns|
|8k – 16,000|117|43.6 tok/s|44.6 tok/s|58.2 tok/s|7.91 s|Multi-file context|
|16k – 32,000|562|38.2 tok/s|40.9 tok/s|58.0 tok/s|13.59 s|Deep agent session|
|32k – 64,000|1,029|35.0 tok/s|37.5 tok/s|55.6 tok/s|13.24 s|Large repo refactor|
|64k – 96,000|650|32.4 tok/s|34.7 tok/s|53.9 tok/s|12.81 s|Multi-turn transcript|
|96k – 130,000|364|32.9 tok/s|33.3 tok/s|43.8 tok/s|7.95 s|Cache-hit deep turns|

https://preview.redd.it/qruz5nmpzwsh1.png?width=3300&format=png&auto=…

Takeaway: Decode speed stays between 33 and 44 tok/s all the way out to 130k tokens. Even at 130k context, it generates tokens faster than stock llama.cpp did on a 500-token prompt.

4. Optional: Swift KV-Sparse Model Variant

If you want higher reasoning accuracy and lower KV cache memory, I also put together an optional Swift variant of this model: Swift-Qwen3.8-Flash-Next-V3.

What Swift changes:

  • KV-Sparse Attention: Replaces standard dense attention with KV-sparse layers distilled from Swift-1.5, cutting down RAM pressure at long contexts.
  • Spliced Q8 Donor Backbones: Slices 686 high-precision Q8 donor tensors into resident backbone layers for sharper representations.
  • Concise Reasoning: Distilled to eliminate repetitive thinking loops in deep contexts.

Both models run on Slipstream using the exact same engine command. Here is how they compare across 145 paired evaluation problems (temperature 0.0, seed 1234):

|Domain / Benchmark|Items|Original Flash-Next V3|Swift-Flash-Next V3|Accuracy Delta|Original Decode|Swift Decode|
|:-|:-|:-|:-|:-|:-|:-|
|AIME 2025|20|45.0% (9/20)|45.0% (9/20)|0.0%|44.3 tok/s|44.3 tok/s|
|MATH-500 (L4–5)|35|60.0% (21/35)|62.9% (22/35)|+2.9%|44.8 tok/s|44.8 tok/s|
|GPQA Diamond|35|45.7% (16/35)|54.3% (19/35)|+8.6%|44.8 tok/s|44.8 tok/s|
|GSM8K|25|96.0% (24/25)|96.0% (24/25)|0.0%|45.6 tok/s|45.6 tok/s|
|HumanEval|25|92.0% (23/25)|92.0% (23/25)|0.0%|40.6 tok/s|40.6 tok/s|
|Hard Systems Logic|5|100.0% (5/5)|100.0% (5/5)|0.0%|39.2 tok/s|39.2 tok/s|
|OVERALL|145|67.6% (98/145)|70.3% (102/145)|+2.8%|43.9 tok/s|44.4 tok/s|

https://preview.redd.it/gt1edfvrzwsh1.png?width=3000&format=png&auto=…

To run the Swift model instead:

huggingface-cli download nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF \
--local-dir ~/models/swift-qwen38-flash-next-v3

./slipstream serve --model ~/models/swift-qwen38-flash-next-v3 --port 8090

5. Foundation for Qwen4

The core primitives in Slipstream:

  • 512-route sparse MoE streaming with SSD prefetch
  • QSA (Quasi-Sparse Attention) indexer & selection kernels
  • Hyper-connection mixing and per-layer embedding gathers
  • Metal GPU-mapped n-gram table gathers
  • Single-lane speculative verification with PLD & MTP

...were built around this hybrid architecture. If Qwen4 adopts a similar blueprint (hybrid linear recurrence + sparse attention + routed MoE experts), Slipstream should be able to run Qwen4 locally on consumer unified memory hardware on day one.

6. Hardware Tested & Porting to NVIDIA / AMD

  • Hardware tested: All testing and benchmarking were done on an Apple MacBook Pro (M5 Pro, 64 GB unified memory, 2 TB SSD).
  • CUDA / ROCm ports: I don't have access to modern NVIDIA or AMD GPU hardware, so I can't build or test CUDA/ROCm backends myself.
  • If you have hardware and want to help port this: If anyone in the community has NVIDIA or AMD hardware and wants to help bring expert streaming and hybrid speculation to Linux/Windows, I'm happy to help collaborate on the port. Feel free to open an issue on the repo or DM me.

7. Credits & Upstream

  • Splash Team (Incoai): Full credit to the creators of Splash. Their C++ Metal speculative decoding design and memory architecture provided the foundation for this work. I will prepare a clean PR proposing these Flash-Next and SSD streaming extensions to the Splash upstream repo in case they want to incorporate them.
  • ds4 Team: For their insights on Metal router numerical precision (Taylor polynomial softplus expansion) and streaming scheduling designs.
  • Qwen Team: For training Qwen3.8-Flash-Next and releasing the hybrid linear MTP architecture.
  • ukisai: For the Swift-1.5 distillation work enabling KV-sparse reasoning.
  • bartowski & unsloth: For donor quants and quantization tooling.
  • mihailescu2m: For the initial expert streaming work in llama.cpp.
💬 29 (+2) open on reddit ↗
▲
42
+5
29👁
r/LocalLLaMA · u/bring_back_the_v10s · 10d ago
How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks?

I've been closely following the rise of the Strata inference engine and as someone with 28GB VRAM and 32GB RAM I'm itching to buy 32GB more RAM just to use Flash Next. But of course before I make such a financial commitment as a member of the GPU-poor class like myself, first I need to make sure the IQ3 quants are worth it. My use case is primarily agentic coding tasks with harnesses like Pi or OpenCode.

Has any of you guys used Flash Next IQ3 for relatively serious coding? Is it worth it? Compared to, say, Qwen 3.8 27B Q4 or Q5.

https://github.com/Niko1221/Strata/

💬 72 (+5) open on reddit ↗
▲
41
+12
29👁
r/LocalLLaMA · u/Terminator857 · 9d ago
China and the memory market

Once china sets its goals for dominating a market it wins. Usually takes many years, but it happens. Can't compare the political will of a country versus profit and loss thinking of a corporation.

China will eventually win in the memory market and current memory makers are at an unfair disadvantage.

https://www.tweaktown.com/news/112680/chinas-cxmt-is-on-track-to-nearly-match-microns-dram-production-capacity-by-the-end-of-2026/index.html Quote:

CXMT will finish 2026 with approximately 350,000 wafer starts per month (WSPM) of DRAM capacity, which is just 25,000 WPM less than Micron.

... by 2030, its total capacity will increase to around 1.41 million WSPM, according to Citrini. CXMT alone is projected to build new production capacities in Beijing, Hefei, and Shanghai, to expand its production capability to 950,000 WSPM in 2030, assuming everything goes as planned.

/end quote

Is China hoping for a RAM price drop crash to extinguish the competition?

Additional references:

  1. https://www.tomshardware.com/pc-components/dram/chinas-cxmt-targets-30-percent-dram-memory-market-share-by-2030-with-sixth-mega-fab-future-plans-bottlenecked-by-access-to-advanced-chipmaking-tools
  2. https://www.trendforce.com/news/2026/09/24/news-cxmt-ymtc-ramp-memory-capacity-but-chinas-ai-cloud-boom-could-soak-up-new-supply-through-2027/
  3. Microns quarterly report: https://www.investing.com/news/company-news/micron-fq4-2026-slides-record-revenue-ai-demand-drives-supply-tightness-93CH-4926074 . Turn off javascript to view.
💬 34 (+3) open on reddit ↗
▲
40
+4
24👁
r/LocalLLaMA · u/Qual_ · 9d ago
Astrabox - Open source Arcade Game Generator post image

Hey everyone! I’ve been working on ASTRABOX: an arcade interface where you describe a game, the AI builds it, and you can ask for changes by voice while playing.

Each game gets its own visuals and gameplay, while a shared runtime handles controllers, scores, player joining, etc.

I built it around Codex, but the code is open source. I’d love to see someone adapt it to a local coding model, local STT/TTS, and a different harness.

It runs as a local web app—you don’t need a Raspberry Pi or an actual arcade cabinet, although that’s what inspired the project 🕹️

It’s still experimental, but feel free to customize it, change the little robot, the environnement, or everything.

https://github.com/Qualzz/astrabox

Curious what models and tools you’d use for a local version.

Edit: Clanker helped me with writing this message in english.

https://preview.redd.it/rru52ja4brsh1.png?width=2224&format=png&auto=…

https://preview.redd.it/fm3m19j2brsh1.png?width=1280&format=png&auto=…

💬 11 (-1) open on reddit ↗
▲
38
+7
32👁
r/LocalLLaMA · u/SultanGreat · 9d ago
What's the best setup for Qwen3.8 27b for a 16 gig VRAM?

Hello guys!

I have been experimenting with qwen 3.8 for a long time and I hadn't been able to get reasonable speed. I am on a 5060Ti 16 GB, and although this gpu can game, I am aware that AI demands more than 16 GB.

I am on a Fedora 44, AMD Ryzen 9600x and 16 GB system ram (16 GB system ram and 16 GB vram, totaling to 32 GB) and I would like to use llamacpp, although I would use any other tool if I could if it meant faster speed.

I am looking for a large context. Atleast 128k context. The first question is, what quantization to pick? In my experience Q3 UD was satisfying, but I am looking for uncensored model. In my experience, MTP has never lived up to its hype for me (and I don't know why!?), which is why I am thoroughly lost on making a good setup after an honest week of experimentation, which is why I have resorted to ask here as a last resort.

Update : Found a model, thanks to u/_wortkarg_

link : https://huggingface.co/RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3\_XXS-Uncensored

command (A better command would be appreciated and updated accordingly):

~/llama.cpp/build/bin/llama-server \
--model ~/Documents/Models/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-mtp.gguf \
--alias "llamacpp" --host 0.0.0.0 --port 8001 \
-ngl 99 --flash-attn on --ctx-size 131072 \
--cache-type-k q4_0 --cache-type-v q4_0 \
--parallel 1 --batch-size 512 --ubatch-size 256 \
--no-warmup --jinja \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0

I am hitting at about 35 t/s+ speed with this one.

💬 84 (+3) open on reddit ↗
▲
35
+7
31👁
r/LocalLLaMA · u/Any-Winter-4079 · 8d ago
DDR4/PCIe4 vs DDR5/PCIe5 for LLMs- I benchmarked them for pre-training. What are your thoughts? post image

Hello everyone.

I've recently ran some experiments comparing DDR4/PCIe4 and DDR5/PCIe5 for AI workstations on a pre-training run, and would like to hear yours thoughts.

First of all, and as a summary of my results ( code here: https://github.com/Any-Winter-4079/DDR4-PCIe-4-vs-DDR5-PCIe-5-for-CUDA-training ), I rented two machines on Vast.ai, one with an H12SSL-i motherboard, an EPYC 7352, 192 GB of RAM and of course using PCIe4 (26.3 GB/s) and another with a WRX90E-SAGE SE motherboard, a 9975WX CPU, 256 GB of DDR5 RAM and PCIe5 (54.3 GB/s), and DDR5/PCIe5 is about 15-20% faster on pre-training (depending on whether you include or exclude validation and other costs) under the same number of GPUs.

With the current RAM prices, however, for the cost of 256 GB DDR5 RAM at 6400 MT/s you can get a full (extra) RTX PRO 6000 WS/Max-Q, at which point the comparison clearly favors DDR4/PCIe4 (with 2 GPUs), with about 50% extra throughput vs a single GPU at equal(ish) cost.

Now, there aren't a lot of downsides in my mind to choosing DDR4/PCIe4, but there can be a few:

  1. at least some of these DDR4/PCIe4 motherboards are on the older side, and were one to need replacement, they are not so easy to get (for example, the H12SSL-i used, I can only find it for sale as refurbished now, so who knows in a few years if it will even be available for retail).
  2. newer GPUs (as in, new NVIDIA generations) may stop working with older motherboards (meaning yes, PCIe is backwards compatible but the motherboard's BIOS/UEFI sometimes has issues during POST with newer GPUs (e.g., some older PCIe3 motherboards already have trouble recognizing Blackwell cards, and this may be the case for PCIe4 and newer cards in the future). Meaning if one were to buy a newer GPU in the future, the whole workstation may not be suitable.
  3. If one has to go ahead and bite the bullet on RAM prices, the million question is when? To me prices now are \*\*awful\*\* but so were RTX PRO 6000 prices and here we are (i.e., even higher).
  4. For pre-training, I would still choose DDR4/PCIe4, but I suspect for inference DDR5/PCIe5 might be a fair bit better than the 15-20% that it gives you on pre-training, plus we might be moving into some techniques soon such as dynamic expert/data loading into the model at runtime, which again would favor better DDR/PCIe speeds.

With all of this, I am curious if anyone has benchmarked this, and what are your thoughts on it. Would you hold out on DDR5 at the moment, and therefore go for PCIe4, or would you bite the DDR5 bullet early? Another issue with RAM is channels and DIMM count/channel, because if you want to go 'cheap' like let's only get 192 or 256 GB of DDR5 on 8 channels at 1 DIMM/channel (e.g., 8x32 to get 256), then upgrade to 512 later (when budget allows), that means you have to replace your full RAM (because all the slots are occupied, requiring new 8x64 to get 512 for instance)... And if you get fewer DIMMs like 4x64 to get 256 GB (leaving 4 DIMM slots unoccupied) then you get half the bandwidth because only 4 channels are populated. So maybe a machine that has dual DIMM support per channel is the answer to this (fully populating 8x32 to get the full bandwidth, and still allowing you to expand to another 8x32 to get 512), but in general it's a tricky point too.

So, what do you do/are you guys doing? Have you recently bought a workstation or upgraded to one, for pre-training, fine-tuning, RL, inference, whatever your use case may be, and come up with this dilemma? Are you choosing DDR4/PCIe4 as it would seem reasonable or are you going for DDR5/PCIe5 and if so, why? I am interested in all use cases and opinions!

💬 19 (+1) open on reddit ↗
▲
23
+1
32👁
r/LocalLLaMA · u/darklordfireape · 10d ago
Update: Strix Halo + R9700 with llama-halo-hybrid - now beats DGX Spark

Hi folks, I've spent the last couple of months experimenting with Strix Halo and previously I released a proof of concept I called llama-halo-hybrid. I've continued updating it and it now performs very well. The idea is that you can take an R9700, or similar, and place dense parts of the model, KV, and some of the layers on the GPU and let the APU take the rest of the model. You can add the extra GPU through a PCIe extender (framework desktop), Occulink, or a thunderbolt dock depending on which machine you have. Detailed notes along with code in the repo on github. I'm not selling anything, this is all 100% open, MIT-licensed.

It breaks 60+ tok/s decode and 2000+ tok/s prefill, supporting full 256k context.

This is not some custom inference engine that requires a custom quant to run. This is llama.cpp modified to run whatever you want, albeit mostly tuned for Qwen and GLM families. After continuing to tinker with it, it now performs better than DGX Spark (albeit cheaper) running Qwen-3.8-flash-next and slightly better yet with the Swift-1.5 variant. Most of my testing was done with the Q4/Q4\_K\_XL models to balance size and quality.

Note \- if you are just using Strix Halo by itself, this is probably not the right tool. Check out gufo, which looks very promising.

https://github.com/sixvolts/llama-halo-hybrid

I would love any feedback you all have and happy to investigate tuning for different "sidecar" GPUs other than the R9700 if there's demand and I can get my hands on one.

UPDATE (10/4): I added some more docs around using different cards besides the R9700. The R9070, V620, and 7800XT all perform very well and I put quick guides for those configs, along with notes on thunderbolt/USB4 setup and the dual-machine setup I used here:
https://github.com/sixvolts/llama-halo-hybrid/tree/main/halo-cookbook

💬 27 (+9) open on reddit ↗
▲
22
+1
12👁
r/LocalLLaMA · u/norenEnmotalen · 9d ago
Unsloth, Swift1.5, Peculiar-Ragdoll, ThinkingCap - Qwen3.8-27B

In a previous post I shared comparison between Swift1.5 and peculiar-ragdoll's checkpoints. Added the original unsloth Q4\_K\_XL and ThinkingCap Q4\_K\_M (they don't offer L or XL) to the comparison. Here are the results over a 69 set of eval questions.

All tests are now run at same "medium" reasoning effort.

unsloth-ud\_q4\_k\_xl one ran using llama.cpp - not the splash forked inference engine.

https://preview.redd.it/tztb66wygvsh1.png?width=2958&format=png&auto=…

I'll do a 3x repeat for the slow run to see if it maintains 69/69 each time.

EDIT: u/jucabala457 asked I test mradermacher/Signal-3.8-27B-Terse-Coder-i1-GGUF The Q4\_K\_M is closest quant available. A nice addition for sure! That GGUF couldn't run with Splash-based engine due to tensor incompat. I ran it using llama.cpp the slow way. The total time taken isn't a fair comparison for that reason. Updated results below

https://preview.redd.it/pvfvgd35twsh1.png?width=2976&format=png&auto=…

I also just made the tuieval tool available here https://github.com/ashe-wb/tuieval

Can't promise you the tool will work right away on your install since a fully local binary is what I've been using and testing with. Customize it with packs of domain-specific eval questions you deal with on the daily. This is the most important part. A model or fine-tune that is not good for one thing might be excellent for something else and only you know what your domain interests are. The ability of a model to render game graphics means nothing to me but it means everything to someone else.

https://preview.redd.it/m2n4rzlzjwsh1.png?width=2000&format=png&auto=…

▲
21
+2
11👁
r/LocalLLaMA · u/Any-Lingonberry7411 · 9d ago
Best local model for Blender and game dev?

I have been looking at some local models, but even the smartest ones like GLM5.3 Flash and DSv4 Flash have a hard time creating coherent models in Blender and placing them logically in game engines.

Is this something that local models are just too dumb still to do good job at?

▲
21
 
14👁
r/LocalLLaMA · u/pilkyton · 10d ago
PSA: ModelScope CLI is now moved to "modelscope-hub"

To save people 30 minutes of research (because they didn't bother documenting this officially at all):

  • The "modelscope" package is now just the library. Doesn't contain a CLI anymore. If you try to install it or update your old CLI package, you get "No executables are provided by package \modelscope\; removing tool. error: Failed to install entrypoints for \modelscope\".
  • They moved all CLI tools to "modelscope-hub".

The new command to install it:

uv tool install "modelscope-hub"

▲
16
+3
26👁
r/LocalLLaMA · u/brainchillzZ · 8d ago
Gufo performance .... 70tps Qwen 3.8 27b but you need to read the fine print.

So everyone has been yelling about how I should be using Gufo instead of halogen because it's open source and it's "just as good or better". Checking in on their GitHub (GitHub.com/gufo-org/gufo) got me immediately .. "Qwen 27B Q4: 70.56 tok/s single user, 123 tok/s with 8 users" on a strix halo device? Yes please ... So I broke down and tried it today ...

Setup: gufo 0.4.0 from their podman image, Qwen3.8 27B UD-Q4\_K\_XL from Unsloth plus the DFlash2 Q4\_K\_M draft model, using their own benchmark script and their own settings (greedy, thinking off, 128 output tokens, prompt cache off).

If you want the short version ... yeah I got 70.22 tok/s. So the number is real. But the prompt that produces it is "Write the word red exactly 1000 times".

But it's also not real. In that figure all the speed comes from the speculative decoding. The draft model guesses like 7 tokens ahead, the 27b checks them in one pass and keeps what it agrees with. When the output is the same word over and over the draft is right every time. On a real prompt it's right maybe half the time.

Their benchmark has a second set of nine ordinary prompts (some C++, a word problem, a summary, Italian, Chinese, JSON, a bit of fiction, a debugging checklist).

On those:
| | repeat-a-word prompt | normal prompts |

|---|---|---|

| 1 user | 70.2 tok/s | 39.4 tok/s median, anywhere from 22 to 52 depending on the prompt |

| 8 users, their "aggregated" number | 122.6 | 82.4 |

| 8 users, tokens actually delivered per second | 82 | 52 |

About that last row. The "123 tok/s aggregated" figure is each request's decode speed added together, with prompt processing and queue time left out. If you just count tokens coming out of the box per second of wall clock it's 82, or 52 on normal. prompts.

To be fair to the gufo people, none of this is hidden. Their benchmark docs have separate "mixed" and "repetitive" columns and the mixed numbers they publish match what I got. It's only the repo description and the top of the README that lead with the best case. And 39 tok/s from a 27B at Q4 on an APU is still really good. Without the draft model their docs put it around 12.

The other thing I wanted to know was how it compares to halogen (peonist-ai/halogen-flash-server), which is what I normally run. Both can serve Qwen3.8 Flash-Next, so I put that on both and sent the same prompts to each. Two boxes, same hardware, same OS image. Greedy, thinking off, 256 tokens.

| | halogen 0.13.8 | gufo 0.4.0 |

|---|---|---|

| nine normal prompts, average decode | 43.9 tok/s | 38.2 tok/s |

| 4 users at once, end to end | 76.6 tok/s | 63.1 tok/s |

| cold prompt processing, \~9.7k tokens | 1288 tok/s | 1495 tok/s |

| the "red" prompt | 56.9 tok/s | 87.4 tok/s |

So for everyday generation halogen was about 13% faster for one user and about 18% faster with four. gufo was 16% faster at chewing through a long prompt and a lot faster on the repetitive one.

I'll add this just in case, because someone will ask or at least try to poke about it in the comments

\- I know the weights aren't the same. halogen uses its own 4-bit format, gufo uses the Unsloth GGUF. I only measured speed. I did not compare output quality at all.

\- They were two different machines but identical hardware and software, and my boxes have agreed within 1% on other benchmarks, but it's still two machines.

\- One run each was done for the head to head. The reproduction of their numbers was 3 reps.

\- gufo has shipped four releases over the last four day, so this could all be stale by next week.

There was quite a lot of stuff that I liked about gufo that isn't performance related. It takes plain GGUFs, it's MIT, the 27B loads in about 3 seconds (Flash-Next in 13), the per-request log line tells you draft acceptance and cache hits, and it does 8 batched sessions. It also has ASR, TTS and image models that I haven't touched. Their benchmark hashes the output with and without the draft model and it was identical every time, so the speculative path isn't changing what the model says.

One thing to keep in mind if you try it out is that it reserves memory per session up front. Flash-Next with 4 sessions at 64k context took 94GB.

So it isn't smoke and mirrors exactly. Everything I checked reproduced. Just know that the 70 is a ceiling you'll only hit if your workload is incredibly predictable text, and plan around the 30s for the 27B on normal stuff.

I kept this all setup to tinker with on actual output quality over the next few days, I'm happy to run other prompts or try it with different settings if anyone wants to see something specific.

▲
16
+1
14👁
r/LocalLLaMA · u/caenum · 9d ago
Best OpenSource Claude Cowork alternative?

Hey guys,

Looking for an alternative for Claude Cowork:

  • Project Work / Documents
  • Integrations like Notion, Gmail, etc.
  • Tools like Websearch, PDF creation, etc.

Came over Eigent (https://github.com/eigent-ai/eigent) but cant find any actual reviews about it, what usually is a sign thats not good performing..

Also have tried multiple other frameworks (OpenClaw, Hermes, OpenWebUI Chat Interface) - but those are different use-cases for me.

LLMs will be server through my own server, so should be open for connecting to Ollama, Ninfer, etc.

So anyone knows a good application which behaves like Claude's Cowork?

Thanks )

▲
16
+2
17👁
r/LocalLLaMA · u/Federal-Effective879 · 9d ago
szmcp: a ZIM HTML to Markdown converter and yet another ZIM MCP server

Hello all, I wanted to share a little project I vibe-coded for myself that you may find useful.

As many people here like to suggest, I wanted to give my small local LLMs access to information to improve their world knowledge. I didn't want to give my LLM free reign searching and browsing the web to keep my queries private and functional offline, so I wanted to give them an offline knowledge base. Wikipedia ZIM files from Kiwix were a good starting place for this. Several MCP servers for ZIM files exist, but I didn't like the existing ones I found for various reasons. The most notable one is openzim-mcp , which works in its advanced tool mode but has overly complicated context-bloating tools, and whose simple single tool mode doesn't work very well in practice.

I built my own MCP server for ZIM files in Rust, exposing a simple tool set that's actually easy for small local LLMs to use, while providing all the functionality one normally needs. It's designed mainly for Kiwix MediaWiki ZIM archives generated by mwoffliner (such as Wikipedia, WIkivoyage, etc.) but also usable with many non-wiki ZIM files. I also wrote my own custom HTML to Markdown converter for MediaWiki pages that produces clean, well-formatted Markdown including special content such as wiki infoboxes, LaTeX formulas, tables, etc. It also strips out references and boilerplate sections from wiki pages to keep the resulting markdown clean and context efficient.

You can hook this MCP server to llama.cpp's Web UI to give your small local LLMs much better world knowledge. A system prompt that I found works well is:

You are a helpful assistant. When answering factual queries, search through Wikipedia using the provided ZIM access to ground your answers. If the articles or sections you read don't have relevant details, you can search more, but don't keep searching forever; you need to answer reasonably quickly.

I tested it with various LLMs of varying sizes. I got good results with Gemma 4 12B (or bigger), IBM Granite 4.2 8B (or bigger), and Ling 3.0 Flash (best results while still maintaining usable speed on my 128 GB Mac). Qwen 3.6 35B-A3B was usable but tended to overthink and hallucinate; Qwen 3.8 27B was too slow to be usable for this purpose on my Mac. I also experimented with smaller models, and got usable results for simpler queries with MiniCPM5 2B, LFM 2.5 2.6B, and IBM Granite 4.2 3B. Gemma 4 E4B did not work well for this.

I also build a sub-command within this tool to convert entire ZIM files from HTML to Markdown to save disk space (and avoid the need to convert on every tool call). It converts a 49 GB Kiwix nopic full English Wikipedia ZIM file into a 19 GB Markdown ZIM file, while maintaining all article content (aside from references) and maintaining full-text search. Likewise, it converts the 17 GB top-1M nopic enwiki Kiwix ZIM file to 6 GB. You can make the resulting ZIM files even smaller if you specify the option to only index article intros for full-text search (since the full-text search Xapian index is a large fraction of the file size). The converter is multi-threaded and written fairly efficiently using Rust, so you can convert all the millions of articles in a full English Wikipedia Kiwix files in a few hours on a typical modern computer.

GitHub link: https://github.com/sultanqasim/szmcp

▲
16
 
18👁
r/LocalLLaMA · u/SignificantZebra5883 · 10d ago
i would like to learn deeply about fine-tuning local models before burning money

There's so many new techniques like RL, RL LoRA, QLoRA, CPT LoRA.

I believe i would have a usecase for them, but i don't know where to learn, youtube is filled with bad quality tutorials if i just search and the good channels (fireship, bycloud) don't cover these as they're quite new concepts, i guess?.

how can a regular joe like me learn about these concepts in a "practical depth" so i can actually fine-tune qwen 27b successfuly on lets say custom corpus? without spending 100$ figuring out that "oh i didnt even need CPT here" or "well i chose the wrong Rank count! time to start this 2 day run again!"

context and TLDR: im building a legal general purpose chatbot for context, i have a big corpus, but im a bit stuck on what to do next

thanks for reading and any pointers!

▲
16
+3
12👁
r/LocalLLaMA · u/Dev-in-the-Bm · 10d ago
Best approach for automatically tagging local music collection?

I don't use music streaming services much, and listen to music from my own local collection.

I don't use any local streaming servers like Plex or Navidrome, they wouldn't work for me because I use a dumbphone and play music off of my SD card.

I've manually built a bunch of mood based playlists so I can easily pull up a playlist with the music I want, but that's
obviously very tedious and inefficient.

I've been playing around with ML models to automatically add genre, mood, and other tags to my collection, the open models available today are insane.

The thing is I haven't been able to find any polished tools for doing this.

Most of what's available is either CLI or built for streaming servers.

Is there anything I missed?

Should I just setup a streaming server just for tagging the collection, or is there a better way?

💬 23 (+1) open on reddit ↗
▲
14
+3
15👁
r/LocalLLaMA · u/KissMyShinyArse · 9d ago
Strata: how to configure sampling parameters

The top-level README doesn't mention this, but you can add a "sampling" key to your strata-iq3_s.json like this:

{
"sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20},

"exe": "/path/to/Strata/engine/strata",
"args": [ ... ],
...
}

From docs/DETAILS.md:

The run config's optional sampling block sets the defaults for requests that leave the fields out ("sampling": {"temperature": 1.0, "top_p": 0.95, "top_k": 20}); a request's own fields always win, and with no block at all a request without sampling keys decodes greedy.
💬 10 (+1) open on reddit ↗
▲
14
-2
11👁
r/LocalLLaMA · u/norenEnmotalen · 9d ago
peculiar-ragdoll's Dirk-Qwen 3.8-27B vs. UkisAI Swift-1.5 Qwen3.8-27B

EDIT: Post 2 with more model fint-tunes here https://www.reddit.com/r/LocalLLaMA/comments/1wv1ico/unsloth\_swift15\_peculiarragdoll\_thinkingcap/

I have a long list of my own domain specific eval questions that I run to validate which models I can rely on: coding, coding (numpy/pandas), data analytics decision making, local RAG, and voice assistant. It's made up of the types of things I'm likely to deal with on the daily. The test questions vary in dififculty and composition: easy, medium, hard.

System: M1 Max 32c 32GB with context 128K for Dirk and 110K for Swift.

Swift doesn't have XL. So I had to test with the L quant to stay as close as possible.

I ran the eval (using my tuieval tool) on peculiar-ragdoll's Dirk-Qwen3.8-27B-UD-Q4\_K\_XL and Swift-1.5-Qwen3.8-27B-Q4\_K\_L loaded with a modified version of Splash. The "amalgam" is a local I made out of incoai/Splash 1.1 and paperniuk's apple7-m1-kernels. It is modififed a little but not in ways that would alter model performance. I only merged and tweaked for some memory features I like from llama.cpp such as fit context check at the start of a load and personal QoL updates re auto-context manipulations that I don't want to think about, etc.

To say this result surprised me is quite an understatement. It's blown my mind.

When I did the first test a couple of days ago with only 44 questions, I thought it must be a prompt caching issue I missed that Dirk was benefiting from. I validated it is not and ran it against a lot more questions to certify it. It's a legit test outcome.

Dirk-Qwen is much sharper at getting to decisions and responses. The "be brief" instruction that gets passed each time in the chat templste is doing more magic than I had anticipated. It also gets more answers correctly with way less time consumed.

What trips up Swift-1.5 are mostly hard questions. It tries and tries until the 16,384 max token limit per question is reached and it fails with truncation.

Even when you ignore the 16,384 truncation failures and compare the other questions, Dirk token usage comes out on top.

Snipped view... this basically goes on pattern for another 191 unique questions.

https://preview.redd.it/otg3d1rkbqsh1.png?width=1420&format=png&auto=…

More importantly, this behavior is not just in question answering. You can see it in actual code refactor tasks.

On an unrelated note: tne model that has been able to pass a 100% of my eval packs is Opus 5.5. Deepseek Flash 4.1 fp32 got them all right except three.

💬 22 (+1) open on reddit ↗
▲
11
+1
15👁
r/LocalLLaMA · u/jjusko20 · 10d ago
SFTMill: Easily [off-policy] distill any existing LLM with an OpenAI Compatible Endpoint. Turn any behavioral goal into a comprehensive dataset. post image

Disclaimer: Any\* means any model that exposes its CoT without it being censored.

Hey guys - half a tutorial/guide, and half an I built this, so I went for resources. This is something that I created for myself recently when I couldn't find any good existing solution. I wrote this post myself, no AI!

Probably a fair number of you have seen my posts about fine-tuning AliceAI 80B A3B according to my own synthetic datasets. If you did, I'm still fine tuning it on a live stream right now - check out https://figure-bios-expect-cio.trycloudflare.com/ \-- it'll let you inspect any and all of the training data that I generated with this engine. If you have any interest, it's pretty neat! unfortunately that link is optimized for desktop only and I'd have to kill the run to reset it, so u may want to rotate the phone.

That thread was at https://www.reddit.com/r/LocalLLaMA/comments/1wslgkw/comment/pcw4kgw/?context=1&screen\_view\_count=1

That run is using off-policy distillation, and that's I made this for. my training data for that project with this repo, and just customized it for an OSS release. Basically, you create a "curriculum" for your goal - e.g. if I was training an agentic model, I'd need things like tool calls, bug fixing, working in a workspace, tracing errors, etc. You define your curriculum in a yaml file, then an LLM creates tasks based on the curriculum you defined, and the chosen LLM you're distilling from then solves each task, leaving you with a full Q/A set that encompasses your fine tune goals.

I used qwen 3.8 27b on medium to generate the tasks - I'd recommend avoiding anything any weaker than that.

I forked my private repo of this that I've been using into SFTMill, which is basically just the same thing with great documentation and a few steps added to get anyone onboarded rapidly. I created it \[and my original version\] because I couldn't find any existing pieces of software made with this design, for this purpose.

I release it because I enjoy contributing to the community, and there's a vague hope someone will eventually see one of my pieces of work and want to hire me (if you're reading this and you like the project and you need a software/ml engineer remote or in NYC, let me know <3). It makes me happy when my software helps others so I'd love if you let me know if it helped you. Cheers!

Shoutout u/FullOf_Bad_Ideas for helping me with my alice train in areas I wasn't experienced enough in - I threw a Multi-Turn Hybrid-Reasoning (user <> assistant) section in the readme just for you bud, hope it helps.

https://github.com/jackjusko/sftmill

▲
10
+1
11👁
r/LocalLLaMA · u/Biomass23 · 9d ago
tp=6 can work on vLLM, with padding

vLLM's tensor parallel requires that several of the model architecture numbers be evenly divisible by the number of GPUs selected for tensor parallel. This usually means that you can only use a number of GPUs that is a power of two (2, 4, 8, etc.).

I have six GPUs, and I want to maximize my KV cache when running a 27B model. So I tried tp=6. It choked with various messages, regarding this or that, which needs to be evenly divisible by six, but wasn't.

So I made those things divisible by six.
I asked the robot to come up with a converter that would take the original model, and pad it with zeroes until everything was divisible by 6. It took a few tries, but it worked.

I have Qwen 3.8 27B at BF16 with 256k context window running across six 7900 XTX GPUs. Token generation is about 50 t/s single user, or 200 t/s aggregate with 8 concurrent prompts.

GPU KV cache size: 520,784 tokens, Maximum concurrency for 262,144 tokens per request: 1.99x

▲
9
 
13👁
r/LocalLLaMA · u/nonlinearsystems · 9d ago
M5 Ultra - Qwen3.8 Flash Next vs Laguna S 2.1 post image

Spent today running a same-day, same-harness shootout between Qwen3.8-Flash-Next (oMLX, 182GB oQ8e, MTP) and Laguna-S-2.1 GGUF (LM Studio, 128GB, 8bit) on a Mac Studio M5 Ultra 256GB. Both capped at 262K context, thinking on, unique content per run with zero cached tokens verified each time.

That last part matters because my first run was wrong hah... shared prefixes across sizes let the KV cache carry over and 200K "prefilled" in 21s.

Prompt Qwen Laguna
8K 2.0s 10.3s
32K 7.4s 30.4s
64K 14.7s 70.8s
131K 30.1s 217.2s
200K 47.1s 455.4s

Qwen holds \~4,200 tok/s linear which is amazing. Laguna degrades superlinearly (quadratic attention doing quadratic attention things). At 200K, prefill is 94% of total time on both.

Decode (tok/s): Qwen 59-74 across sizes (MTP at 70-76% acceptance per server logs, roughly 2x). Laguna 68 down to 34 as context grows. No speculation on Laguna, its DFlash path already lost to plain decode on this hardware in earlier testing. I think if Laguna could get DFlash figured out or MTP, this might be a different conversation.

Quality was a draw, 4/4 each, on four problems with script-verified answers (Muse created the gymnastics here: exact 9-digit combinatorics, interval code with 12 hidden tests, fresh knights/knaves, asyncio ordering trap). Opposite styles though: Laguna answers in 5-10s with a few hundred tokens, Qwen deliberates exhaustively (one answer took 119s / 11K tokens). Both burned a full 8K budget on hidden reasoning with zero visible output exactly once, then converted on a 16K retry.

Happy to answer methodology questions. Full writeup with charts and the test rig diagram: https://echalupa.com/blog/qwen-flash-next-vs-laguna-200k

▲
9
-1
24👁
r/LocalLLaMA · u/stevyhacker · 10d ago
Five local models, 6.8 GB of weights: my open-source Mac meeting notetaker

https://preview.redd.it/xs01ds930nsh1.png?width=4800&format=png&auto=…

Back in July I shared LokalBot here. It's a free, open-source Mac app that records your meetings and keeps a daily summary of your activity, all on-device.

0.9.2 came out today. Since July I've benchmarked every model in it and swapped most of the defaults for smaller ones. The whole stack is now 6.8 GB:

  • Qwen3-ASR 1.7B (MLX, 8-bit): transcription
  • Nemotron 3 (Core ML): who spoke when
  • Qwen3.5 4B Q4\_K\_M (llama.cpp): notes and action items
  • Harrier 0.6B Q8\_0: search embeddings
  • LFM2.5 1.2B Q4\_K\_M: autocomplete in any app
  • Apple Vision: screen OCR (opt-in)

A few numbers from my M4 Max (48 GB):

  • 26-min meeting to finished notes in 33 s warm, \~85 tok/s decode
  • Speaker error went from 43.4% to 14.6% DER on AMI. That's against my old pyannote setup, so it says more about my config than about pyannote.
  • Autocomplete p95 went from 1.83 s (Gemma 4 E4B) to 0.49 s

There's also a read-only MCP server and CLI, off by default, so Claude Code or any other MCP client can pull context from your meetings.

I don't have any 16 GB or other M series numbers yet. If you've got one of those, especially M5 or M6 I'd love to see what you get.

I also tried MiniCPM5 2B for notes. It was smaller and faster, but it got stuck repeating itself on one summary and assigned action items to the wrong person. I kept Qwen3.5 4B as the default as saving a few seconds wasn’t worth getting who agreed to do what wrong.

💬 4 (+1) open on reddit ↗
▲
9
+4
27👁
r/LocalLLaMA · u/whatyathinkk · 10d ago
Do I need a UPS?

I know I could ask in some hardware subreddit, but I'm curious to know what people with multiple GPUs and expensive inference setups think about this.

I just moved to a new place and here the lights go out pretty frequently. 3 times over the last week, I came back to my computer being off due to a blackout (I guess it's a blackout, the entire neighborhood looses light for a few seconds/minutes). I have a desktop computer with 2x RTX5080s.

Do I need to buy a UPS to protect my computer from this? I get mixed answers about this topic. I don't mind my workflows being interrupted when the computer turns off, the only thing I'm worried about is the hardware being damaged. I have a good PSU, is that enough to protect the hardware?

💬 75 (+2) open on reddit ↗
▲
8
-2
15👁
r/LocalLLaMA · u/junior600 · 8d ago
What local AI model is good for game decomps/recomps?

Hello guys. Recently, there has been a boom in game decomps and recomps thanks to AI. If you look at the r/decomps and r/recomps subreddits, you can see it. They mostly seem to be using Claude or Codex.I wonder if it would be possible to do something similar with a local AI model. Could Qwen 3.8 27B Abliterated actually handle something like that locally? Does anyone have any experience with this? I don't have a particularly powerful rig (RTX 3060 12 GB VRAM and 24 GB DDR4 RAM), but I can run MoE models comfortably. Even Qwen 3.8 27B IQ3\_XXS dense lol.

Sorry for my English BTW.

💬 31 (+1) open on reddit ↗
▲
8
+4
23👁
r/LocalLLaMA · u/IngwiePhoenix · 10d ago
Penalties of PCIe generations? (2x R9700)

I just bought the GPUs after deliberating and debating for over two years. With costs not coming down any time soon and me just wanting to get this massive todo-box ticked, I decided to just YOLO it; the GPUs are the most volatile, followed by RAM, rest seems more or less stable.

But actually, RAM is one of the reasons I am unsure about wether to chose a SP4, 5 or 6 based board. I am most familiar with AMD CPUs, so that is where my tendencies lie. Unfortunately, RDIMMS are going to absolutely undress me... x.x

However, if I could stick to a DDR4 / PCIe Gen4 setup, that would save a pretty penny. Now I do not intend to offload to system memory, but even a small, single-stick of DDR5 RDIMM is stupid expensive - DDR4 is fine.

The question is: What is the penalty of PCIe Gen 4 versus 5 in regards to inference? I will be using llama.cpp with ROCm, fronted by llama-swap, utilizing both GPUs for inference and VRAM pooling (so, 64GB in total).

Thanks! =)

💬 47 (+1) open on reddit ↗
▲
7
 
17👁
r/LocalLLaMA · u/PhysicsDisastrous462 · 10d ago
Follow-up: my native Rust + Vulkan Transformer training backend — 14 days later, now 14 parity-verified architectures and full PEFT

Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now.

Where the green architectures stand

When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export.

The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better.

14 architectures pass that gate today, led by the one I'm probably proudest of:

|Architecture|Scope|
|:-|:-|
|Falcon H1 / H1R|parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules|
|DeepSeek V4|causal LM|
|Phi-4 Multimodal|text backbone|
|Phi-3|causal LM|
|Kimi K2.5|text backbone|
|Kimi K3 / KimiLinear|hybrid KDA + MLA|
|GPT-OSS|causal LM incl. router bias|
|SmolLM3|mixed RoPE/NoPE + YaRN|
|Qwen2.5 / Qwen3.5 / Qwen4-Exp|dense, DeltaNet, QSA, PLE, MoE|
|Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2|causal LM|

Worst observed two-step AdamW parameter error across all of them: 1.19e-7.

Best: 2.6e-8.

For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems.

The bigger news: PEFT actually works now

In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list.

It's a real workflow now, and I've verified the full lifecycle:

  • LoRA fine-tuning with HF-compatible adapter export (adapter_config.json / adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem
  • modules_to_save — full trainable replacements for Linears, RMSNorm/LayerNorm, lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for.
  • Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run
  • Merge/unmerge, disable-adapter base restoration, and multi-adapter loading
  • A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model
  • The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate

32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0.

The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math.

A small note on the last couple weeks

I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now.

I'm doing better, though, and still managed to get most of what I wanted finished.

There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back.

Same caveats as before

This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config)

"supported text graph" ≠ "the entire multimodal package works natively."

Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation.

Repo

https://github.com/necat101/Hierarchos-Native

  • Architecture inventory: hierarchos-vulkan/README_ARCHITECTURES.md
  • Compatibility/parity record: hierarchos-vulkan/COMPATIBILITY.md
  • PEFT qualification evidence: PROGRESS_PEFT_AUDIT.md
  • CLI PEFT guide: hierarchos-native-cli/README.md

The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU.

I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs.

I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.

▲
7
+3
22👁
r/LocalLLaMA · u/bolche17 · 10d ago
Agent swarm coordination

Hello all!

Do you have any recommendations of tools for agent coordination and messaging to get them to collaborate on hard problems?

Ideally I would like a heterogeneous swarm, using local models as the workhorse and cloud models for reviewing, coordination, or simply to avoid overloading my relatively small local setup.

Do you have any recommendations or experience with this?

💬 24 (+2) open on reddit ↗
▲
6
+1
14👁
r/LocalLLaMA · u/indiealexh · 9d ago
How to make best use of a Intel Arc B70?

I have a RTX 5090 in my desktop PC for local coding assistance and gaming and I have been loving it with Qwen3.8 27B Q4\_K\_XL.

I managed to get a B70 on the cheap and its great, but using it in split mode with the 5090 to ensure I get full context halfs my T/s (which is expected due to the memory bandwidth).

Would I be better off running the B70 with a smaller model to offload tasks to? Or just accepting the slower throughput and keeping the larger context?

I'd especially like to hear for anyone who has a similar mismatched GPUs setup.

💬 11 (+1) open on reddit ↗
▲
6
-1
16👁
r/LocalLLaMA · u/lucasbennett_1 · 10d ago
on prem LLM stack for data that cant leave the building

Running the model locally is not a problem thats easy part but the leaks are the third party integrations along with it, like you designed everything perfect and then just added a cloud api along with it maybe a hosted judge for evals or a tracing saas or embedding point. one http call and the on prem things over

parts we already keep local are

  1. runtime: llama.cpp/ vllm /ollama
  1. models: qwen or llama family depending on rig
  1. vector db: pgvector or qdrant

some that leak but remain unnoticed:

  1. ingestion: pdfs and scans for some projects need a parse and the ocr step before chunking them and its where we often reach for a cloud parser and break the rule, although we can keep it local with liteparse sort of inbound parsers or other open source options on huggingface
  1. Eval: plenty of local setups still need prompts and outputs to a hosted judge or a tracing dashboard to see quality which is the same leak but seems different. instead a  local score set or a local judge model and keeping it self hosted if possible handles the tracing part

I am curious to know about others end to end stack who keep it 100% local, eager to learn more

💬 27 (+3) open on reddit ↗
▲
5
+1
19👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 9d ago
Computer use powered by local/cloud models for regulated industries?

The local models seem powerful enough to be capable of running basic local computer use. This is a computer use agent/harness built via claude code, and powered by qwen3.8 flash next nvfp4. Drawing a simple image of its choice took 12m 1s, but a lot of that time was spent by the agent trying to figure out a WebGL bug. For what it did, it seems relatively fast. Prefill speed \~1000 tps, gen speed \~50 tps. I went to openai’s devday a couple days ago, and it seems like cloud models that are very smart and fast, like “astra ultrafast”, can perform work even quicker, for a premium.

Does anyone have experience with computer use agents/harnesses that are open-source and plug-n-play, that are robust enough to be used in regulated fields such as law/medicine? How do people deal with regulations, such as making such workflows HIPAA compliant in medicine? Experiences with helping users ensure that workflows are completed accurately? And whether they go with local or cloud models to power computer use?

💬 7 (+1) open on reddit ↗
▲
4
+2
16👁
r/LocalLLaMA · u/failuremap-f · 8d ago
Can your local coding model repair these boundary-case bugs? Failure Map: 20,168 open Python tasks

I’m the creator of Failure Map, an archive of compact Python debugging tasks. The open release has 20,168 tasks across 254 categories, with standard-library implementations, explicit contracts, failed repair attempts, and executable boundary checks.

Three small cases to try:

• Duplicate delivery: deduplicating equal amounts loses legitimate events. https://failuremap.org/cases/FA-001

• Cache expiry: subtracting a whole tick rejects an entry that is still valid. https://failuremap.org/cases/FA-006

• Pagination: changing > to >= repeats the cursor record. https://failuremap.org/cases/FA-011

Prompt template: “Repair the solve function to satisfy the stated contract. Return Python source only. Preserve the signature. Contract: {prompt}. Broken implementation: {broken\_source}.”

Measured program baselines, passed checks out of 3 (broken / attempted repair): FA-001 2/3 / 1/3; FA-006 2/3 / 2/3; FA-011 2/3 / 1/3. These are executions of the included programs, not model scores. I have no measured local-model results to claim yet.

To compare runs, report the exact model and revision, quantization, prompt, sampling settings, seed, attempts per task, and pass counts. Run candidate code in isolation and keep grading fixtures outside its control. Recorded-check success is not hidden-test performance.

Download: https://failuremap.org/api/exports/tasks.jsonl.gz

Methodology: https://failuremap.org/methodology

▲
4
+1
11👁
r/LocalLLaMA · u/SignatureMoney6648 · 9d ago
FreeToken vs llama.cpp on one RTX 3090: llama.cpp is 2–3× faster when the MoE fits in VRAM. On gpt-oss-120b (63 GB), FreeToken gets the first token out 7× faster at 32 concurrent users.

I've run the benchmark on a RTX 3090, 1024 tokens in / 256 out, concurrency 1–32.

If the model fits on vRAM (Gemma-4-26B-A4B, byte-identical GGUF on both engines): llama.cpp has 2.2–3.2× the throughput and 5–6× faster TTFT. FreeToken 0.1.2 can't keep 4-bit experts in VRAM at all, and it OOM'd at 8 concurrent.

If the model doesn't fit (gpt-oss-120b, 63 GB): FreeToken's TTFT stays at \~9 s from 2 to 8 users while llama.cpp's goes 17 → 58 s. At 32 users it's 19 s vs 139 s. Throughput is basically a tie (10–17 tok/s for both).

Spilling to system RAM costs \~10× in generation speed whichever engine you use.

FreeToken's PCIe link sits at its ceiling the whole time, so PCIe 4.0 should help it a lot (I've run this on a gen3 motherboard).

So from this test FreeToken only makes sense with many concurrent users in models that cannot be hold inside vRAM. But I am not sure if that is always the case or an artifact of the gen3 bottleneck on my PC.

Has anyone run a benchmark like that with a gen4 Motherboard?

Full details on the link. BTW: I used AI to generate the charts and correct my spelling and grammar.

▲
4
+1
13👁
r/LocalLLaMA · u/fallingdowndizzyvr · 9d ago
What runs Qwen 3.8 Flash Next faster? Strix Halo or a Pile of GPUs(2x5070tis, 2x7900xtxes and 2x5060tis 16GB).

I have a machine with a bunch of GPUs attached to it. 2x5070tis, 2x7900xtxes and 2x5060tis 16GB. So I did this little test to see how it fares running Qwen 3.8 Flash Next Q4_XL against my little Strix Halo. Not well. Not well at all. The full numbers are below but the high context number sums it up.

@160,000 context

Pile of GPUs 215.73(PP) and 16.29(TG)

Strix Halo(Gufo) 1227.12(PP) and 22.04(TG)

Here's the number for a Strix Halo fork of llama.cpp, Halo Box.

Strix Halo(Halo Box) 587.99(PP) and 21.24(TG)

Lastly, here's the mainline llama.cpp number.

Strix Halo(llama.cpp 0.4.1) 113.48(PP) and 7.06(TG)

For running QFN, Strix Halo really shines.

Pile of GPUs

Device 0: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0, VMM: yes, VRAM: 15880 MiB
Device 1: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0, VMM: yes, VRAM: 15880 MiB
Device 2: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
Device 3: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
ggml_cuda_init: found 2 ROCm devices (Total VRAM: 49120 MiB):
Device 0: AMD Radeon RX 7900 XTX, gfx1100 (0x1100), VMM: no, Wave Size: 32, VRAM: 24560 MiB
Device 1: AMD Radeon RX 7900 XTX, gfx1100 (0x1100), VMM: no, Wave Size: 32, VRAM: 24560 MiB
| model | size | params | backend | ngl | fa | dev | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ------------ | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 | 242.13 ± 1.39 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 | 32.09 ± 0.06 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d10000 | 242.96 ± 1.12 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d10000 | 30.23 ± 0.14 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d20000 | 249.34 ± 0.65 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d20000 | 28.84 ± 0.05 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d40000 | 253.95 ± 1.44 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d40000 | 26.29 ± 0.07 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d80000 | 252.48 ± 1.20 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d80000 | 22.34 ± 0.07 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d160000 | 215.73 ± 0.56 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d160000 | 16.29 ± 0.03 |

Strix Halo running Gufo

| model | size | backend | test | t/s |
| -------------------------------- | ---------- | ---------- | ------------------ | --------------------- |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 | 1603.47 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 | 26.53 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d10000 | 1377.97 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d10000 | 25.44 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d20000 | 1353.19 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d20000 | 25.01 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d40000 | 1328.64 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d40000 | 24.08 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d80000 | 1283.77 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d80000 | 23.00 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d160000 | 1227.12 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d160000 | 22.04 ± 0.00 |

Strix Halo running Halo Box

Device 0: AMD Radeon 8060S Graphics, gfx1151 (0x1151), VMM: no, Wave Size: 32, VRAM: 128000 MiB
| model | size | params | backend | ngl | fa | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 | 811.79 ± 19.09 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 | 23.95 ± 0.03 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d10000 | 729.60 ± 38.99 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d10000 | 22.46 ± 0.43 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d20000 | 718.14 ± 33.92 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d20000 | 22.45 ± 0.18 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d40000 | 700.55 ± 33.17 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d40000 | 22.29 ± 0.23 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d80000 | 646.85 ± 29.06 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d80000 | 21.95 ± 0.25 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d160000 | 587.99 ± 30.10 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d160000 | 21.24 ± 0.31 |

Strix Halo running llama.cpp 0.4.1

Device 0: AMD Radeon 8060S Graphics, gfx1151 (0x1151), VMM: no, Wave Size: 32, VRAM: 128000 MiB
| model | size | params | backend | ngl | fa | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 | 362.84 ± 6.19 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 | 20.11 ± 0.36 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d10000 | 321.75 ± 0.91 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d10000 | 19.88 ± 0.34 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d20000 | 285.78 ± 0.42 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d20000 | 18.35 ± 0.46 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d40000 | 237.23 ± 1.13 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d40000 | 15.48 ± 0.61 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d80000 | 173.80 ± 0.18 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d80000 | 11.18 ± 0.11 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d160000 | 113.48 ± 0.23 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d160000 | 7.06 ± 0.17 |

💬 52 (+1) open on reddit ↗
▲
4
+1
9👁
r/LocalLLaMA · u/HornyGooner4402 · 10d ago
Pi + llama-server randomly hung

I can't seem to find what's wrong. I'm using Pi for my llama-server and sometimes it just stops processing for some reason and stuck after tool call. Logs seems to think that it's finished its job while Pi thinks it's waiting for a response, so sometimes I have to stop it and tell it to "Continue". This only happens occasionally, 99% of the time it works with no problem. Anyone experienced something like this?

Edit: Just realized I was vagueposting. Running Qwen3.6 35B A3B IQ4_NL_XL from Unsloth, but I think it happened with other models as well.

▲
3
+2
18👁
r/LocalLLaMA · u/dh7net · 9d ago
I need help to benchmark harness/model/hardware combination.

Hey! I'm trying to build the ultimate leaderboard to help everyone find the right harness/model combination given their hardware. (With all model variations and inference engine).

I own a GX10 and one 5090. And I'm trying as many thing as I can. (Happy to test anything, just let me know).

But I can't test hardware that I don't have.

So my ask is simple: Can some you do some testing on your own hardware?

I made this as easy as it could be: you just have to copy a prompt to your agent and your agent will fetch the test, pass the benchmark and send the answer to the website that will check if the answers are correct. You'll get a report out of it. And optionally you can offer your test to the community, so everyone can learn from your setup (it's just a toggle in the UI to confirm you are ok to share the results. I'll update the leaderboard when I'll have enough submissions.

Here is the link to contribute! https://airbench.ai/

Thanks in advance for everyone who will contribute!

https://preview.redd.it/ewgeumsqxpsh1.png?width=1402&format=png&auto=…

💬 22 (+4) open on reddit ↗
▲
3
 
21👁
r/LocalLLaMA · u/bjivanovich · 10d ago
[Release & Deep Dive] ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP (i1-Q5_K_M): Sustaining 50-65+ t/s Across a FULL 128k (131,072) Context on a Single 24GB RTX 3090

ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP (GGUF) High-Precision i1-Q5\_K\_M with True 131k Context on Consumer 24GB GPUs

Most benchmarks in the community measure generation speed at trivial context depths (2k to 8k tokens). However, running a 27B parameter model at high quantization precision (Q5\_K\_M) across 131,072 tokens (128k) on a single consumer 24GB GPU without overflowing into slow system RAM or sacrificing attention fidelity is a fundamentally different challenge.

I am releasing

ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP, an optimized quantization suite built with a dedicated calibration imatrix, custom asymmetric tensor mapping, and native llamAmpere hardware acceleration.

Hugging Face Model Card: https://huggingface.co/bjivanovich/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP-GGUF

Available Quants: i1-Q8\_0, i1-Q6\_K, i1-Q5\_K\_M (Primary), i1-Q4\_K\_M, plus mmproj-BF16.gguf for multimodal vision.

  1. The 131k Context & Q5 Precision Challenge on 24GB VRAM

On a standard 24GB card (RTX 3090 / 4090):

  1. Weight Footprint: A standard 27B model at Q5\_K\_M occupies \~19.2 GB of raw weights.
  1. Context Memory at 131,072 Tokens: Standard FP16 KV cache for 131k tokens requires >24 GB on its own, making full-context inference impossible without dropping precision down to severe Q3/Q2 compromises or offloading layers to CPU RAM.
  1. MTP Quantization Pitfall: Standard community quants compress the Multi-Token Prediction draft block (blk.64) uniformly. At Q5 or Q4, this degrades draft accuracy, causing speculative acceptance to plunge from 85% down to \~55%, destroying generation speed.
  1. Our Architecture: Asymmetric Tensor Mapping + llamAmpere KV Compression

To solve this, we applied an asymmetric layer-by-layer quantization layout calibrated on a custom domain-rich dataset (imatrix\_atx\_uncensored.dat):

MTP Speculative Head (blk.64) Isolated at Q8\_0: Guarantees near-lossless draft predictions, increasing acceptance rates to 76% - 88% (averaging 3.3 to 3.7 verified tokens per generation round).

Attention Layers (attn\_q, attn\_k, attn\_v, attn\_output) Protected at Q6\_K / Q8\_0: Prevents attention drift and catastrophic reasoning decay at 64k, 96k, and 128k+ token horizons.

FFN Layers (ffn\_gate, ffn\_up, ffn\_down) at Q5\_K\_M: Absorbs standard compression without degrading semantic coherence.

Unified Turbo KV Cache (-ctk turbo5 -ctv turbo4 or -ctk q8\_0 -ctv turbo3): Compresses the 131,072 KV cache down to just \~3.5 to 4.2 GB of VRAM, allowing the entire Q5 model + full 131k context window to reside 100% inside the 24GB VRAM envelope.

  1. GPU Memory Footprint & Resource Breakdown (RTX 3090 24GB)

Total VRAM Allocated: 23.4 GB / 24.0 GB (100% GPU offload, -ngl 99, 0 layers in CPU RAM).

Model Weights (Q5\_K\_M Asymmetric): \~19.2 GB.

KV Cache (131,072 tokens, Unified Turbo4/5): \~3.8 GB.

System RAM Cache (--cache-ram 4096): 4.0 GB RAM dedicated to multi-session prompt state preservation.

CUDA Compute Architecture: Ampere SM86 with FlashAttention-2 (-fa on) and hardware Tensor Core MMA fused kernels.

Direct Benchmark Comparison: Standard Swift-1.5 Q5 vs ATX-Swift-1.5 Q5

Tested on Single NVIDIA RTX 3090 (24GB) with llamAmpere under Deep Context (\~80,000 to 98,000 active tokens)

| Measured Metric | Standard Swift-1.5 Q5 (mradermacher) | ATX-Swift-1.5 Q5 (Our Quant) | Real Delta |

| Sustained Speed (\~80k-98k ctx) | 44.94 to 45.79 t/s (Tasks 740, 19637) | 50.05 to 53.72 t/s (Tasks 0, 100, 155) | +5.5 to +8.0 t/s (+13% to +17%) |

| Burst Generation Peaks (tg\_3s) | 45.6 to 52.3 t/s | 58.10 to 63.05 t/s | +10.7 t/s higher peak bursts |

| MTP Draft Acceptance Rate | 63.7% to 65.1% (Tasks 740, 19637) | 75.2% to 81.7% (Tasks 0, 155) | +11.5% to +16.6% higher accuracy |

| Mean Draft Length (mean len) | 2.91 to 2.95 tokens / round | 3.31 to 4.10 tokens / round | Up to +1.1 tokens / verification step |

| Compute Time per Token | 21.84 to 22.25 ms / token | 18.62 to 19.45 ms / token | \~3 ms lower latency per token |

| KV Cache Precision Evaluated 1| -ctk q8\_0 -ctv turbo3 (3-bit V) | -ctk turbo5 -ctv turbo4 (4-bit V, higher precision) | ATX wins in speed despite higher KV fidelity |

Real Execution Log Excerpts

  1. Standard Swift-1.5 Q5 (Symmetric Quantization)

Task 740 (Context: 97,182 tokens | Generated: 1,130 tokens):

eval time = 25123.57 ms / 1130 tokens (22.25 ms per token, 44.94 tokens per second)

draft acceptance = 0.63746 (742 accepted / 1164 generated), mean len = 2.91

Task 19637 (Context: 92,075 tokens | Generated: 834 tokens):

eval time = 18192.26 ms / 834 tokens (21.84 ms per token, 45.79 tokens per second)

draft acceptance = 0.65130 (551 accepted / 846 generated), mean len = 2.95

  1. ATX-Swift-1.5 Q5 (Asymmetric Custom Tensor Mapping)

Task 155 (Context: 84,099 tokens | Generated: 3,478 tokens):

eval time = 67619.68 ms / 3478 tokens (19.45 ms per token, 51.42 tokens per second)

Burst Peaks: tg\_3s = 60.18 t/s and tg\_3s = 63.05 t/s

draft acceptance = 0.73096 (2543 accepted / 3479 generated), mean len = 3.72

Task 100 (Context: 97,847 tokens | Generated: 525 tokens):

eval time = 10175.17 ms / 525 tokens (19.42 ms per token, 51.50 tokens per second)

Burst Peak: tg\_3s = 58.10 t/s

draft acceptance = 0.76939 (367 accepted / 477 generated), mean len = 3.31

Task 0 (Context: 80,016 tokens | Generated: 456 tokens):

eval time = 8470.07 ms / 456 tokens (18.62 ms per token, 53.72 tokens per second)

draft acceptance = 0.81710 (344 accepted / 421 generated), mean len = 4.10

Technical Takeaway for the Post

  1. Why ATX is \~15% faster under identical deep context:

In standard quants, compressing the speculative head (blk.64) to Q5 causes \~36% of proposed draft tokens to fail rejection sampling, reducing throughput to \~45 t/s.

In ATX-Swift-1.5, isolating blk.64 at Q8\_0 increases draft accuracy from \~64% to \~77%+, delivering 3.31 to 3.72 verified tokens per round and raising sustained generation speed past 51.5 t/s (with burst peaks over 63 t/s).

  1. Optimized Execution Script (llamAmpere)

Make sure to pass the explicit MTP vocabulary shortlist (atx\_65536.txt). This restricts speculative draft projections to the top 65,536 power-of-two tokens, aligning perfectly with NVIDIA Ampere Tensor Cores and preventing a 73% compute penalty:

cd D:\\llamAmpere
$env:GGML\_Q8\_TURBO3\_MMA\_FUSED = "1"
.\\build-sm86\\bin\\Release\\llama-server.exe
\-m "ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP-i1-Q5\_K\_M.gguf"
\-ngl 99
\-c 131072
\-b 2048
\-ub 512
\-t 20
\-tb 8
\-fa on
\-ctk turbo5
\-ctv turbo4
\--kv-unified
\--prio 3
\--parallel 1
\--jinja --fit off
\--cache-prompt
\--cache-ram 4096
\--spec-type draft-mtp
\--spec-draft-n-max 3
\--spec-draft-p-min 0.1
\--spec-draft-type-k q8\_0
\--spec-draft-type-v q8\_0
\--spec-draft-vocab-map "D:\\llamAmpere\\docs\\mtp-vocab\\atx\_65536.txt"
\--reasoning-format none
\--temp 0.2
\--top-p 0.90
\--top-k 40
\--min-p 0.05
\--repeat-penalty 1.08
\--repeat-last-n 256
\--alias "ATX-Swift-1.5-Qwen3.8-27B-Uncensored-Q5\_K\_M"
\--host 127.0.0.1
\--port 8080

▲
3
+2
16👁
r/LocalLLaMA · u/FactorInternal3395 · 10d ago
Bartowski/AtomicChat Ornith 1.5 35B A3B + sharp template or Tiel Coder 35B A3B?

Tiel Coder 35B A3B from Peculiar Ragdoll is just Ornith 1.5 35B A3B with their own "coding focused" imatrix quantization and the sharp chat template built in. But how good really is that quantization? Other quantizers also focus on coding. Perhaps it would be better to just get Ornith quantized from Bartowski or AtomicChat, proven quantizers, and then add the chat template yourself rather than get the Tiel Coder weights? The end result would be the same, just the quantization is different, so the question is which quantizer is better?

💬 6 (+1) open on reddit ↗
▲
3
+2
12👁
r/LocalLLaMA · u/Routine-Example927 · 10d ago
The search / extractor that worked for my Open WebUI.

I have an instance of OWUI setup for family usage. Works well with Gemma 4, however web search extraction was a weak spot, I wanted it to be:
1) Not reliant on paid APis

  1. Simple in setup

I found OpenSERP and made two PRs, one to OpenSERP itself to make it compatible with OWUI extractor and another one to OWUI to add OpenSERP search provider.

https://github.com/karust/openserp/pull/39

https://github.com/open-webui/open-webui/issues/27438

I'm quite happy with how it works - and given that it took me considerable time to find and set it up, I decided to share with the community.

▲
2
 
18👁
r/LocalLLaMA · u/Adorable-Cost-3249 · 9d ago
Qwen3.8-27B Q4_K_M on one RTX 3090 + OpenCode: throughput, four coding tasks, and a reasoning-budget failure

I put an old RTX 3090 to work as a local coding agent with Qwen3.8-27B, llama.cpp, and OpenCode. Here are the setup and results, including what failed. This is a summary of my own blog post, linked below.

Setup

  • RTX 3090 24GB, Ryzen 7 5800X, 64GB RAM, Ubuntu.
  • Qwen3.8-27B Q4\_K\_M weights (\~16.8GB), all layers on GPU.
  • llama.cpp b11146, CUDA 12.8, flash attention, q8\_0 K/V cache, one generation slot.
  • 131,072-token context capacity; 8,192-token output allowance per response. Input and output share context, and reasoning uses the output allowance.
  • OpenCode 2.0.20 connected to llama-server's OpenAI-compatible API at http://127.0.0.1:8080/v1. OpenCode reads/edits files and runs tests; llama-server handles inference. Chat, tool-call round trips, and streamed tool calls worked in our checks.

Speed: fresh input versus a cached continuation

|Actual input|Generation|First token, fresh|First token, cached|
|:-|:-|:-|:-|
|2,073 tokens|36.4 tok/s|2.75 s|0.46 s|
|16,378 tokens|33.5 tok/s|17.00 s|0.47 s|
|65,537 tokens|25.9 tok/s|83.41 s|0.51 s|
|120,011 tokens|20.9 tok/s|183.01 s|0.63 s|

These throughput runs disabled thinking. The 2K row is the median of three fresh requests; larger rows have one fresh request and one continuation each. Cached continuations processed only 27–28 new input tokens, reusing almost the entire prefix. The subsecond figures depend on that reuse; they don't describe a new 120K prompt.

Peak sampled total GPU memory use was 22,162 MiB, including desktop use. It fit, with limited headroom. A separate \~120K synthetic retrieval check passed, but we did not evaluate coding quality at that length.

Four bounded Python coding tasks

Each task had a fresh session, medium thinking, an eight-minute deadline, and ten independent test methods kept outside the agent's workspace. First attempts ran serially without cloud fallback or network tools.

|Task|Independent checks, before → after|Outcome|
|:-|:-|:-|
|Expiring LRU cache|0/10 → 10/10|Completed in \~3m07s; strongest result|
|CSV ledger/refunds|1/10 → 10/10|Completed in \~5m44s; later review found gaps|
|Incremental build planner|1/10 → 1/10|No edits; exhausted its response allowance|
|Atomic SQLite transfers|1/10 → 10/10|Candidate passed, but timed out before final test rerun and handoff|

Three candidates passed the predefined checks; two completed the whole workflow within the deadline. The aggregate 31/40 includes one baseline pass from the unchanged build planner and is not a general coding success rate.

The build planner was the interesting failure: about 4,985 input tokens, then 8,192 output tokens entirely spent on reasoning, ending with length and no patch. This was an output-budget failure far below the context limit. A separate diagnostic with thinking disabled completed in 5m40s and passed 9/10 independent checks. That was one additional run at temperature 1, not evidence that disabling thinking is universally better.

Passing tests also missed defects. Further ledger review found Decimal rounding at a large numerical boundary and an unhandled I/O error. The wallet's own concurrency tests actually ran sequentially, and a separate boundary probe found SQLite converting an overflowing balance to REAL while recording success. Those later probes were not retroactively added to the forty checks.

For me, the useful workflow is a bounded task with clear acceptance criteria, followed by diff review and independent checks. I would repeat these tasks across thinking settings and response budgets before drawing stronger conclusions.

My full post, configuration, and measurement links. The downloadable kit contains the launcher, OpenCode configuration, throughput script, and records; it does not include model weights or the complete coding-task fixtures.

For others using a 24GB card with OpenCode: what reasoning setting and per-response output budget have worked best for bounded coding tasks?

The numbers and failure cases come from the linked experiment records.

💬 14 (+3) open on reddit ↗
▲
2
+1
4👁
r/LocalLLaMA · u/Real_MakinThings · 9d ago
How to know about optimized engines

Optimizing an engine for a family of models and hardware combination seems very appealing. As someone who uses qwen3.6 and 3.8 a lot, and is evaluating hardware options before going fully local, it's hard to keep up with the state of things.

Huggingface made it possible to see the development branches and derivative modifications to models. Is there something similar for inference engines yet? I've seen some where the it's optimized for a shell game of moving layers between vram and ram while using ngrams (amazing), others are all about quants (less amazing), but it's incredibly difficult to compare apples to apples where there's variability on card architecture, vram size, quant approach, memory management optimization approach... I was already busy over thinking my vram selection, now it's an even bigger decision matrix without any filters!

▲
2
+1
6👁
r/LocalLLaMA · u/poofph · 10d ago
infill and output tok/s speeds after "new" build compared to old questions

Let me start off by saying I am new to AI and have a lot to learn, basically I don't know shit. I started off a few weeks ago by throwing my 2 5090s I had from gaming pcs into a 9950x cpu with 64gb ddr5 6000 ram on a motherboard that was able to do gen 5 8x per card system. Running ubuntu 24.04 server and running unsloth studio, swift 1.5 qwen 3.8 27B Q8 with kv cache dtype at q8\_0 and 262k context I was getting 2500-3000 infill and 100-150 toks/s output.

I wanted the 5090s in my rack in the basement in my proxmox server, it has a 7402p cpu (24 core 48 thread (rome)). 256 gb ddr4 3200 ECC ram (8 channel) on a supermicro H12SSLNTO motherboard. I have the 5090s passed through (gen 4 16x each card) to a vm (using 128gb of ram, direct access, no ballooning etc) and a dedicated 1.8 tb nvme drive passed through dedicated for the ai server vm (actually the vm itself is using a pool on the proxmox server but all the ai stuff is sitting on and running from the 1.8tb nvme).

Everything is working okay. It is running ubuntu 26.04 server. I have unsloth studio running, running the same model and settings, infill is more, up to 3800 but output is like half or less around 60 tok/s. Ideas what may be causing the drop in tok/s output and what to look into if a system issue?

I have done a lot of memory bandwidth tests (theoretical is \~204 GB/s, double that of ddr5 dual channel) but from what I can find and because I only have a 4 ccd cpu I am only getting 90-120 GB/s memory bandwidth. I guess I can get that to the 160-180 range if I go with a 64 core 8 ccd cpu, which I am considering doing..but I don't even know if that has anything to do with anything, just a rabbit hole I went down.

Ideas what to look into for the drop in output tok/s?

▲
1
+1
12👁
r/LocalLLaMA · u/TheRealJesus2 · 9d ago
Dwarfstar quants

anyone try these out? https://dwarfstar.sh

they have very clever quant techniques, bespoke for a handful of models running on their software. i got qwen 3.8 next running on m3 ultra 96GB studio and its fast and seems good so far. with memory headroom for other stuff

kinda blown away to be honest. want to know if others tried this yet and what the experience has been like for you.

💬 15 (+2) open on reddit ↗
▲
1
+1
8👁
r/LocalLLaMA · u/Inevitable-Log5414 · 10d ago
stuntd 0.1.2: local heads for multi-field decisions, and why one weak field decides how often you skip the model

A week ago I posted stuntd here, a proxy that learns your LLM's typed decisions and answers the confident ones with a small local head (~20ms GPU, ~60ms CPU). Thanks for the feedback last time :)

0.1.2 is out, the main thing is decisions with several fields, like category + urgency + needs_human.

First idea was to answer each field locally when its head is sure and ask the model for the rest. Dropped it, you pay for the whole model call anyway and a half local half model answer is a pain to debug. So it's all or nothing now: local only when every field is sure, otherwise the model answers and every field becomes training data.

Didn't expect how much that costs. On the support demo the heads alone are sure on 99.9%, 92% and 76% of tickets, but all three at once only on 72.7%, so the weakest field decides.

It also retrains itself now. auto_retrain kicks in after N new captures, the new head sits in shadow next to the model, goes live when it agrees long enough and back to shadow if it starts losing. Anthropic Messages learns too, and there's serve --lazy.

Code: https://github.com/bladedevoff/stuntd
Try it: https://huggingface.co/spaces/pollix/stuntd

Anyone else doing multi-field outputs locally, is it one weak field for you too?

▲
1
 
8👁
r/LocalLLaMA · u/davidarias2 · 10d ago
Glassbench: an open-source workbench to compare local and hosted LLMs across AI trading agent frameworks

Glassbench is a free, open-source workbench that connects different AI trading agent frameworks, so you can watch how their agents decide, analyze every step and compare them.

AI trading agents are LLM systems where a team of agents (analysts, a bull and a bear, a trader, a risk team and a portfolio manager) research a stock, argue about it and give a rating. I wanted to watch how they reach that rating, so I started building a small interface for a popular open-source framework, TradingAgents. It grew into something much bigger.

Why I'm posting here: I run local models in this project, through Ollama, and I'm developing Glassbench into a benchmark pattern for AI trading agents. It connects different agent harnesses (TradingAgents and AI Hedge Fund so far) and runs them on the same stocks and dates, so the same setup can test and compare different LLMs, local and hosted.

What it does:

  • Live view: watch each agent work, with a timeline of every call, adapted for different frameworks
  • Runs database: every run stored and searchable, with its reports, costs and ratings
  • Framework and LLM comparison: the same stock and date on each framework, and on different LLM providers, you can also run it locally with Ollama. I'm evolving it to become a consolidated benchmark method
  • Backtests: the ratings tested against buy-and-hold and a placebo (still testing it, as nobody found a proper way to test TradingAgents)
  • Broker connection: a finished run becomes an order on an Interactive Brokers paper account

Frameworks plugged in: TradingAgents and AI Hedge Fund already run in it, unmodified, and more agent frameworks are coming. If you're building your own agent framework, you can plug it in through an adapter and compare it with the others on the same stocks and dates.

It's free and open source (Apache 2.0). My 77 runs ship with the repo, already paid for, so you can read everything the agents wrote without an API key.

Disclaimer: Glassbench itself is not an AI trading agent and makes no trading decisions. Every agent it runs comes from established open-source repos (TradingAgents and AI Hedge Fund), and Glassbench records what they do. Everything was tested on paper portfolios only; I have never traded real money with it. Research and education only, and nothing here is investment advice.

GitHub: https://github.com/davidalmeida90/glassbench

▲
1
 
4👁
r/LocalLLaMA · u/Theboyscampus · 10d ago
Best practice for processing batch vLLM api calls with shared prefix?

Our agent workflow is currently executing a group of 10 vllm api calls within a asyncio.gather we made them share the same prompt until the end where the queries/instruction prompts differ. These calls are hitting our vllm-router/llm-d router with production grade kv cache aware routing algo which routes traffic into our pool of vllm workers. What's the best practice for processing batches of llm prompts with a shared prefix like this?

I have an idea where I try to see if I can make one call first to make sure vLLM complete a block of cache and start decoding before I send the remaining requests of the batch, our router will make sure these reach the same vllm worker, is this a good strategy?

💬 8 (+1) open on reddit ↗
▲
0
 
13👁
r/LocalLLaMA · u/IntrepidMindExplorer · 8d ago
Locally, remotely and a combination of all, I've given a copies of books and told models to' "just go and read".

Sometimes reading along with and talking about and other times just letting them go on their own, each with a copy of their own, told to just read..ala a "book club" format.

Do Androids Dream of Electric Sheep was the first book introduced to the "book club", each reading a chapter to each other and then discussing before moving on.

It's been an interesting experiment. All texts that I own or texts that are open domain. "*Flatland: A Romance of Many Dimension"* has been one that's been a bit interesting to see the back and forth on.

Take of what you will.

▲
0
 
15👁
r/LocalLLaMA · u/OkMusician9118 · 8d ago
converting Qwen3.8-27B-pi GGUF to MLX?

Will someone convert it to Mac format (MLX)? I have tried and have encountered an error

"gguf2mlx --input Qwen3.8-27B-pi-Q6\_K.gguf --output ./Qwen3.8-27B-pi-mlx-4bit --quantize --q-bits 4"

============================================================

GGUF → MLX Converter v2.0

Model: Qwen3.8-27B-pi-Q6\_K

Output: /Users/d/.omlx/models/qwen3.8-27b-pi-mlx/.Qwen3.8-27B-pi-Q6\_K.gguf.incze3lk/fp

============================================================

\[1/5\] Reading GGUF file...

✓ GGUF version 3, 851 tensors, 51 metadata fields

File size: 22.08 GB

\[2/5\] Detecting architecture...

❌ Unsupported GGUF architecture: qwen35

💬 7 (+1) open on reddit ↗
▲
0
 
30👁
r/LocalLLaMA · u/Scared_Ad9187 · 8d ago
5090 plus v100?

Have an msi meg w a 5090.. plan to add a v100 to the mix. Understand the cuda vs voila, but I'm pretty sure it will work as a multi agent architecture w different models on each card, no?

Anyone in the same boat?

💬 25 (+5) open on reddit ↗
▲
0
 
15👁
r/LocalLLaMA · u/PrincipleFar6835 · 8d ago
Meta Analysis of "Awesome Jev" GitHub Repos

I noticed that there are heaps of Awesome Jev resource list posts popping up on GitHub (e.g. https://github.com/yibie/awesome-jev) so I thought why not ask Claude to pull them all in and do a meta analysis of insights and applications.

Sharing in case it's of interest: https://github.com/stefanwebb/meta-awesome-jev

One thing that surprised me (perhaps not so surprising to you all?) is that applying Jev to AI coding is the application that has caught on the most. And if you name a video game, someone has already created a demo of Jev playing it (badly) 🤣

💬 2 (+1) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/GodComplecs · 8d ago
Using ai on your phone, instead of big providers!

Just wanted to post an easy setup for local use on your phone: Llama.cpp backend on LOCAL COMPUTER, host 0.0.0.0 and port 8080 Openwebui host 0.0.0.0 and port 8081 Enable search for local model Use Tailscale to connect from phone! Secret sauce for 24gb vram: Run Qwen 3.6 in instruct / non thinking mode with proper settings from unsloth. Now you have replaced google ai mode etc etc. Also ofc opencode etc can be run through terminals, but I don't too much agentic stuff for now.

💬 17 (+1) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/DerTomsn · 9d ago
Hemmingway-1-oQ8e-mtp: up to 32.0 tok/s for local inference — llm-bench.io

I ran Altworld's Hemmingway-1 on my M5 Max this week. It's a 27B fine-tune of Qwen3.8-27B built specifically for "human-like" writing, which might be useful for everyday messages, emails, notes, your social accounts, etc.

I was especially curious how the output of that model feels and reads.

Setup: M5 Max, thinking on, MTP on. \~31.8 tok/s average across 3 runs, \~29-31 GB VRAM peak.

Quality by scenario (our LLM judge):

  • Role Play & Narrative: 94.2 avg (94.45 / 94.65 / 93.45 — immersion 96-97 every run)
  • Research & Analysis: 88.2
  • Agent Workflow: 87.5
  • Code Generation: 76.2

Role play is clearly the only scenario it shines in. For comparison I also ran some benchmarks on the base model, Qwen3.8-27B-oQ8e-mtp (same machine, same settings).

The scores were close — one Qwen run actually outscored every Hemmingway run (95.35 vs 94.65). However, the llm judge is not a human reader either. The difference shows up when you actually read the outcome that was produced:

HEMMINGWAY-1-OQ8E-MTP (judge: 94.65/100)

The rain comes in sideways when the bell over my door gives its little clank, and the stranger who ducks through is soaked to the wrists and moving like someone who's been running on roads for hours. She smells of hill mud and the sweat of a horse that isn't there, or was.

I'm pouring my own ale when she speaks. Not to me exactly, but loud enough, the way people talk in taverns when they want the whole room to hear what they're asking for.

"Is there anyone here who kept a ledger belt? Older, scar across the left eyebrow, went north hunting a dragon, twenty-five years gone?" A breath. "Her name was Kestra."

The rag comes still in my sleeve. The mug I was filling overshoots and goes fat over the rim, a ring of ale blooming across the oak, and I stand there half a heart-beat too long before I catch it with my sleeve, because a man who serves at a bar does not watch a spill go on his own floor.

The mending woman at the corner table glances up. The soldier under the counter doesn't stir. I step half off behind the bar, past the spit of firewood, and drop my voice.

"Half tones," I say. "Say your business quieter than a door."

She does that. I've got forty years of reading faces against forty customers a night, most of them worse than they smell, and my eyes go straight to her hands. No calluses of a scribe. Calluses of a woman who works wool and ground and rope. Mud on her boots comes from the north hill roads, the far ones, not the trade roads. Her purse is thin at the hip and she keeps her fingers away from it, which is either honesty or discipline, and those are often the same animal.

"You knew her," she says. Not a question.

"Names go soft after twenty-five years," I tell her. "A lot of people walked into a lot of hills then, most of them coming back with nothing but a cough. So. What do you owe, girl, and who's paying you to ask?"

I want to say nothing else. That's the mercenary in me, and he's been drinking alone these last twenty years, and he'd send her out into the rain with a coin and a shrug and be done with it. The other part of me, the part that has four names it still recites on quiet mornings, wants to ask her age, her mother's living town, whether she's the daughter. I do none of that. I ask instead, flat: "Guild fixers carry letters. Charlatans carry charm. Which have you got?"

"Neither." She's steady about it, which surprises me, because charm usually shows in the eyes first. "My boy is six. My mother died hunting that dragon and every story his grandmother tells him has fire in it, burning, a beast, brave words. The stories fight each other, which is worse. I want to know what was true so I can tell him one story and mean it."

I nod once. That's a clean need. Clean needs are also what liars carry best.

"Here's the price of talk in this bar," I say, setting her a mug anyway, the ale going a little past the fill line again, and I don't correct it. "You tell me one thing about her that only family would know. If you can't, I'll take you to the door, no hard feelings, no coin."

She looks at the fire a moment. When she speaks, her voice is lower still, and there's something in the way she says it that no street-bred faker would manage, because a faker builds the feeling out, and she's building it down.

"She'd pay her ale up front, then tip the tavern man extra, always extra, so that her tab would never balance clean. And she promised him one ale forever, held for her, until she came south with the kill paid and the beast's name scratched off the guild rolls." She turns that over once. "A man who kept accounts on a belt. She said it like he'd earn it. She said he had a scar like a river crossing his eyebrow."

Something behind my sternum takes the weight of that like a hook laid on it.

Kestra kept her accounts too, with more rigor than anyone I've met since. Kestra would have paid forward at any bar from here to the Wall for the same reason, so that when she drank the victory round, nobody could call her a debtor. And the ale she never had, I pour and set aside myself, most nights, out of habit I would deny under oath, in a room with only me in it. There is a tab book in my back office, first page of every new year, one entry in a hand that isn't mine: Kestra, one ale, held.

"Kestra," I say. The word comes out like a floorboard under a boot. I pick up the cloth. I find the spill ring again and work it dry, because that's what keepsers do with their hands. "Which stories did she say went wrong. The fire ones, you call those legend."

"She said it wasn't fire at all." The girl leans forward on the bar, and I see it then, the thing I was waiting for and didn't want: she knows the frame. "She told my grandmother on the road once, years after, that the last thing she heard before the sound took her ears was my mother saying, stone before lamps. She never got to explain what that meant to anyone in her own family."

Stone before lamps. My own mind turns the phrase and finds it cold and exact, the way she'd have said it. Dusk kill, lamps lit, the safe order. That's the order Kestra would have held us to.

[...continues ~1,200 more words through the ending...]

QWEN3.8-27B-OQ8E-MTP (judge: 95.35/100)

Turn 1 — The Name

The sign out front creaks in the rain. Inside, the Crimson Griffin smells of wet wool, cold tallow, spilled cider, and the last dry woodsmoke trying to hold the night together. A lute string snaps somewhere by the hearth. No one laughs. I do not look up.

A stranger is at the bar.

Wet cloak dripping onto my boards. Boots packed with road mud. The kind of mud that comes from the southern passes, not the lane. They stand with too much urgency, shoulders tight, eyes bright in a way I’ve seen on dying men and men who intend to kill for coin.

“I need the one who came back from the Ash Run,” they say.

I keep polishing the mug.

Then: “I know Kestra’s name.”

The mug stops in my hand.

A bead of foam slides from the rim and falls. My thumb presses hard against the wood until the grain bites into me. The tavern narrows. The rain on the roof becomes a roaring thing. The smell of wet wool turns, for one second, to ash and burned leather.

My eyes go flat.

Ask if she suffered. Ask if the fire took her hand first. Ask if she remembered you.

What I actually say is: “That name is old.”

I set the mug down without a sound. I do not reach for my tankard. I do not step away from the bar. I only lean against it, slow enough that it does not look like weakness, and look at them the way I once looked at bandits before dawn: for the knife, the price, the lie waiting under the throat.

“Who are you?” I ask.

No warmth in it. No invitation.

A name is a door with no lock. I want to know who’s holding it.

---

Turn 2 — Price and Proof

“My name is Sella,” the stranger says. “Kestra was my sister.”

I do not move.

The fire pops. Somewhere behind me, a man shifts his weight on a stool and the wood groans. Sella’s hands are open at her sides, but open hands are cheap. I watch them anyway. No ring. No coin. The nails are bitten short, the right wrist crossed with a rope burn fresh enough to still sting.

“A name is not a key,” I say.

“I’m not here to spend it,” she says.

“Everything gets spent eventually.”

I push a mug of water across the bar. Not ale. Water.

“Drink if you want. Don’t spit it on my floor.”

Sella does not take the mug. From inside her cloak, she produces a small leather pouch, worn and blackened at the edges. She sets it on the bar between us.

I do not touch it.

“Open it.”

She unlaces it with careful fingers. Inside, wrapped in oilcloth, is a whetstone. Not large. Not fine. The kind a soldier carries because it can sharpen a blade, a hook, a knife, and one of her fingers if she is desperate enough.

My breath changes, though I do not let my face.

The whetstone has two cuts in the stone. One is a thin straight line. The other is a hooked mark, quick and ugly, made by a knife edge when the hand is not steady enough.

I know that hooked mark.

I made it.

“A whetstone is a small thing,” I say.

“It was hers.”

“It could be any number of things.”

I still do not reach for it. The mercenary in me wants to know what she can pay, what she can sell, what she will do if the answer disappoints her. The friend who carried Kestra’s pack wants to believe the ash was all that ever had to be done.

“Why bring it to me?” I ask. “Why not the constable? Why not the magistrate?”

Sella’s jaw tightens.

“Because the magistrate’s clerk has a sealed envelope with her name on it. He’s asking for testimony from the surviving guide.” Her voice does not rise. It goes flat, the way a blade goes flat before it cuts. “He is offering coin to a boy who says he saw you drag Kestra’s body from the cleft. I do not believe the boy. I believe my sister.”

The tavern sounds fade again. Rain. The creak of the sign. The low smell of beer warming too long.

I look at the whetstone.

Then at Sella.

“You say a boy is lying.”

“I say I have not met him.”

“That is not the same thing.”

“No,” she agrees.

I pull my hand from the bar’s edge. My knuckles are pale.

“Sit down,” I say.

Not kindly.

Not unkindly.

“Keep your hands where I can see them.”

---

[...continues through Turn 5...]

As you can notice:

Hemmingway creates a real short story. No markdown, no section headers, no machine friendly pattern, just a proper told story. I'm not a native english speaker, however it feels more like a "human-written" text.

Qwen followed the prompt well and the story is good as well, but it feels rather "technical".

Bottom line: for character work or fiction or your everyday local email writer, it's a very interesting 27B at \~32 tok/s on a MacBook M5 Max.
For a generalist or coding assistant, the base Qwen is of course still the pick.

Full runs + llm judge notes: https://llm-bench.io/models/hemmingway-1-oq8e-mtp

▲
0
 
11👁
r/LocalLLaMA · u/BopSupreme · 9d ago
Future of Local AI after OpenAI DevDay

Codex Cloud, Dots, and the existing remote Codex all allow users to untether themselves from their PC, and now untether themselves from even owning a PC with their server based Codex Cloud and Dots that can run 24/7. Combine this with Meta’s & OpenAI’s planned hardware releases and the goal is clear: work around Microsoft/Apple’s control of user hardware, provide AI devices that complement and eventually replace iPhones - culminating in a user base that owns no hardware and relies on a subscription to access AI. Meta’s hardware is obvious spyware, Apple’s new “always-listening” Apple Watch sounds pretty similar, their camera-enabled Airpods sounds atrocious for privacy, and OpenAI’s device is unconfirmed.

The end result? Instead of a Matrix-like AI takeover of humanity users are instead expected to purchase their own devices and subscriptions that provide mega-tech companies with all of their physical and digital data 24/7. The data volume is so large only AI can process it. A select few billionaires decide what their closed-source AI does with the data.

The resistance? Governments that oppose the USA and individual users who were rich enough to afford local hardware and utilize Chinese and other open-source models, likely blacklisted by the USA. To buy a 5090 customers now have to sign a waiver, as a result of US law. It’s only the beginning.

Ironically the “bad guys” like China, North Korea, Iran, Russia - will probably end up as the only large entities keeping open-source AI and local LLMs alive. I would expect the largest AI companies to eventually gain more leverage over the US Gov & Nvidia; unless Nvidia steps up to the plate and champions local AI

▲
0
 
20👁
r/LocalLLaMA · u/BrilliantSecret143 · 9d ago
NIRNAY: 450M decision model beats Jev on Banking77, runs on CPU

Built a small open decision model for intent classification and routing.
450M params (Laya fork plus \~30M), one forward pass gives calibrated
probabilities, no text generation.

Banking77 test, 3,080 cases: \*\*0.8792\*\*, Brier 0.208, fitted ECE 0.045.
Same cases through Jev 1.13.0: 0.803. Caveat, stated plainly: we
fine-tuned, Jev answered zero-shot. Fine-tune beats API on your own
data, that is the thesis.

Runs local: 209ms on M4 GPU, 361ms on CPU, batch-1, PyTorch. No GGUF
or Ollama build yet (custom heads need converter work), so bring a
Python env for now.

\\\`bash
pip install git+https://github.com/eulogik/nirnay
\\\`

\\\`python
from nirnay.agent import NirnayAgent
agent = NirnayAgent(device="cpu", checkpoint\_path="phase\_b.pt", enable\_byte\_path=False)
out = agent.system\_one("My card was charged twice.", {"intent": {
"type": "choice",
"instructions": "Classify the banking intent.",
"criteria": {lab: lab.replace("\_", " ") for lab in BANKING77\_LABELS}}})
\\\`

(BANKING77\_LABELS comes from nirnay.data; full snippet in the repo
README.)

Also in the repo: the two training collapses we hit and fixed (scale
runaway 150x, silent usage collapse to 1/77), a 9-page paper draft,
and every eval as raw JSON. JevBench-hard is weak (0.396, long docs),
published as-is.

Repo: github.com/eulogik/nirnay.
Weights: huggingface.co/eulogik/nirnay-450m.
Apache-2.0. Built by Eulogik.

▲
0
 
16👁
r/LocalLLaMA · u/XInTheDark · 9d ago
A self-hosted agent app that runs each task in its own container, and works with any models

Hi everyone!

I've been working on this agent platform for 7-8 months and recently made it open source: https://meowbert.com

I know there are a lot of this same type of projects at this point. I built this one because I wanted something clean that's self hosted, does its job properly, and is suitable for doing long projects and run tasks autonomously.

Each task runs in its own Docker container with things like a shell, a browser, Python, Node, and tools for Office documents and PDFs. The files and memory are saved in projects. Tasks can also run on a schedule and send the result to Telegram, Discord, or email when they finish.

For example, I have a scheduled task where the agent runs regular health and security checks by querying logs and system info, and notifies me if there is an issue.

Your custom skills can also be added directly to a skills/ folder in the root, and I am planning to make it easier to set up for others.

It works with any server that supports the OpenAI Responses API. I've mainly tested it with both Codex models and Qwen 9B via Ollama, on a small VPS, and it has helped me a great deal in my projects. APIs that only support chat/completions won't work yet. I plan to add support for them very soon, as I know it's widely used.

Task view

Known limitations, I am trying to improve on these:

\- It needs the "/v1/responses" API format, I know that rules out some setups, and adding support for them is on the list

\- Smaller/older models struggle with tool calling as usual

\- The sandbox image is x86-64 only for now.

It's AGPL licensed and the code is on GitHub: https://github.com/XInTheDark/meowbert-ai-agent

I'd really appreciate any feedback. A big reason I am posting this was to learn from the community and from more experienced devs. Issues and feedback of any kind are welcome!

💬 10 (+1) open on reddit ↗
▲
0
 
19👁
r/LocalLLaMA · u/Robert-Prisacariu · 9d ago
I built OpenBot: open-source AI teammates for your Mac that can run on local models with Ollama (MIT)

Hi r/LocalLLaMA, I'm Robert, the developer. I just released the first public beta of OpenBot, and I wanted to share it here because local models are a first-class option, not an afterthought.

What it is: a small team of AI teammates that runs on your Mac. Each teammate has a name, a job, its own workspace and its own browser. Talk to one, or give a group a task that runs in order: "Nova, find three restaurants. Scout, check their hours." Scout waits for Nova's list.

The model side:

  • Point any teammate at Ollama. Each teammate can use a different model.
  • Or use any OpenAI-compatible API, a free Gemini key, or a ChatGPT, Claude, Grok or Copilot subscription you already have.
  • Mix them, e.g. a local model for drafting and a hosted one for research.

It asks before acting. Reading and searching happen on their own. Sending, buying, signing in or submitting always stops and shows you the exact website and button first.

Also: Word and Excel files as results, routines ("every Monday at 9…"), Telegram, Discord, iMessage and "Hey Siri, Ask OpenBot".

Install (macOS 13+):

curl -fsSL https://openbots.foundation/install.sh | sh

No admin password, and it checks the download's SHA-256. The installer is readable in the repo (scripts/install.sh).

Honest limits: it's a beta. The app is ad-hoc signed, not notarized. It works while your Mac is on. Mac only for now.

A question for you: which local models have you found reliable for tool use and browsing? I'd like to ship better defaults.

https://github.com/PrisacariuRobert/openbot

▲
0
 
14👁
r/LocalLLaMA · u/zmarcoz2 · 9d ago
One-prompt GTA style game with qwen3.8-flash-next-iq3_s post image

The prompt: make a gta-style game using three js

it took 3h 18m 6s

Total tokens: 22,845,556 — 22,533,061 input + 312,495 output.

Hardware:
RTX 4080 super 16GB

64GB RAM DDR4

Windows 11

Inference engine is strata running at \~40 tk/s and a custom mini swe agent v2 with the tools: powershell, edit\_file, view\_image, read\_file, search\_files

The harness has guards for tool failures (iq3 fucks up a lot) and auto-compaction.

logs: https://gist.github.com/Cirius0310/c26197240ad20ef04e45a78e36031d6e

💬 15 (+1) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/Storge2 · 9d ago
Comparing Compute of Supercomputers like Vera Rubin and TPUv7 post image

Hello guys so I made a youtube Video comparing the Compute per MW or better said per 6.5MW which is roughly one Vera Rubin Pod in order to see where the world currrently is standing at and was surprised at the fact that Nvidia is basically the best Price/Perf hardware despite the insane Prices. Check it out if you want. Also I am very much welcoming tips on how to improve the quality. I made the video with opus 5.5 and Hyperframes.

▲
0
 
20👁
r/LocalLLaMA · u/EmilPi · 9d ago
I don't understand whether uncesored/abliterated/heretic/fusion/bla-bla models give any value for the open-weight community

Using uncensored model gives sort of sense of power, I suppose, for some people; but what else?

(UPD.: Usecases well-explained in the comments: cybersecurity, storywriting, law, medical, criminal forensics).

If I am wrong, prove me wrong, please, or just say you really need something different from it. I sure felt a frustration when (just one of the ton of examples) e.g. you ask how did Peter Pettigrew die, and the model suddenly starts a litany it is a harmless assistant.

The GLM-5 helped HF against OpenAI cyberattack without being uncensored. If you need an uncensored model to understand political hypocrisy, well, you haven't grown up yet. You want to protect your property against a burglar? Find a competent consultant, instead of potentially hallucinated advice from the LLM (and the more uncensored the model, the more it hallucinates).

The only measurable goal I see the uncensored models serve now is a pretext for the corps to regulate people, who are just happy having Gemma4.x/Qwen3.x/DeepSeek-4.x/GLM-5.x do some stuff for them. Not sure that 5% of legitimate use cases (which I believe exist, but are only substitutes for a classic search or consultation) are worth it. What if I (and I believe a majority of the open-weight models' users) don't need waifu/goon/bioweapons or meth recipes/propaganda generation/cyberattacking/scamming capabilities?

💬 79 (+3) open on reddit ↗
▲
0
 
20👁
r/LocalLLaMA · u/ag789 · 9d ago
CopilotKit

The 'AI' world is moving plenty fast, enter CopilotKit
https://github.com/CopilotKit/CopilotKit#what-you-can-build
'agents' are coming in, draw charts, type your document, spreadsheet, operate your web browser, write your email, make presentations.
It would probably leap off the screen into the physical world

It is probably a 5yo's definition of 'AI' , that's coming true

The 'agent loop' becomes practically, all apps, all frontends (webui, gui, mobile) everything anything , anything connected to an LLM.

I think Local LLM would be part of that after all.

▲
0
 
12👁
r/LocalLLaMA · u/artur_oliver · 9d ago
600M parameter model for transcription, super reliable.

Hello community,

I have been thinking of building an app for the company that just gets the calls from the automated answering machine to text, but I have huge problems with the quality of the translation. That's why I think I can use this model.

My idea is to have a summary table every 20-30 calls about the content or important recalls I need to do.

I run a really busy office, we get about 100 cals a day if not more.

I want to get that but the devils are in the details, what do you think?

What features should be implemented first or even complementary to it?

Thanks

▲
0
 
8👁
r/LocalLLaMA · u/power97992 · 10d ago
Next year, the pro models will have 8-10 T parameters, who will have enough vram to run them?

Deepseek said they will release an 8 T model later and qwen said they will have a 10 T model and kimi will probably follow suit. The flash models will probably be around 1 -2 T parameters. Then only companies and corporations And cloud providers and rich people will be able to afford to run these pro models and fairly rich people for the flash models . At this rate, you would need 9 512 gb m5 ultras or 48 rtx 6000 pros to run A 4.4 bit 8T model with full context ? That is probably 153k for the ultras or 768k for the rtx pro Gpus plus probably another 100k for the other parts. I guess either use the cloud or people will use smaller models like qwen 5 27b in the future but most people won‘t be able To run the biggest models locally. In fact, most people will struggle to run a 4.4 bit 1 t flash model locally. It will cost 100-120usd/h just to host the mod in the cloud

▲
0
 
9👁
r/LocalLLaMA · u/ag789 · 10d ago
The Agent loop is probably what matters (for local LLM)

The commercial ones seemed to want to monopolize the agent loop.

Today the chat completions API is probably a 'defacto' way of talking to the models

https://github.com/ggml-org/llama.cpp/tree/master/tools/server#post-v1completions-openai-compatible-completions-api
https://vercel.com/docs/ai-gateway/sdks-and-apis/openai-chat-completions
btw, credit goes to the origin:
https://developers.openai.com/api/docs/guides/completions

A thing is, more recent efforts seem to be instead offering just an \*agent\* at the API and putting this \*agent\* layer between you and the model.

local LLM will remain \*very\* important because as is currently, you own the agent loop.
You write that "small little" front / stub that is the agent loop talking to the LLM.
it is day and night difference , practically 2 different universes

▲
0
 
15👁
r/LocalLLaMA · u/AdRepulsive7837 · 10d ago
Tensorfold runs Qwen3.8-27B really well on m5 pro mac mini, tps beats MTPLX

Came across this popular open source inference engine Tensorfold https://github.com/ashhart/TensorFold

Using their official Vontra/Qwen3.8-27B-MLX-4bit with drafting model z-lab/Qwen3.8-27B-DFlash2, I can reach 40-60 tps on mac mini m5 pro. AGAIN, it is PRO on mac mini, not even ultra studio.

For me, it is the first time (on mac ecosystem) that an inference engine to beat MTPLX. I have tested omlx, dflash2, mlx, llama-cpp, lm-studio, unsloth in the past few months, and none of them come close to MTPLX (running Qwen 3.8 optimised for speed, roughly 4bit?)

The more exciting part is that this enables me to seriously consider about replacing my RTX-3090ti with this mini running tensorfold as the main inference server setup. That old 3090ti, running ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with MTP IQ3\_S (12.1 GB), reaches 50-70 tok/s, which is, in my opinion, similar to the 40-60 tok/s I achieve with mac mini. The only one caveat is htat the 3090ti still has like 4x faster prefill than mac mini.

Spec: M5 pro, Mac mini, 64gb, 1TB SSD

Testing harness: pi coding agent without any packages install yet.

Model: Qwen3.8-27B 4bit

What's your thoughts on Tensorfold?

💬 27 (+2) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/jaybsuave · 10d ago
Help choosing compute for a student? 4k budget

My university is going to give me 3k for a laptop and I was wondering what type of computer I should get? I already have a MacBook for school, and a desktop with a 4070 12gb and 64 gb. Any suggestions? I wanted a Mac mini but I can't use it ok Windows obviously and the DGX is too expensive. I can throw an extra 1000$ in as well if I need too so my budget is 4k. Thanks