111 posts · 1 sub · RSS
← prev Sep 29, 2026 → Sep 30, 2026 next →
2026-09-29 → 2026-09-30 hourdayweekmonthyearall
allr/LocalLLaMA
▲
2455
+162
102👁
r/LocalLLaMA · u/BannedGoNext · 10d ago
Anthropic just dropped the greatest advertisement for GLM ever.

Like.. yea bro, I knew GLM was cool. Now everyone does.

💬 531 (+26) open on reddit ↗
▲
1326
+76
72👁
▲
1220
+66
75👁
r/LocalLLaMA · u/Dany0 · 10d ago
AMD's new 256 core EPYC has 16-channel DDR5-12800, 91% memory bandwidth of an RTX 5090

Here are some inspirational quotes you can put into the comments:

  • God is dead and we killed him
  • I am become death
  • All this for 1.5 tok/s?
  • Sir this is LocalLLaMA not RichPeopleofLocalLLaMA
  • Sweet! A 2TB DDR5-12800 RDIMM kit is going to cost only 2 kidneys and a small micronation's GDP
💬 246 (+1) open on reddit ↗
▲
731
+57
70👁
r/LocalLLaMA · u/xenovatech · 9d ago
We just open-sourced the world's fastest WebGPU kernels for local AI on Hugging Face post image

The collection includes kernels for more than 200 common ML operations, all of which can run entirely locally in your browser on WebGPU. We're also working to upstream these optimizations to Transformers.js, ONNX Runtime Web, LiteRT.js, and more!

Kernels: https://huggingface.co/kernels?platform=webgpu
Blog: https://huggingface.co/blog/webgpu-kernels

💬 45 (+2) open on reddit ↗
▲
229
+25
45👁
r/LocalLLaMA · u/Significant-Price695 · 9d ago
Oído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)

I'm part of the Lokutor team that built this.

Model: NVIDIA Conformer-CTC Small (13M params, int8). It runs on an ESP32-S3 with 8 MB PSRAM, no GPU or NPU. LibriSpeech WER is 3.7 / 8.2, versus 6.3 / 15.9 for Whisper tiny.en on a laptop. Under real noise (DEMAND: car, kitchen, cafeteria) plus babble and reverb, mean WER is 8.4 vs 12.1 for Whisper tiny.en. You can try the exact chip arithmetic on your laptop mic with live_demo.py. https://github.com/lokutor-ai/oido

💬 40 (+1) open on reddit ↗
▲
458
+24
68👁
r/LocalLLaMA · u/Rombodawg · 9d ago
Least to most expensive (Somewhat modern) GPU's with 32gb of vram (Under $1600) Based on ebay listings post image

I was researching prices on ebay and fed claude a bunch of images of listings. I had it make a chart and thought it would be useful to share.

💬 260 (+24) open on reddit ↗
▲
262
+21
45👁
r/LocalLLaMA · u/jacek2023 · 10d ago
Reflection 70B was released two years ago (September 2024)

You may think that jev, OpenClaw or TurboQuant are super cool, but actually the coolest LLM invention happened two years ago

As we all know, the best source of reliable information about LLMs is YouTube:

https://preview.redd.it/4tv0ctbwyfsh1.png?width=2544&format=png&auto=…

Back in September 2024, Reflection 70B appeared out of nowhere and was announced as an open-source model that supposedly destroyed GPT-4o

There was only one small problem. People downloaded it. And tested it :(

https://preview.redd.it/kamhrp9izfsh1.png?width=1514&format=png&auto=…

It turned out that Reflection 70B was basically a Llama 3.1

https://preview.redd.it/ghmof4l50gsh1.png?width=1524&format=png&auto=…

but at the end the mystery was solved

https://preview.redd.it/44eb3jw90gsh1.png?width=1530&format=png&auto=…

Let this be a moment of reflection on the current hypes in LocalLLaMA.

great summary by Maziyar PANAHI https://x.com/MaziyarPanahi/status/1838559480658710982

▲
270
+20
26👁
r/LocalLLaMA · u/Spiritual_Impress_30 · 9d ago
Thank You, Mradermacher.

best iq quants in the biz, got me gemma 4 26b to run 75tok/s tg and 1500 pp on 2x 4060 8gb using lmstudio serving to hermes, much work has been done.

▲
338
+14
58👁
r/LocalLLaMA · u/politefella0 · 10d ago
Deepseek Harness app is out now!!!

Downloading now.

💬 91 (+5) open on reddit ↗
▲
54
+14
37👁
r/LocalLLaMA · u/Designer_Cost8989 · 9d ago
Index-Translate: 150 text languages, plus document translation, multilingual subtitles and dubbing

Quick update: we’ve opened a free public API for Index-Translate-35B-A3B! It’s OpenAI-compatible, and you can get started with our Python script—no extra dependencies needed.

I'm part of the BiliBili Index LLM team. We're sharing Index-Translate and its companion models for translating text, documents, and videos.

Index-Translate supports 150 text languages, with 2B, 9B, and 35B-A3B (preview) options. You can specify terminology, writing style, and output format—for example, keeping product names consistent, translating in a casual tone, or preserving JSON and placeholders during localization.

There are also models for more specific workflows:

  • Index-NativeLong: translate whole documents, using their context to help keep names and terminology consistent across passages.
  • Index-Homura: set a syllable budget for translated lines, useful for fitting a dubbing script.
  • Index-Echo: generate multilingual subtitles or translate speech into speech, using the source speaker's voice as a reference.

The attached video shows English → Japanese dubbing, followed by an English clip with subtitles in six languages. The 150-language coverage applies to the text models; Echo supports a smaller set of language pairs.

https://reddit.com/link/1wugf2t/video/9zajzj72zpsh1/player

Code and released weights are Apache-2.0.

Try the demo · GitHub · Models

What would you try it on—video subtitles, game localization, or documents? We'd especially appreciate examples where it gets your language pair wrong.

💬 22 (+3) open on reddit ↗
▲
312
+12
58👁
r/LocalLLaMA · u/writesfw · 10d ago
Are you worried about a potential ban of Chinese open weight models?

Anthropic released the GLM article today. Trump is getting very involved.

Do you foresee Chinese open weight models getting banned soon?

💬 506 (+12) open on reddit ↗
▲
245
+12
29👁
r/LocalLLaMA · u/LambdaHominem · 9d ago
AI CEO Interviews (2026) post image
💬 34 (+1) open on reddit ↗
▲
37
+11
24👁
r/LocalLLaMA · u/Defiant_Ranger607 · 10d ago
What kinds of problems are still fundamentally hard for LLMs?

I played a game of a custom chess variant against an claude opus 5.5, and it beat me.
The game combined several rule changes: the board wraps around from the h-file to the a-file, knights move three squares in one direction and one sideways, and captured pieces can be dropped back onto the board, as in crazyhouse. I gave the model the rules, the starting position, and a board diagram, then asked it to reply with one legal move at a time.
Also I recreated this board game https://nika-game.com/ and played with claude, and it still win, although I believe it is really unpopuplar and old game without much training data available (claude didn't even know the rules initially)

I used to think chess exposed a fundamental limitation of LLMs: keeping track of a changing board, following exact rules, and planning ahead seemed like a poor fit for a language model. This game made me reconsider that assumption.
So I’m curious: what broad classes of problems do you think LLMs still can’t solve reliably? Are there limitations you consider fundamental/archiectural, rather than problems that might improve with better models, more computation, or tools? What would be a good test?

💬 87 (+6) open on reddit ↗
▲
101
+9
27👁
▲
218
+8
47👁
r/LocalLLaMA · u/jacek2023 · 9d ago
add GLM-5.3-Flash (GLM5-Next) support by timkhronos · Pull Request #27773 · ggml-org/llama.cpp

now you can use GLM-5.3-Flash on your home computer

💬 89 (+10) open on reddit ↗
▲
184
+8
34👁
▲
184
+7
31👁
r/LocalLLaMA · u/Educational_Sun_8813 · 9d ago
Preorder for new AMD Ryzen™ AI Max 400 Series 192GB from framework just started

Framework Desktop
Framework Desktop DIY Edition (AMD Ryzen™ AI Max 400 Series) 192GB

💬 129 (-1) open on reddit ↗
▲
123
+7
24👁
r/LocalLLaMA · u/autoencoder · 10d ago
Another case of censorship post image

I was optimizing my diet using a cloud AI provider, and I found my Pi agent stuck like this. Looks like I had too many mushrooms lol.

💬 24 (+1) open on reddit ↗
▲
82
+7
34👁
r/LocalLLaMA · u/Skyline34rGt · 9d ago
BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B

"AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise.

AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories.

  • Architecture: Dense Qwen3.8-compatible multimodal model
  • Parameters: 27B
  • Context length: 262,144 tokens

[](https://huggingface.co/BAAI/AREX-2#key-features)Key features

  • Long-horizon self-improvement: turns extra test-time rounds into useful solution refinement.
  • Feedback-driven reflection: reads scores, logs, errors, and timings to decide what to change next.
  • Cross-domain performance: training on coding and machine-learning tasks also improves the model's deep-research performance.
  • Long-horizon reasoning: sustains productive iteration as the task budget grows."

Gguf's - https://huggingface.co/mradermacher/AREX-2-GGUF

💬 29 (+2) open on reddit ↗
▲
426
+6
47👁
▲
163
+6
31👁
r/LocalLLaMA · u/WebAssemblyMan · 9d ago
DeepSeek now trained on Ascend 950 post image

26 months ago Liang Wenfeng said:
"Someone must step onto the frontier."

Now they are training their models on Ascend 950

▲
136
+6
20👁
r/LocalLLaMA · u/yahbluez · 10d ago
Is AI Profitable Yet?
▲
89
+6
44👁
r/LocalLLaMA · u/pubudeux · 11d ago
First few days of qwen3.8-flash-next on 4x R9700 - it's been really interesting so far post image

Here's a metric dashboard giving an idea of the last few days.

Been testing with a variety of different agentic coding use-cases, mostly using a pi harness.

qwen3.8-flash-next has seriously exceeded my expectations (used https://huggingface.co/tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8)

Both speed and quality have surprised me, given that I can get 3-5 concurrent streams going with \~100t/s gen each, and single stream easily gets to 150+t/s. Prefill is 10k+t/s

💬 67 (+2) open on reddit ↗
▲
49
+6
27👁
r/LocalLLaMA · u/netherreddit · 10d ago
Inference Engines will become a series of one-offs

ninfer, dwarfstar, Splash, llamAmpere, gufo, etc.

We've all seen them popping up, great tok/s, people loving them. Forks of llama.cpp or another engine, or made from scratch.

For better or worse, the list will continue to grow

They work so well because they dodge a main difficulty of software, generality, and just implement for a single model/hardware combo (or a few), and then optimize kernels/compute graph for that one case. Highly 'overfit' codebases that beat well-known inference engines (llama.cpp, vLLM, etc) and incidentally will be completely forgotten in 6 months.

But new ones will take their place...

THESIS

One-off engines will become the norm. Llama.cpp, vllm, etc, will not make sense for most people to use, because they're slower

A few axioms you probably accept:

  1. the more general a codebase, the harder it is to cleanly fit new features in over time. This reduces the pace of innovation. The smaller, the faster
  2. AI coding is getting better and cheaper. Thus, the barrier to creating an inference engine is dropping
  3. many coding tasks are difficult to completely give to AI (or a human) because they are not fully specified. But "Make tok/s go up" in an inference engine fork for one hardware/model combo is fully specified, and is therefore a great candidate for 100% autonomous implementations to be perfectly fine in terms of quality (as long as correctness tests are included, which is trivial). No human bottleneck.

String these axioms together, and I arrive at

  1. general engines, like llama.cpp, will perpetually lag behind these one-offs in development speed, and therefore token speed
  2. none of the one-offs will be able to maintain generality and dev speed over time
  3. one-off inference engines for specific hardware/model combinations will continue to proliferate, and be loved

Thank you for coming to my ted talk

IMPLICATIONS

  1. This thesis brings up an interesting question: What elements of inference engines WILL remain in common?

Most obvious example: it would be annoying to have a different usage API for every engine, so we already standardized on OpenAI API compatibility years ago.

Is that also true for cli arguments/configs? The packaged gui (llama-server)? Benchmarking tools (llama-bench)? Logging, model format, Etc?

One-off engines that replicate the experience of everything wrapping the inference itself will be more seamless to adopt. Case in point, the main reason I haven't tried any of these new one-off engines myself is it was annoying enough to figure out how to drive llama.cpp properly. Don't want to do that again unless it's really worth it.

There's probably a place for an open source project that standardizes all of this and makes it easy for one-off engines to adopt.

  1. Maybe we'll see more 'half-general' inference engines that just target one hardware platform. So still general on the dimension of models, but not on hardware. Splash could be an example.
  1. Nobody wants to continuously scan github/reddit/x for the best inference engine for their model/rig. Some will just have their agent custom make one. But I think a larger number will not do that. So, hardware-specific communities will form. Think r/appleM2Max32gbLLM and r/4090And64gbRamLLM, etc (however that actually ends up organizing. exaggerating a bit on the names.)

ALTERNATIVE FUTURES

Scenarios where the one-off future doesn't happen:

  1. General inference engines find a way to 'plugin-ify' the model/hardware specific kernels and compute graph so you can swap them at runtime. So you'd download not just a .gguf, but also an .inference\_recipe to go with it, which contains the optimizations for your specific hardware, for that specific model. Maybe those optimizations will make it into llama.cpp mainline in 3 months, but you can use them today, without a fork.
  2. General inference engines find a way to AI-ify their workflow so much that they maintain quality and codebase coherence but also achieve the same development velocity for each model/hardware platform as the one-offs. I think this is the best for everyone involved.
  3. The full vision of something like MLIR, or Mojo is realized to a sufficient degree. ie writing hardware-optimized kernels is fully and invisibly done by compilers, no hardware-specific tinkering needed anymore for each silicon platform) Then, inference engines that cover all hardware/models would be much more manageable to maintain and add features to. btw, if you really want to have an impact, solve this. The world will thank you for centuries to come. Unfortunately not many people have even conceptualized this as a goal.

P.S. there's growth in a dimension separate from single model/hardware engines which is more like "frontrunning a general inference engine's features because it's slower to pull in PRs". Freetoken, BeeLlama, etc. Not as model- or hardware- specific as the other examples I've given. Haven't thought much about that dimension.

💬 185 (+8) open on reddit ↗
▲
89
+5
18👁
▲
50
+5
22👁
▲
29
+5
19👁
r/LocalLLaMA · u/XiRw · 9d ago
What was the mafia meeting of the tech criminals about at the White House?

Or should I assume it was mainly about trying to stop/ban/regulate China and open models,

▲
42
+5
29👁
r/LocalLLaMA · u/bring_back_the_v10s · 9d ago
How smart is the IQ3 family of Qwen 3.8 Flash Next for coding tasks?

I've been closely following the rise of the Strata inference engine and as someone with 28GB VRAM and 32GB RAM I'm itching to buy 32GB more RAM just to use Flash Next. But of course before I make such a financial commitment as a member of the GPU-poor class like myself, first I need to make sure the IQ3 quants are worth it. My use case is primarily agentic coding tasks with harnesses like Pi or OpenCode.

Has any of you guys used Flash Next IQ3 for relatively serious coding? Is it worth it? Compared to, say, Qwen 3.8 27B Q4 or Q5.

https://github.com/Niko1221/Strata/

💬 72 (+5) open on reddit ↗
▲
5
+5
15👁
r/LocalLLaMA · u/politefella0 · 9d ago
What’s better? A very small quant or a large model distilled into a small one?

I see people asking (begging, take it as humor) for 0.000001 bit quants but isn’t a lower quant essentially going to hurt model’s quality and tool calls?

💬 20 (+2) open on reddit ↗
▲
60
+5
27👁
▲
55
+5
18👁
r/LocalLLaMA · u/jacek2023 · 10d ago
IQuestLab/IQuest-Q1 · Hugging Face

IQuest-Q1 is a Mixture-of-Experts (MoE) model developed by IQuest for agentic coding, reasoning, and multi-step tool use. It comprises approximately 320B total parameters, with an estimated 15B parameters activated per token.

▲
218
+4
32👁
r/LocalLLaMA · u/Acceptable-Cycle4645 · 10d ago
Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models post image

I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.

Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.

And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.

The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.

All figures here: https://github.com/0xShug0/audio.cpp/tree/main/assets/figure/

💬 24 (+1) open on reddit ↗
▲
105
+4
32👁
r/LocalLLaMA · u/Loginhe · 10d ago
[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw post image

We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed.

Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16.

What's inside

  • Four quantized GGUFs, 2.40 to 3.50 bpw (66.4 to 83.6 GB), and the BF16 vision projector
  • Expert-pruned Coder GGUF, 58.4 GB in total, of which 29.6 GB must remain resident
  • GSQ (Gumbel-Softmax Quantization): post-training scalar quantization that jointly learns grid assignments and group scales, closing most of the gap between scalar and vector quantization at low bit-widths while remaining deployable in standard GGUF types
  • RCO (Riemannian Constrained Optimization): enforces exact budgets by gradient descent on the task loss, without per-constraint tuning. It serves two roles in this release: assigning a quantization type to every tensor, and selecting which experts to retain in the Coder build, where it enforces several exact budgets simultaneously, one per layer

Results:

At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.

  • IQ3\_S (3.50 bpw, 83.6 GB): AIME25 100.00, GPQA-Diamond 92.93 against 91.92 for BF16, LiveCodeBench v6 86.86 against 87.43. Task average 93.26 against 93.12.
  • IQ3\_XXS (3.00 bpw, 75.8 GB): AIME25 100.00, GPQA-Diamond 91.41, LiveCodeBench v6 86.29
  • Q2\_0 (2.40 bpw, 66.4 GB): zero-shot average 78.00, above the BF16 value of 76.94, at approximately one fifth of the size

Coder (capability pruned model):

Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits.

The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.

  • SWE-bench Verified: 75.60 against 82.80 for BF16, retaining 91.3%
  • LiveCodeBench v6: 86.28 against 87.43, retaining 98.7%

Both measured at xhigh reasoning effort.

Links

Both repositories ship the complete per-tensor RCO allocation.

The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome.

From the ISTA Deep Algorithms and Systems Lab.

💬 77 (+5) open on reddit ↗
▲
8
+4
23👁
r/LocalLLaMA · u/IngwiePhoenix · 9d ago
Penalties of PCIe generations? (2x R9700)

I just bought the GPUs after deliberating and debating for over two years. With costs not coming down any time soon and me just wanting to get this massive todo-box ticked, I decided to just YOLO it; the GPUs are the most volatile, followed by RAM, rest seems more or less stable.

But actually, RAM is one of the reasons I am unsure about wether to chose a SP4, 5 or 6 based board. I am most familiar with AMD CPUs, so that is where my tendencies lie. Unfortunately, RDIMMS are going to absolutely undress me... x.x

However, if I could stick to a DDR4 / PCIe Gen4 setup, that would save a pretty penny. Now I do not intend to offload to system memory, but even a small, single-stick of DDR5 RDIMM is stupid expensive - DDR4 is fine.

The question is: What is the penalty of PCIe Gen 4 versus 5 in regards to inference? I will be using llama.cpp with ROCm, fronted by llama-swap, utilizing both GPUs for inference and VRAM pooling (so, 64GB in total).

Thanks! =)

💬 47 (+1) open on reddit ↗
▲
9
+4
27👁
r/LocalLLaMA · u/whatyathinkk · 9d ago
Do I need a UPS?

I know I could ask in some hardware subreddit, but I'm curious to know what people with multiple GPUs and expensive inference setups think about this.

I just moved to a new place and here the lights go out pretty frequently. 3 times over the last week, I came back to my computer being off due to a blackout (I guess it's a blackout, the entire neighborhood looses light for a few seconds/minutes). I have a desktop computer with 2x RTX5080s.

Do I need to buy a UPS to protect my computer from this? I get mixed answers about this topic. I don't mind my workflows being interrupted when the computer turns off, the only thing I'm worried about is the hardware being damaged. I have a good PSU, is that enough to protect the hardware?

💬 75 (+2) open on reddit ↗
▲
161
+3
50👁
r/LocalLLaMA · u/MLDataScientist · 10d ago
Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine

I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling and the developer fixed them.

Inference engine (only runs on Nvidia for now; AMD support is experimental): https://github.com/Niko1221/Strata

Here are some metrics with screenshots. My laptop has 5070ti 12GB VRAM, 64GB ddr5 RAM, Intel 275HX CPU, gen4 SSD.

Aquarium test \(unsloth studio connected via local API\)

The model I used was https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/tree/main/IQ3\_XXS which has a good quality for its size. Above, the model generated the aquarium test. At 43k context depth, it was running at 51 t/s. Stock llama.cpp reached only 23t/s with the same quant.

32k context read at 1500t\/s \(unsloth studio via local API\)

This quant could only reach 100t/s PP with stock llama.cpp using the same quant. Strata was reading 32k context text at 1500t/s. This is way above my expectation. This quant can load with up to 200k context at 8bit. However, I was only using 131k context.

Memory utilization

As you can see it is utilizing 11GB VRAM and 56GB RAM (includes system/OS programs).

This engine is specifically built for one model only and only select ggufs (ISTA-DASLab) work with it. You can use IQ3\_S from ISTA-DASLab which they claim recovers full model's performance on coding benchmarks. I tested IQ3\_XXS for some time and I would say it is an excellent model.

I never thought 12GB VRAM would be enough to run frontier models from 6 months ago locally on a laptop. What a time to be alive!

💬 143 (+3) open on reddit ↗
▲
70
+3
20👁
r/LocalLLaMA · u/Elouakili_Flexy · 9d ago
Another Ling model comes out, same receipt, 2 weeks free to use, then open source. Chinese labs do contribute a lot to open source community

Ling-3.1-flash: \~560B total params, \~25B active/token, up to 1M-token context.

Across work, coding & healthcare: 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional.

▲
64
+3
30👁
r/LocalLLaMA · u/ciprianveg · 10d ago
Do you need some extra memory on your DGX Spark? post image

​

I created this repo to help the DGX Spark users that have a spare 10-24 GB GPU at home to squeeze some extra memory out of a single Spark or a Sparks cluster.

It moves the spec-decode draft model off your Sparks onto that GPU: the freed GB of memory can be used for extra context, or better quant quality. Supports both TCP and RDMA, shipped as eugr-vllm compatible mods:

https://github.com/ciprianveg/gb10-vllm/tree/main/remote-dspark

▲
54
+3
22👁
r/LocalLLaMA · u/jjusko20 · 9d ago
Update: Yandex/AliceAI 80B-A3B fine tune progress

loss curve \(taken from the last micro of every step, to explain the variation\)

some help from gemini 3.8 flash high

About 40% of the way done with the initial fine tune. The loss is so spiky because I accidentally used the last loss of each micro, rather than the average of each step

The training live stream is at: https://figure-bios-expect-cio.trycloudflare.com/ \- and it allows you to inspect any and all of the training data I'm using, if you're interested - I can also provide those roughly 3.5k examples as a dataset. It was generated from sftmill

Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wslgkw/watch\_me\_posttrain\_aliceaifoundation80ba3b\_from/

▲
7
+3
22👁
r/LocalLLaMA · u/bolche17 · 9d ago
Agent swarm coordination

Hello all!

Do you have any recommendations of tools for agent coordination and messaging to get them to collaborate on hard problems?

Ideally I would like a heterogeneous swarm, using local models as the workhorse and cloud models for reviewing, coordination, or simply to avoid overloading my relatively small local setup.

Do you have any recommendations or experience with this?

💬 24 (+2) open on reddit ↗
▲
8
+3
12👁
▲
16
+3
12👁
r/LocalLLaMA · u/Dev-in-the-Bm · 10d ago
Best approach for automatically tagging local music collection?

I don't use music streaming services much, and listen to music from my own local collection.

I don't use any local streaming servers like Plex or Navidrome, they wouldn't work for me because I use a dumbphone and play music off of my SD card.

I've manually built a bunch of mood based playlists so I can easily pull up a playlist with the music I want, but that's
obviously very tedious and inefficient.

I've been playing around with ML models to automatically add genre, mood, and other tags to my collection, the open models available today are insane.

The thing is I haven't been able to find any polished tools for doing this.

Most of what's available is either CLI or built for streaming servers.

Is there anything I missed?

Should I just setup a streaming server just for tagging the collection, or is there a better way?

💬 23 (+1) open on reddit ↗
▲
30
+3
20👁
r/LocalLLaMA · u/thebadslime · 10d ago
I created a personality test for models, need more TESTS!!

First 3 disposition results.

Hi there!

My name is Jerry and I recently built Enclosure, e deterministic environment for testing LLMs. It's a simulated office enfironment https://github.com/openconstruct/Enclosure with tools agents are used to, like slack, calendar, mail, chat and more. For conversations is uses AIML instead of a model, so it is totally deterministic.

The first test I made(had Claude make) is the one I had in mind when I designed Enclosure, a personality test for models. It's not a winnable benchmark, but rather a tool to help people find the mdoels that fit their workstyle best. It's called disposition and you can find it here: https://github.com/openconstruct/disposition

I tested the cheapest models on Alibab modelstudio already, after I refill my opencode go next month, I will probably test more. I am asking the community to test some models if you think it's a cool project.

Instructions are in the Disposition repo, and the submission repo is here: https://github.com/openconstruct/disposition

▲
21
+2
11👁
r/LocalLLaMA · u/Any-Lingonberry7411 · 9d ago
Best local model for Blender and game dev?

I have been looking at some local models, but even the smartest ones like GLM5.3 Flash and DSv4 Flash have a hard time creating coherent models in Blender and placing them logically in game engines.

Is this something that local models are just too dumb still to do good job at?

▲
3
+2
18👁
r/LocalLLaMA · u/dh7net · 9d ago
I need help to benchmark harness/model/hardware combination.

Hey! I'm trying to build the ultimate leaderboard to help everyone find the right harness/model combination given their hardware. (With all model variations and inference engine).

I own a GX10 and one 5090. And I'm trying as many thing as I can. (Happy to test anything, just let me know).

But I can't test hardware that I don't have.

So my ask is simple: Can some you do some testing on your own hardware?

I made this as easy as it could be: you just have to copy a prompt to your agent and your agent will fetch the test, pass the benchmark and send the answer to the website that will check if the answers are correct. You'll get a report out of it. And optionally you can offer your test to the community, so everyone can learn from your setup (it's just a toggle in the UI to confirm you are ok to share the results. I'll update the leaderboard when I'll have enough submissions.

Here is the link to contribute! https://airbench.ai/

Thanks in advance for everyone who will contribute!

https://preview.redd.it/ewgeumsqxpsh1.png?width=1402&format=png&auto=…

💬 22 (+4) open on reddit ↗
▲
3
+2
16👁
r/LocalLLaMA · u/FactorInternal3395 · 9d ago
Bartowski/AtomicChat Ornith 1.5 35B A3B + sharp template or Tiel Coder 35B A3B?

Tiel Coder 35B A3B from Peculiar Ragdoll is just Ornith 1.5 35B A3B with their own "coding focused" imatrix quantization and the sharp chat template built in. But how good really is that quantization? Other quantizers also focus on coding. Perhaps it would be better to just get Ornith quantized from Bartowski or AtomicChat, proven quantizers, and then add the chat template yourself rather than get the Tiel Coder weights? The end result would be the same, just the quantization is different, so the question is which quantizer is better?

💬 6 (+1) open on reddit ↗
▲
3
+2
10👁
r/LocalLLaMA · u/Dupliss18 · 9d ago
Locally runnable AI writing detection?

Is there any model or tool that can run locally to detect AI writing? Preferably something more up to date.

▲
3
+2
12👁
r/LocalLLaMA · u/Routine-Example927 · 9d ago
The search / extractor that worked for my Open WebUI.

I have an instance of OWUI setup for family usage. Works well with Gemma 4, however web search extraction was a weak spot, I wanted it to be:
1) Not reliant on paid APis

  1. Simple in setup

I found OpenSERP and made two PRs, one to OpenSERP itself to make it compatible with OWUI extractor and another one to OWUI to add OpenSERP search provider.

https://github.com/karust/openserp/pull/39

https://github.com/open-webui/open-webui/issues/27438

I'm quite happy with how it works - and given that it took me considerable time to find and set it up, I decided to share with the community.

▲
157
+1
32👁
r/LocalLLaMA · u/soyalemujica · 9d ago
If one hour of AI is costing me 0.12€ is paying for frontier a cheaper option?

Running Qwen flash next of even Qwen 27b dense, I can do any,burning sticking to flash due to its speed, and the kwh cost is at 0.25€ where I live in, ranging from 0.11€ to 0.35€, so I used chatgpt to help me calculate the total kwh consumption on my 7900xtx plus 9800x3D, and well that is the result.

Judging by this, if deepseek flash is indeed then faster to use per 1m token, does it mean that frontier is cheaper for me or am I calculating something wrong ?

💬 175 (-1) open on reddit ↗
▲
6
+1
14👁
r/LocalLLaMA · u/indiealexh · 9d ago
How to make best use of a Intel Arc B70?

I have a RTX 5090 in my desktop PC for local coding assistance and gaming and I have been loving it with Qwen3.8 27B Q4\_K\_XL.

I managed to get a B70 on the cheap and its great, but using it in split mode with the 5090 to ensure I get full context halfs my T/s (which is expected due to the memory bandwidth).

Would I be better off running the B70 with a smaller model to offload tasks to? Or just accepting the slower throughput and keeping the larger context?

I'd especially like to hear for anyone who has a similar mismatched GPUs setup.

💬 11 (+1) open on reddit ↗
▲
4
+1
13👁
r/LocalLLaMA · u/fallingdowndizzyvr · 9d ago
What runs Qwen 3.8 Flash Next faster? Strix Halo or a Pile of GPUs(2x5070tis, 2x7900xtxes and 2x5060tis 16GB).

I have a machine with a bunch of GPUs attached to it. 2x5070tis, 2x7900xtxes and 2x5060tis 16GB. So I did this little test to see how it fares running Qwen 3.8 Flash Next Q4_XL against my little Strix Halo. Not well. Not well at all. The full numbers are below but the high context number sums it up.

@160,000 context

Pile of GPUs 215.73(PP) and 16.29(TG)

Strix Halo(Gufo) 1227.12(PP) and 22.04(TG)

Here's the number for a Strix Halo fork of llama.cpp, Halo Box.

Strix Halo(Halo Box) 587.99(PP) and 21.24(TG)

Lastly, here's the mainline llama.cpp number.

Strix Halo(llama.cpp 0.4.1) 113.48(PP) and 7.06(TG)

For running QFN, Strix Halo really shines.

Pile of GPUs

Device 0: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0, VMM: yes, VRAM: 15880 MiB
Device 1: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0, VMM: yes, VRAM: 15880 MiB
Device 2: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
Device 3: NVIDIA GeForce RTX 5060 Ti, compute capability 12.0, VMM: yes, VRAM: 15888 MiB
ggml_cuda_init: found 2 ROCm devices (Total VRAM: 49120 MiB):
Device 0: AMD Radeon RX 7900 XTX, gfx1100 (0x1100), VMM: no, Wave Size: 32, VRAM: 24560 MiB
Device 1: AMD Radeon RX 7900 XTX, gfx1100 (0x1100), VMM: no, Wave Size: 32, VRAM: 24560 MiB
| model | size | params | backend | ngl | fa | dev | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ------------ | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 | 242.13 ± 1.39 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 | 32.09 ± 0.06 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d10000 | 242.96 ± 1.12 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d10000 | 30.23 ± 0.14 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d20000 | 249.34 ± 0.65 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d20000 | 28.84 ± 0.05 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d40000 | 253.95 ± 1.44 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d40000 | 26.29 ± 0.07 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d80000 | 252.48 ± 1.20 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d80000 | 22.34 ± 0.07 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | pp2048 @ d160000 | 215.73 ± 0.56 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm | -1 | 1 | CUDA0/CUDA1/ROCm0/ROCm1/CUDA2/CUDA3 | none | tg128 @ d160000 | 16.29 ± 0.03 |

Strix Halo running Gufo

| model | size | backend | test | t/s |
| -------------------------------- | ---------- | ---------- | ------------------ | --------------------- |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 | 1603.47 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 | 26.53 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d10000 | 1377.97 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d10000 | 25.44 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d20000 | 1353.19 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d20000 | 25.01 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d40000 | 1328.64 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d40000 | 24.08 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d80000 | 1283.77 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d80000 | 23.00 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | pp2048 @ d160000 | 1227.12 ± 0.00 |
| Qwen3.8 Flash Next | 103.69 GiB | ROCm (HIP) | tg128 @ d160000 | 22.04 ± 0.00 |

Strix Halo running Halo Box

Device 0: AMD Radeon 8060S Graphics, gfx1151 (0x1151), VMM: no, Wave Size: 32, VRAM: 128000 MiB
| model | size | params | backend | ngl | fa | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 | 811.79 ± 19.09 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 | 23.95 ± 0.03 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d10000 | 729.60 ± 38.99 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d10000 | 22.46 ± 0.43 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d20000 | 718.14 ± 33.92 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d20000 | 22.45 ± 0.18 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d40000 | 700.55 ± 33.17 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d40000 | 22.29 ± 0.23 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d80000 | 646.85 ± 29.06 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d80000 | 21.95 ± 0.25 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | pp2048 @ d160000 | 587.99 ± 30.10 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | ROCm | -1 | 1 | none | tg128 @ d160000 | 21.24 ± 0.31 |

Strix Halo running llama.cpp 0.4.1

Device 0: AMD Radeon 8060S Graphics, gfx1151 (0x1151), VMM: no, Wave Size: 32, VRAM: 128000 MiB
| model | size | params | backend | ngl | fa | lm | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | --: | ---------: | --------------: | -------------------: |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 | 362.84 ± 6.19 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 | 20.11 ± 0.36 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d10000 | 321.75 ± 0.91 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d10000 | 19.88 ± 0.34 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d20000 | 285.78 ± 0.42 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d20000 | 18.35 ± 0.46 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d40000 | 237.23 ± 1.13 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d40000 | 15.48 ± 0.61 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d80000 | 173.80 ± 0.18 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d80000 | 11.18 ± 0.11 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | pp2048 @ d160000 | 113.48 ± 0.23 |
| qwen4exp A3B Q4_K - Medium | 103.68 GiB | 176.94 B | CUDA,ROCm,Vulkan | -1 | 1 | none | tg128 @ d160000 | 7.06 ± 0.17 |

💬 52 (+1) open on reddit ↗
▲
23
+1
32👁
r/LocalLLaMA · u/darklordfireape · 9d ago
Update: Strix Halo + R9700 with llama-halo-hybrid - now beats DGX Spark

Hi folks, I've spent the last couple of months experimenting with Strix Halo and previously I released a proof of concept I called llama-halo-hybrid. I've continued updating it and it now performs very well. The idea is that you can take an R9700, or similar, and place dense parts of the model, KV, and some of the layers on the GPU and let the APU take the rest of the model. You can add the extra GPU through a PCIe extender (framework desktop), Occulink, or a thunderbolt dock depending on which machine you have. Detailed notes along with code in the repo on github. I'm not selling anything, this is all 100% open, MIT-licensed.

It breaks 60+ tok/s decode and 2000+ tok/s prefill, supporting full 256k context.

This is not some custom inference engine that requires a custom quant to run. This is llama.cpp modified to run whatever you want, albeit mostly tuned for Qwen and GLM families. After continuing to tinker with it, it now performs better than DGX Spark (albeit cheaper) running Qwen-3.8-flash-next and slightly better yet with the Swift-1.5 variant. Most of my testing was done with the Q4/Q4\_K\_XL models to balance size and quality.

Note \- if you are just using Strix Halo by itself, this is probably not the right tool. Check out gufo, which looks very promising.

https://github.com/sixvolts/llama-halo-hybrid

I would love any feedback you all have and happy to investigate tuning for different "sidecar" GPUs other than the R9700 if there's demand and I can get my hands on one.

UPDATE (10/4): I added some more docs around using different cards besides the R9700. The R9070, V620, and 7800XT all perform very well and I put quick guides for those configs, along with notes on thunderbolt/USB4 setup and the dual-machine setup I used here:
https://github.com/sixvolts/llama-halo-hybrid/tree/main/halo-cookbook

💬 27 (+9) open on reddit ↗
▲
1
+1
8👁
r/LocalLLaMA · u/Inevitable-Log5414 · 9d ago
stuntd 0.1.2: local heads for multi-field decisions, and why one weak field decides how often you skip the model

A week ago I posted stuntd here, a proxy that learns your LLM's typed decisions and answers the confident ones with a small local head (~20ms GPU, ~60ms CPU). Thanks for the feedback last time :)

0.1.2 is out, the main thing is decisions with several fields, like category + urgency + needs_human.

First idea was to answer each field locally when its head is sure and ask the model for the rest. Dropped it, you pay for the whole model call anyway and a half local half model answer is a pain to debug. So it's all or nothing now: local only when every field is sure, otherwise the model answers and every field becomes training data.

Didn't expect how much that costs. On the support demo the heads alone are sure on 99.9%, 92% and 76% of tickets, but all three at once only on 72.7%, so the weakest field decides.

It also retrains itself now. auto_retrain kicks in after N new captures, the new head sits in shadow next to the model, goes live when it agrees long enough and back to shadow if it starts losing. Anthropic Messages learns too, and there's serve --lazy.

Code: https://github.com/bladedevoff/stuntd
Try it: https://huggingface.co/spaces/pollix/stuntd

Anyone else doing multi-field outputs locally, is it one weak field for you too?

▲
4
+1
9👁
r/LocalLLaMA · u/HornyGooner4402 · 9d ago
Pi + llama-server randomly hung

I can't seem to find what's wrong. I'm using Pi for my llama-server and sometimes it just stops processing for some reason and stuck after tool call. Logs seems to think that it's finished its job while Pi thinks it's waiting for a response, so sometimes I have to stop it and tell it to "Continue". This only happens occasionally, 99% of the time it works with no problem. Anyone experienced something like this?

Edit: Just realized I was vagueposting. Running Qwen3.6 35B A3B IQ4_NL_XL from Unsloth, but I think it happened with other models as well.

▲
2
+1
6👁
r/LocalLLaMA · u/poofph · 10d ago
infill and output tok/s speeds after "new" build compared to old questions

Let me start off by saying I am new to AI and have a lot to learn, basically I don't know shit. I started off a few weeks ago by throwing my 2 5090s I had from gaming pcs into a 9950x cpu with 64gb ddr5 6000 ram on a motherboard that was able to do gen 5 8x per card system. Running ubuntu 24.04 server and running unsloth studio, swift 1.5 qwen 3.8 27B Q8 with kv cache dtype at q8\_0 and 262k context I was getting 2500-3000 infill and 100-150 toks/s output.

I wanted the 5090s in my rack in the basement in my proxmox server, it has a 7402p cpu (24 core 48 thread (rome)). 256 gb ddr4 3200 ECC ram (8 channel) on a supermicro H12SSLNTO motherboard. I have the 5090s passed through (gen 4 16x each card) to a vm (using 128gb of ram, direct access, no ballooning etc) and a dedicated 1.8 tb nvme drive passed through dedicated for the ai server vm (actually the vm itself is using a pool on the proxmox server but all the ai stuff is sitting on and running from the 1.8tb nvme).

Everything is working okay. It is running ubuntu 26.04 server. I have unsloth studio running, running the same model and settings, infill is more, up to 3800 but output is like half or less around 60 tok/s. Ideas what may be causing the drop in tok/s output and what to look into if a system issue?

I have done a lot of memory bandwidth tests (theoretical is \~204 GB/s, double that of ddr5 dual channel) but from what I can find and because I only have a 4 ccd cpu I am only getting 90-120 GB/s memory bandwidth. I guess I can get that to the 160-180 range if I go with a 64 core 8 ccd cpu, which I am considering doing..but I don't even know if that has anything to do with anything, just a rabbit hole I went down.

Ideas what to look into for the drop in output tok/s?

▲
11
+1
15👁
r/LocalLLaMA · u/jjusko20 · 10d ago
SFTMill: Easily [off-policy] distill any existing LLM with an OpenAI Compatible Endpoint. Turn any behavioral goal into a comprehensive dataset. post image

Disclaimer: Any\* means any model that exposes its CoT without it being censored.

Hey guys - half a tutorial/guide, and half an I built this, so I went for resources. This is something that I created for myself recently when I couldn't find any good existing solution. I wrote this post myself, no AI!

Probably a fair number of you have seen my posts about fine-tuning AliceAI 80B A3B according to my own synthetic datasets. If you did, I'm still fine tuning it on a live stream right now - check out https://figure-bios-expect-cio.trycloudflare.com/ \-- it'll let you inspect any and all of the training data that I generated with this engine. If you have any interest, it's pretty neat! unfortunately that link is optimized for desktop only and I'd have to kill the run to reset it, so u may want to rotate the phone.

That thread was at https://www.reddit.com/r/LocalLLaMA/comments/1wslgkw/comment/pcw4kgw/?context=1&screen\_view\_count=1

That run is using off-policy distillation, and that's I made this for. my training data for that project with this repo, and just customized it for an OSS release. Basically, you create a "curriculum" for your goal - e.g. if I was training an agentic model, I'd need things like tool calls, bug fixing, working in a workspace, tracing errors, etc. You define your curriculum in a yaml file, then an LLM creates tasks based on the curriculum you defined, and the chosen LLM you're distilling from then solves each task, leaving you with a full Q/A set that encompasses your fine tune goals.

I used qwen 3.8 27b on medium to generate the tasks - I'd recommend avoiding anything any weaker than that.

I forked my private repo of this that I've been using into SFTMill, which is basically just the same thing with great documentation and a few steps added to get anyone onboarded rapidly. I created it \[and my original version\] because I couldn't find any existing pieces of software made with this design, for this purpose.

I release it because I enjoy contributing to the community, and there's a vague hope someone will eventually see one of my pieces of work and want to hire me (if you're reading this and you like the project and you need a software/ml engineer remote or in NYC, let me know <3). It makes me happy when my software helps others so I'd love if you let me know if it helped you. Cheers!

Shoutout u/FullOf_Bad_Ideas for helping me with my alice train in areas I wasn't experienced enough in - I threw a Multi-Turn Hybrid-Reasoning (user <> assistant) section in the readme just for you bud, hope it helps.

https://github.com/jackjusko/sftmill

▲
6
+1
12👁
r/LocalLLaMA · u/Gold-Bat-3225 · 10d ago
Does post training make LLMs funnier? post image

We did a study: does post training actually make LLMs funnier?

We used open models that publish every stage of post training, so we could compare a base model with the future models it became: Tulu 3 (on Llama 3.1 70B), OLMo 3.1 32B and Qwen2.5. We tracked 11 stages, 100 joke prompts, 64 human raters and 2,330 head-to-head judgments.

What we found: post training makes models funnier, but reduces diversity of response.

\- In 5 of 7 training steps, the later model's jokes were judged funnier. Jokes also got 10–20 words shorter after early post training, so they get to the punchline faster.

\- In 6 of 7 steps, the jokes a model wrote for the same prompt got more similar to each other. Ask for eight jokes on one premise and you get eight versions of the same joke. The biggest drop was Qwen2.5 base to instruct.

\- Asking the model to plan a line or two before the joke cut variety in all 4 models we tried, with no reliable gain in funniness.

\- A comedian persona won back a little variety in all 4 models, but only made the jokes funnier in 2 of them.

Humans judged the base versus final. A model judge calibrated on those votes compares the stages in between.

Full report and paper below. Which open models should we run through this next?

https://laugh.so/research/humor-tax/

💬 12 (+1) open on reddit ↗
▲
11
+1
10👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 10d ago
What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects?

For example, directly using claude code or code, which is then hooked up to automatically delegate the actual code writing tasks to a local model like qwen 3.8 flash next, to save on cloud usage limits.

I’m imagining the loop would be:
User writes prompt
Claude/codex thinks about it and the plan
Claude/codex sends the specific and bounded coding instructions to the local model+harness (opencode, pi, etc) via api endpoint or MCP, with clear instructions on a defined endpoint
One the local model+harness hits the clear endpoint/“done” step, it sends a ping back to claude/codex
Claude/codex then verifies the output and then thinks about next steps to instruct the local model+harness on

Does this actually lead to improved savings on the cloud model usage while preserving code quality? Or does this end up being unnecessarily complex and not saving on any cloud usage

💬 21 (-1) open on reddit ↗
▲
8
+1
8👁
r/LocalLLaMA · u/DeliciousBelt9520 · 10d ago
Forlinx 20-TOPS M.2 AI accelerator supports PCIe cascading for local LLM inference

Forlinx Embedded has listed an M.2 AI accelerator card based on Rockchip’s RK1820 and RK1828 processors, providing 20 TOPS of INT8 computing performance and up to 5GB of integrated DRAM. The module uses an M.2 2280 interface and is designed to handle local AI inference, including large language models, vision-language models, and computer vision workloads on embedded Linux and Android systems.

https://linuxgizmos.com/forlinx-20-tops-m-2-ai-accelerator-supports-pcie-cascading-for-local-llm-inference/

▲
10
+1
16👁
r/LocalLLaMA · u/dh7net · 10d ago
Distributed Local Agents Benchmark.

I wanted a setup where I can compare all the harnesses with all the local models.

It turned out to be a rabbit hole. For instance, you would not only have to test all the harnesses (being sure that they are well configured), but also all models, with all their flavors, and this for all kinds of hardware.

Everyone can do their share, but no one can pretend to do all possible tests extensively.

To solve this, I created a website where everyone can test the configurations they want and share the results if they want. You can try it here: airbench.ai

There is a leaderboard where I share the tests I'm making, but I hope I can populate it with tests from others. https://airbench.ai/leaderboard?k=poL

My hope is to turn this into a fully distributed Agent Benchmark.

Let me know what you think.

▲
12
+1
9👁
r/LocalLLaMA · u/ButtercupLyn100 · 10d ago
I’m building an open-source browser agent that can run locally with LM Studio/Ollama — including a 450M browser VLM

I’ve been working on an open-source project called WebBrain that gives LLMs the ability to see and operate a browser.

One thing I really wanted to avoid was making the browser agent dependent on a single cloud model/provider.

So WebBrain can work with local models through things like LM Studio and Ollama, as well as cloud APIs if you want them.

I also ended up training a small vision model specifically for browser tasks:
webbrain-vl-2-450M

It’s based on LFM-2.5-VL-450M and fine-tuned on browser screenshots/tasks. The idea is that instead of sending every screenshot to a giant multimodal model, some browser perception can happen with a very small model locally.

It can run through WebGPU directly on the user's machine.

The agent itself combines screenshots with the browser accessibility tree rather than relying entirely on DOM parsing.

Current architecture is roughly:
• screenshot + accessibility tree for perception
• browser-specialized tiny VLM where useful
• model-agnostic planner
• local models via LM Studio/Ollama
• Chrome / Edge / Firefox / Chromium support
• optional cloud execution
• open source

I'm especially interested in figuring out how far browser agents can realistically go with small local models rather than GPT/Claude-scale models.

Repo: https://github.com/webbrain-one/webbrain
Model: https://huggingface.co/webbrain-one/webbrain-vl-2-450M
Dataset: https://huggingface.co/datasets/webbrain-one/webbrain-vl-2-450M-dataset

Would be very interested in feedback from people here running smaller Qwen/LFM/MiniCPM/etc. models locally — particularly what model you would try as the planner.

▲
110
 
40👁
r/LocalLLaMA · u/rmonsurate · 9d ago
Two open-weights releases: Victoria (Qwen3.8-Flash-Next with 44% of experts cut, 70% Terminal-Bench 2.1, GGUF included) and Maple (a Canada-first fine-tune)

We had a Dell B300 in the lab for a few weeks and used it to create two fine tunes of Qwen Flash Next.

Victoria (coding and agents)

  • Qwen3.8-Flash-Next cut down by 44% using a paper / technique called REAP: 512 down to 288 per layer.
  • Retrained at 4-bit (NVFP4) afterwards, so it's trained for the format it ships in rather than just quantized after the fact.
  • Terminal-Bench 2.1: 70.0%, averaged over 3 runs with an 8h per-task timeout. Our previous NVFP4 build scored 62.5%.
  • HumanEval: 159/164.
  • 48.0 GiB of weights, including the draft head. The 95.4 GiB n-gram table is separate and not counted in that number.
  • 280 tok/s single stream on one B300 with the draft head, versus 135 without it.
  • GGUF Q4_K_M is 49.17 GiB. It scored 75.3% on Terminal-Bench (a single run, so treat it as noisy) and 93.2% on HumanEval (averaged over 5 runs).
  • Uses 35% fewer output tokens than our previous build.

Maple (Canadian questions)

Most models answer questions about taxes, benefits and regulations as if you live in the US. Maple is fine-tuned to default to Canada. On 600 held-out questions, with search:

  • Cites an official Canadian source: 6.0% before fine-tuning, 62.9% after.
  • Fully correct answers: 6.6% before, 21.8% after.
  • "No answer" responses: 47.2% before, 23.7% after.
  • It pushes Canada onto people who said they live somewhere else less often: 2.9% before, 1.0% after.

Coding holds up: 157/164 on HumanEval. Grading was done by an AI judge panel; human review hasn't happened yet.

Links:
https://huggingface.co/rmonsurate/Victoria
https://huggingface.co/rmonsurate/Maple

Happy to answer questions about running them.

Edit: llama.cpp users. The GGUF carries our draft head, and mainline llama.cpp doesn't know about it yet, so it fails with "expected 1256, got 1224". Your download is fine. For now, build from our fork: github.com/rmonsurate/llama.cpp, branch qwen4exp-mtp. Prebuilt binaries are on the way. Thanks to the reader who caught this.

💬 42 (+3) open on reddit ↗
▲
57
 
34👁
r/LocalLLaMA · u/jwestra · 9d ago
1248 GB/s on 5060ti (+40%) with +5500 memory overclocks

edit: sorry should be 5070ti

There now is unlock for higher memory overclocks called mlock. And apparently the GDDR7 has a lot of headroom:
https://www.reddit.com/r/overclocking/comments/1wsnllh/finally\_unlocked\_gddr7\_memory\_overclocking\_with/
Of course this can help massively for local inference, especially token generation.

💬 57 (+1) open on reddit ↗
▲
0
 
20👁
r/LocalLLaMA · u/EmilPi · 9d ago
I don't understand whether uncesored/abliterated/heretic/fusion/bla-bla models give any value for the open-weight community

Using uncensored model gives sort of sense of power, I suppose, for some people; but what else?

(UPD.: Usecases well-explained in the comments: cybersecurity, storywriting, law, medical, criminal forensics).

If I am wrong, prove me wrong, please, or just say you really need something different from it. I sure felt a frustration when (just one of the ton of examples) e.g. you ask how did Peter Pettigrew die, and the model suddenly starts a litany it is a harmless assistant.

The GLM-5 helped HF against OpenAI cyberattack without being uncensored. If you need an uncensored model to understand political hypocrisy, well, you haven't grown up yet. You want to protect your property against a burglar? Find a competent consultant, instead of potentially hallucinated advice from the LLM (and the more uncensored the model, the more it hallucinates).

The only measurable goal I see the uncensored models serve now is a pretext for the corps to regulate people, who are just happy having Gemma4.x/Qwen3.x/DeepSeek-4.x/GLM-5.x do some stuff for them. Not sure that 5% of legitimate use cases (which I believe exist, but are only substitutes for a classic search or consultation) are worth it. What if I (and I believe a majority of the open-weight models' users) don't need waifu/goon/bioweapons or meth recipes/propaganda generation/cyberattacking/scamming capabilities?

💬 79 (+3) open on reddit ↗
▲
0
 
20👁
r/LocalLLaMA · u/ag789 · 9d ago
CopilotKit

The 'AI' world is moving plenty fast, enter CopilotKit
https://github.com/CopilotKit/CopilotKit#what-you-can-build
'agents' are coming in, draw charts, type your document, spreadsheet, operate your web browser, write your email, make presentations.
It would probably leap off the screen into the physical world

It is probably a 5yo's definition of 'AI' , that's coming true

The 'agent loop' becomes practically, all apps, all frontends (webui, gui, mobile) everything anything , anything connected to an LLM.

I think Local LLM would be part of that after all.

▲
9
 
13👁
r/LocalLLaMA · u/nonlinearsystems · 9d ago
M5 Ultra - Qwen3.8 Flash Next vs Laguna S 2.1 post image

Spent today running a same-day, same-harness shootout between Qwen3.8-Flash-Next (oMLX, 182GB oQ8e, MTP) and Laguna-S-2.1 GGUF (LM Studio, 128GB, 8bit) on a Mac Studio M5 Ultra 256GB. Both capped at 262K context, thinking on, unique content per run with zero cached tokens verified each time.

That last part matters because my first run was wrong hah... shared prefixes across sizes let the KV cache carry over and 200K "prefilled" in 21s.

Prompt Qwen Laguna
8K 2.0s 10.3s
32K 7.4s 30.4s
64K 14.7s 70.8s
131K 30.1s 217.2s
200K 47.1s 455.4s

Qwen holds \~4,200 tok/s linear which is amazing. Laguna degrades superlinearly (quadratic attention doing quadratic attention things). At 200K, prefill is 94% of total time on both.

Decode (tok/s): Qwen 59-74 across sizes (MTP at 70-76% acceptance per server logs, roughly 2x). Laguna 68 down to 34 as context grows. No speculation on Laguna, its DFlash path already lost to plain decode on this hardware in earlier testing. I think if Laguna could get DFlash figured out or MTP, this might be a different conversation.

Quality was a draw, 4/4 each, on four problems with script-verified answers (Muse created the gymnastics here: exact 9-digit combinatorics, interval code with 12 hidden tests, fresh knights/knaves, asyncio ordering trap). Opposite styles though: Laguna answers in 5-10s with a few hundred tokens, Qwen deliberates exhaustively (one answer took 119s / 11K tokens). Both burned a full 8K budget on hidden reasoning with zero visible output exactly once, then converted on a 16K retry.

Happy to answer methodology questions. Full writeup with charts and the test rig diagram: https://echalupa.com/blog/qwen-flash-next-vs-laguna-200k

▲
0
 
12👁
r/LocalLLaMA · u/artur_oliver · 9d ago
600M parameter model for transcription, super reliable.

Hello community,

I have been thinking of building an app for the company that just gets the calls from the automated answering machine to text, but I have huge problems with the quality of the translation. That's why I think I can use this model.

My idea is to have a summary table every 20-30 calls about the content or important recalls I need to do.

I run a really busy office, we get about 100 cals a day if not more.

I want to get that but the devils are in the details, what do you think?

What features should be implemented first or even complementary to it?

Thanks

▲
0
 
8👁
r/LocalLLaMA · u/power97992 · 9d ago
Next year, the pro models will have 8-10 T parameters, who will have enough vram to run them?

Deepseek said they will release an 8 T model later and qwen said they will have a 10 T model and kimi will probably follow suit. The flash models will probably be around 1 -2 T parameters. Then only companies and corporations And cloud providers and rich people will be able to afford to run these pro models and fairly rich people for the flash models . At this rate, you would need 9 512 gb m5 ultras or 48 rtx 6000 pros to run A 4.4 bit 8T model with full context ? That is probably 153k for the ultras or 768k for the rtx pro Gpus plus probably another 100k for the other parts. I guess either use the cloud or people will use smaller models like qwen 5 27b in the future but most people won‘t be able To run the biggest models locally. In fact, most people will struggle to run a 4.4 bit 1 t flash model locally. It will cost 100-120usd/h just to host the mod in the cloud

▲
8
 
14👁
r/LocalLLaMA · u/thatscoolbutno123 · 9d ago
48GB VRAM + 64GB RAM, anything worth trying except Q38 fn/27b?

Basically the title.
Just got myself a R9700(32GB) additionally to my rx7800 (16GB) and im already experimenting with qwen3.8 fn and 27b, but im interested wheter they are any other models compatible and comparable.
I want to use it with hermes agent mainly.

💬 25 (+3) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/ag789 · 9d ago
The Agent loop is probably what matters (for local LLM)

The commercial ones seemed to want to monopolize the agent loop.

Today the chat completions API is probably a 'defacto' way of talking to the models

https://github.com/ggml-org/llama.cpp/tree/master/tools/server#post-v1completions-openai-compatible-completions-api
https://vercel.com/docs/ai-gateway/sdks-and-apis/openai-chat-completions
btw, credit goes to the origin:
https://developers.openai.com/api/docs/guides/completions

A thing is, more recent efforts seem to be instead offering just an \*agent\* at the API and putting this \*agent\* layer between you and the model.

local LLM will remain \*very\* important because as is currently, you own the agent loop.
You write that "small little" front / stub that is the agent loop talking to the LLM.
it is day and night difference , practically 2 different universes

▲
22
 
24👁
r/LocalLLaMA · u/StartupTim · 9d ago
What's going on with DGX Spark? Price up $2k in 1 week?

I need to buy 2x DGX Spark but can't find a seller. Any hope or idea where I can get 2?

The price seems to have soared. Local Microcenter had 25+ then 0 the next day. They pulling stock?

💬 73 (+1) open on reddit ↗
▲
15
 
8👁
r/LocalLLaMA · u/parepeg · 9d ago
Gliner2.5-Decide (Jev style model)

I was perusing huggingface trending and was surprised that this hadn't been posted to localllama. Looks like it's been out for about a week.

▲
16
 
18👁
r/LocalLLaMA · u/SignificantZebra5883 · 9d ago
i would like to learn deeply about fine-tuning local models before burning money

There's so many new techniques like RL, RL LoRA, QLoRA, CPT LoRA.

I believe i would have a usecase for them, but i don't know where to learn, youtube is filled with bad quality tutorials if i just search and the good channels (fireship, bycloud) don't cover these as they're quite new concepts, i guess?.

how can a regular joe like me learn about these concepts in a "practical depth" so i can actually fine-tune qwen 27b successfuly on lets say custom corpus? without spending 100$ figuring out that "oh i didnt even need CPT here" or "well i chose the wrong Rank count! time to start this 2 day run again!"

context and TLDR: im building a legal general purpose chatbot for context, i have a big corpus, but im a bit stuck on what to do next

thanks for reading and any pointers!

▲
3
 
21👁
r/LocalLLaMA · u/bjivanovich · 9d ago
[Release & Deep Dive] ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP (i1-Q5_K_M): Sustaining 50-65+ t/s Across a FULL 128k (131,072) Context on a Single 24GB RTX 3090

ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP (GGUF) High-Precision i1-Q5\_K\_M with True 131k Context on Consumer 24GB GPUs

Most benchmarks in the community measure generation speed at trivial context depths (2k to 8k tokens). However, running a 27B parameter model at high quantization precision (Q5\_K\_M) across 131,072 tokens (128k) on a single consumer 24GB GPU without overflowing into slow system RAM or sacrificing attention fidelity is a fundamentally different challenge.

I am releasing

ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP, an optimized quantization suite built with a dedicated calibration imatrix, custom asymmetric tensor mapping, and native llamAmpere hardware acceleration.

Hugging Face Model Card: https://huggingface.co/bjivanovich/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP-GGUF

Available Quants: i1-Q8\_0, i1-Q6\_K, i1-Q5\_K\_M (Primary), i1-Q4\_K\_M, plus mmproj-BF16.gguf for multimodal vision.

  1. The 131k Context & Q5 Precision Challenge on 24GB VRAM

On a standard 24GB card (RTX 3090 / 4090):

  1. Weight Footprint: A standard 27B model at Q5\_K\_M occupies \~19.2 GB of raw weights.
  1. Context Memory at 131,072 Tokens: Standard FP16 KV cache for 131k tokens requires >24 GB on its own, making full-context inference impossible without dropping precision down to severe Q3/Q2 compromises or offloading layers to CPU RAM.
  1. MTP Quantization Pitfall: Standard community quants compress the Multi-Token Prediction draft block (blk.64) uniformly. At Q5 or Q4, this degrades draft accuracy, causing speculative acceptance to plunge from 85% down to \~55%, destroying generation speed.
  1. Our Architecture: Asymmetric Tensor Mapping + llamAmpere KV Compression

To solve this, we applied an asymmetric layer-by-layer quantization layout calibrated on a custom domain-rich dataset (imatrix\_atx\_uncensored.dat):

MTP Speculative Head (blk.64) Isolated at Q8\_0: Guarantees near-lossless draft predictions, increasing acceptance rates to 76% - 88% (averaging 3.3 to 3.7 verified tokens per generation round).

Attention Layers (attn\_q, attn\_k, attn\_v, attn\_output) Protected at Q6\_K / Q8\_0: Prevents attention drift and catastrophic reasoning decay at 64k, 96k, and 128k+ token horizons.

FFN Layers (ffn\_gate, ffn\_up, ffn\_down) at Q5\_K\_M: Absorbs standard compression without degrading semantic coherence.

Unified Turbo KV Cache (-ctk turbo5 -ctv turbo4 or -ctk q8\_0 -ctv turbo3): Compresses the 131,072 KV cache down to just \~3.5 to 4.2 GB of VRAM, allowing the entire Q5 model + full 131k context window to reside 100% inside the 24GB VRAM envelope.

  1. GPU Memory Footprint & Resource Breakdown (RTX 3090 24GB)

Total VRAM Allocated: 23.4 GB / 24.0 GB (100% GPU offload, -ngl 99, 0 layers in CPU RAM).

Model Weights (Q5\_K\_M Asymmetric): \~19.2 GB.

KV Cache (131,072 tokens, Unified Turbo4/5): \~3.8 GB.

System RAM Cache (--cache-ram 4096): 4.0 GB RAM dedicated to multi-session prompt state preservation.

CUDA Compute Architecture: Ampere SM86 with FlashAttention-2 (-fa on) and hardware Tensor Core MMA fused kernels.

Direct Benchmark Comparison: Standard Swift-1.5 Q5 vs ATX-Swift-1.5 Q5

Tested on Single NVIDIA RTX 3090 (24GB) with llamAmpere under Deep Context (\~80,000 to 98,000 active tokens)

| Measured Metric | Standard Swift-1.5 Q5 (mradermacher) | ATX-Swift-1.5 Q5 (Our Quant) | Real Delta |

| Sustained Speed (\~80k-98k ctx) | 44.94 to 45.79 t/s (Tasks 740, 19637) | 50.05 to 53.72 t/s (Tasks 0, 100, 155) | +5.5 to +8.0 t/s (+13% to +17%) |

| Burst Generation Peaks (tg\_3s) | 45.6 to 52.3 t/s | 58.10 to 63.05 t/s | +10.7 t/s higher peak bursts |

| MTP Draft Acceptance Rate | 63.7% to 65.1% (Tasks 740, 19637) | 75.2% to 81.7% (Tasks 0, 155) | +11.5% to +16.6% higher accuracy |

| Mean Draft Length (mean len) | 2.91 to 2.95 tokens / round | 3.31 to 4.10 tokens / round | Up to +1.1 tokens / verification step |

| Compute Time per Token | 21.84 to 22.25 ms / token | 18.62 to 19.45 ms / token | \~3 ms lower latency per token |

| KV Cache Precision Evaluated 1| -ctk q8\_0 -ctv turbo3 (3-bit V) | -ctk turbo5 -ctv turbo4 (4-bit V, higher precision) | ATX wins in speed despite higher KV fidelity |

Real Execution Log Excerpts

  1. Standard Swift-1.5 Q5 (Symmetric Quantization)

Task 740 (Context: 97,182 tokens | Generated: 1,130 tokens):

eval time = 25123.57 ms / 1130 tokens (22.25 ms per token, 44.94 tokens per second)

draft acceptance = 0.63746 (742 accepted / 1164 generated), mean len = 2.91

Task 19637 (Context: 92,075 tokens | Generated: 834 tokens):

eval time = 18192.26 ms / 834 tokens (21.84 ms per token, 45.79 tokens per second)

draft acceptance = 0.65130 (551 accepted / 846 generated), mean len = 2.95

  1. ATX-Swift-1.5 Q5 (Asymmetric Custom Tensor Mapping)

Task 155 (Context: 84,099 tokens | Generated: 3,478 tokens):

eval time = 67619.68 ms / 3478 tokens (19.45 ms per token, 51.42 tokens per second)

Burst Peaks: tg\_3s = 60.18 t/s and tg\_3s = 63.05 t/s

draft acceptance = 0.73096 (2543 accepted / 3479 generated), mean len = 3.72

Task 100 (Context: 97,847 tokens | Generated: 525 tokens):

eval time = 10175.17 ms / 525 tokens (19.42 ms per token, 51.50 tokens per second)

Burst Peak: tg\_3s = 58.10 t/s

draft acceptance = 0.76939 (367 accepted / 477 generated), mean len = 3.31

Task 0 (Context: 80,016 tokens | Generated: 456 tokens):

eval time = 8470.07 ms / 456 tokens (18.62 ms per token, 53.72 tokens per second)

draft acceptance = 0.81710 (344 accepted / 421 generated), mean len = 4.10

Technical Takeaway for the Post

  1. Why ATX is \~15% faster under identical deep context:

In standard quants, compressing the speculative head (blk.64) to Q5 causes \~36% of proposed draft tokens to fail rejection sampling, reducing throughput to \~45 t/s.

In ATX-Swift-1.5, isolating blk.64 at Q8\_0 increases draft accuracy from \~64% to \~77%+, delivering 3.31 to 3.72 verified tokens per round and raising sustained generation speed past 51.5 t/s (with burst peaks over 63 t/s).

  1. Optimized Execution Script (llamAmpere)

Make sure to pass the explicit MTP vocabulary shortlist (atx\_65536.txt). This restricts speculative draft projections to the top 65,536 power-of-two tokens, aligning perfectly with NVIDIA Ampere Tensor Cores and preventing a 73% compute penalty:

cd D:\\llamAmpere
$env:GGML\_Q8\_TURBO3\_MMA\_FUSED = "1"
.\\build-sm86\\bin\\Release\\llama-server.exe
\-m "ATX-Swift-1.5-Qwen3.8-27B-Uncensored-MTP-i1-Q5\_K\_M.gguf"
\-ngl 99
\-c 131072
\-b 2048
\-ub 512
\-t 20
\-tb 8
\-fa on
\-ctk turbo5
\-ctv turbo4
\--kv-unified
\--prio 3
\--parallel 1
\--jinja --fit off
\--cache-prompt
\--cache-ram 4096
\--spec-type draft-mtp
\--spec-draft-n-max 3
\--spec-draft-p-min 0.1
\--spec-draft-type-k q8\_0
\--spec-draft-type-v q8\_0
\--spec-draft-vocab-map "D:\\llamAmpere\\docs\\mtp-vocab\\atx\_65536.txt"
\--reasoning-format none
\--temp 0.2
\--top-p 0.90
\--top-k 40
\--min-p 0.05
\--repeat-penalty 1.08
\--repeat-last-n 256
\--alias "ATX-Swift-1.5-Qwen3.8-27B-Uncensored-Q5\_K\_M"
\--host 127.0.0.1
\--port 8080

▲
0
 
8👁
r/LocalLLaMA · u/akumaburn · 9d ago
CadetCoder Version 1 (Another Coding CLI - Pure Java Implementation)

Check if it out if you're interested.

Github: https://github.com/akumaburn/CadetCoder

License: Apache 2.0

https://i.redd.it/r2l11vracosh1.gif

Built with a re-engineered version of the SCHEMA harness that aced ARC-AGI-3 (read more here: https://schema-harness.github.io/ )

Feature additions/bug fixes and pull requests welcome.

▲
7
 
17👁
r/LocalLLaMA · u/PhysicsDisastrous462 · 9d ago
Follow-up: my native Rust + Vulkan Transformer training backend — 14 days later, now 14 parity-verified architectures and full PEFT

Follow-up to my post from about two weeks ago. A lot has changed since then, so I wanted to post an update on where the backend is now.

Where the green architectures stand

When I posted last time, 7 architectures had verified full training support. Everything is now held to the same strict harness: a pinned local Hugging Face Transformers source tree as the oracle, forward logits + gradients + two full AdamW steps compared, and every named parameter checked again after export.

The hard ceiling is 2e-7 absolute error. No loosening tolerances and no rounding numbers afterward to make the README look better.

14 architectures pass that gate today, led by the one I'm probably proudest of:

|Architecture|Scope|
|:-|:-|
|Falcon H1 / H1R|parallel GQA/RoPE attention + Mamba2 in every layer; full training, full fine-tuning, LoRA, saved modules|
|DeepSeek V4|causal LM|
|Phi-4 Multimodal|text backbone|
|Phi-3|causal LM|
|Kimi K2.5|text backbone|
|Kimi K3 / KimiLinear|hybrid KDA + MLA|
|GPT-OSS|causal LM incl. router bias|
|SmolLM3|mixed RoPE/NoPE + YaRN|
|Qwen2.5 / Qwen3.5 / Qwen4-Exp|dense, DeltaNet, QSA, PLE, MoE|
|Mistral 4, MiniMax M3, Gemma 3/4, MiniMax M2|causal LM|

Worst observed two-step AdamW parameter error across all of them: 1.19e-7.

Best: 2.6e-8.

For hardware context, all of the local Vulkan validation I've been reporting was run on my ASUS ROG Ally Z1 Extreme, using its AMD RDNA 3 integrated GPU. So the RDNA 3 results here are from that specific machine rather than testing across several different AMD systems.

The bigger news: PEFT actually works now

In the last post, "LoRA/PEFT-style fine-tuning" was basically one line in a feature list.

It's a real workflow now, and I've verified the full lifecycle:

  • LoRA fine-tuning with HF-compatible adapter export (adapter_config.json / adapter_model.safetensors), so adapters can round-trip with the PEFT ecosystem
  • modules_to_save — full trainable replacements for Linears, RMSNorm/LayerNorm, lm_head, and input embeddings, including named-adapter switching and bank isolation. Adapter A leaking into adapter B is explicitly tested for.
  • Exact resume — adapter weights + AdamW moments + step + dropout RNG state restore bit-identically against an uninterrupted run
  • Merge/unmerge, disable-adapter base restoration, and multi-adapter loading
  • A parameter-budget flag that automatically chooses the largest LoRA rank that fits within a requested percentage of the base model
  • The CLI fails closed if you try to use saved modules on an architecture that hasn't passed its corresponding gate

32 architecture surfaces across 20 families pass all three PEFT stages — LoRA, saved modules, and adapter switching — under the same 2e-7 gate, with frozen-base drift exactly 0.0.

The validation harness also fingerprints the pinned Transformers source alongside my shaders and binaries now, so a qualification run can't silently end up testing against different reference math.

A small note on the last couple weeks

I didn't get quite as many working days out of the last two weeks as I normally would have. Partway through this I got covid, then when that started to go away, it became a secondary nasty ear infection that ended up perforating my eardrum. I was running fevers around 104°F at one point and eventually went to the hospital, so I lost a few days to that and I'm on antibiotics now.

I'm doing better, though, and still managed to get most of what I wanted finished.

There are still things I want to clean up and expand, but I figured this was a good point to get the current work in front of people rather than holding the update back.

Same caveats as before

This is deterministic FP32 tiny-model correctness against a reference implementation, which should in theory ensure total mathematical parity for training and finetuning larger models with this backend, however, things are currently bound to FP32 training runs still, I eventually plan to work on MXFP4 weight tying to reduce memory footprints while fine-tuning (plastic parameters will still be trained in FP32 with PEFT in this config)

"supported text graph" ≠ "the entire multimodal package works natively."

Unsupported functionality is supposed to fail closed rather than silently falling back to an approximation.

Repo

https://github.com/necat101/Hierarchos-Native

  • Architecture inventory: hierarchos-vulkan/README_ARCHITECTURES.md
  • Compatibility/parity record: hierarchos-vulkan/COMPATIBILITY.md
  • PEFT qualification evidence: PROGRESS_PEFT_AUDIT.md
  • CLI PEFT guide: hierarchos-native-cli/README.md

The hardware I've personally validated this on is an ASUS ROG Ally Z1 Extreme with its AMD RDNA 3 GPU.

I'm very interested in criticism, compatibility reports, and especially results from people trying it on other hardware — NVIDIA, Intel, or other AMD GPUs.

I'd also love people to stress-test the PEFT resume/merge paths specifically. That's some of the newest code in the project, so it's probably the most useful area to try to break right now.

▲
1
 
8👁
r/LocalLLaMA · u/davidarias2 · 9d ago
Glassbench: an open-source workbench to compare local and hosted LLMs across AI trading agent frameworks

Glassbench is a free, open-source workbench that connects different AI trading agent frameworks, so you can watch how their agents decide, analyze every step and compare them.

AI trading agents are LLM systems where a team of agents (analysts, a bull and a bear, a trader, a risk team and a portfolio manager) research a stock, argue about it and give a rating. I wanted to watch how they reach that rating, so I started building a small interface for a popular open-source framework, TradingAgents. It grew into something much bigger.

Why I'm posting here: I run local models in this project, through Ollama, and I'm developing Glassbench into a benchmark pattern for AI trading agents. It connects different agent harnesses (TradingAgents and AI Hedge Fund so far) and runs them on the same stocks and dates, so the same setup can test and compare different LLMs, local and hosted.

What it does:

  • Live view: watch each agent work, with a timeline of every call, adapted for different frameworks
  • Runs database: every run stored and searchable, with its reports, costs and ratings
  • Framework and LLM comparison: the same stock and date on each framework, and on different LLM providers, you can also run it locally with Ollama. I'm evolving it to become a consolidated benchmark method
  • Backtests: the ratings tested against buy-and-hold and a placebo (still testing it, as nobody found a proper way to test TradingAgents)
  • Broker connection: a finished run becomes an order on an Interactive Brokers paper account

Frameworks plugged in: TradingAgents and AI Hedge Fund already run in it, unmodified, and more agent frameworks are coming. If you're building your own agent framework, you can plug it in through an adapter and compare it with the others on the same stocks and dates.

It's free and open source (Apache 2.0). My 77 runs ship with the repo, already paid for, so you can read everything the agents wrote without an API key.

Disclaimer: Glassbench itself is not an AI trading agent and makes no trading decisions. Every agent it runs comes from established open-source repos (TradingAgents and AI Hedge Fund), and Glassbench records what they do. Everything was tested on paper portfolios only; I have never traded real money with it. Research and education only, and nothing here is investment advice.

GitHub: https://github.com/davidalmeida90/glassbench

▲
0
 
15👁
r/LocalLLaMA · u/AdRepulsive7837 · 9d ago
Tensorfold runs Qwen3.8-27B really well on m5 pro mac mini, tps beats MTPLX

Came across this popular open source inference engine Tensorfold https://github.com/ashhart/TensorFold

Using their official Vontra/Qwen3.8-27B-MLX-4bit with drafting model z-lab/Qwen3.8-27B-DFlash2, I can reach 40-60 tps on mac mini m5 pro. AGAIN, it is PRO on mac mini, not even ultra studio.

For me, it is the first time (on mac ecosystem) that an inference engine to beat MTPLX. I have tested omlx, dflash2, mlx, llama-cpp, lm-studio, unsloth in the past few months, and none of them come close to MTPLX (running Qwen 3.8 optimised for speed, roughly 4bit?)

The more exciting part is that this enables me to seriously consider about replacing my RTX-3090ti with this mini running tensorfold as the main inference server setup. That old 3090ti, running ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF with MTP IQ3\_S (12.1 GB), reaches 50-70 tok/s, which is, in my opinion, similar to the 40-60 tok/s I achieve with mac mini. The only one caveat is htat the 3090ti still has like 4x faster prefill than mac mini.

Spec: M5 pro, Mac mini, 64gb, 1TB SSD

Testing harness: pi coding agent without any packages install yet.

Model: Qwen3.8-27B 4bit

What's your thoughts on Tensorfold?

💬 27 (+2) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/jaybsuave · 9d ago
Help choosing compute for a student? 4k budget

My university is going to give me 3k for a laptop and I was wondering what type of computer I should get? I already have a MacBook for school, and a desktop with a 4070 12gb and 64 gb. Any suggestions? I wanted a Mac mini but I can't use it ok Windows obviously and the DGX is too expensive. I can throw an extra 1000$ in as well if I need too so my budget is 4k. Thanks

▲
1
 
4👁
r/LocalLLaMA · u/Theboyscampus · 9d ago
Best practice for processing batch vLLM api calls with shared prefix?

Our agent workflow is currently executing a group of 10 vllm api calls within a asyncio.gather we made them share the same prompt until the end where the queries/instruction prompts differ. These calls are hitting our vllm-router/llm-d router with production grade kv cache aware routing algo which routes traffic into our pool of vllm workers. What's the best practice for processing batches of llm prompts with a shared prefix like this?

I have an idea where I try to see if I can make one call first to make sure vLLM complete a block of cache and start decoding before I send the remaining requests of the batch, our router will make sure these reach the same vllm worker, is this a good strategy?

💬 8 (+1) open on reddit ↗
▲
0
 
10👁
▲
21
 
14👁
r/LocalLLaMA · u/pilkyton · 10d ago
PSA: ModelScope CLI is now moved to "modelscope-hub"

To save people 30 minutes of research (because they didn't bother documenting this officially at all):

  • The "modelscope" package is now just the library. Doesn't contain a CLI anymore. If you try to install it or update your old CLI package, you get "No executables are provided by package \modelscope\; removing tool. error: Failed to install entrypoints for \modelscope\".
  • They moved all CLI tools to "modelscope-hub".

The new command to install it:

uv tool install "modelscope-hub"

▲
1
 
2👁
r/LocalLLaMA · u/OvertaxedOne · 10d ago
Good setup for QFN on 48GB Ampere GPU (A40)

Anyone have a good config they've found for QFN on a A40 (or similar Ampere GPU(s) with 48GB VRAM)? The system the card is in has 128GB of RAM (DDR3); right now it's running 27B but I'm curious if there's a way to move to QFN to maybe get better speed/a little more smarts. TY in advance!

💬 7 (+1) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/One_Temperature5983 · 10d ago
Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included

TypeSafe's Jev answers typed questions (yes/no, pick one label, pick a rubric level) with a probability per answer instead of text. Its docs say it takes text only: "Images, audio, and video are not supported (yet)." I wanted the same kind of answer about photos, from an open model on my own card, so I built typevet (MIT, Python 3.12).

How it works. No sampling, no parsing. typevet composes the native Gemma 4 turn itself, ends the prompt with the empty-thought no-thinking prefill, and reads the next-token distribution over the allowed answer tokens only. What comes back is a label plus a probability for every option, or an error. Images go to llama.cpp's /completion as base64 in prompt.multimodal_data with one media marker per image; on vLLM they go as image_url blocks. There's also a JSON path that returns an object passing your JSON Schema, or raises.

Local setup. Gemma 4 31B, my 24 GiB vramfit pack (byte-identical to the file on HF, projector sidecar for vision), llama.cpp b11223, one RTX 4090. Nothing leaves the machine.

The receipt test. 6 real receipts from CORD v2 (CC BY 4.0), 3 synthetic expense claims each: right total, two digits swapped, digits masked with ?. One three-label Choice: match / mismatch / insufficient_evidence. Each claim sent as text only, then with the receipt photo.

  • Right total, said match: 6/6 text only, 6/6 with the photo
  • Swapped total, said mismatch: 0/6 text only, 6/6 with the photo
  • Masked total, said insufficient: 6/6 text only, 6/6 with the photo

Example: claim says 646329, receipt says 664,329. Text only: match at 0.99964. With the photo: mismatch at 0.99999. Every swapped total was caught at 0.99998 or higher, and the masked ones abstained every time. The tests also check the image actually arrived: each photo added 228 to 1,108 prompt tokens here, and the gate fails if the count doesn't grow.

Hosted. Same code against vLLM 0.30.0, BF16 Gemma 4 31B, one H100: the receipt test went 18/18, and reversing the label order flipped 0 of 18 answers. Throughput on 480 Banking77 records with two questions each: 0.24 s median per record at 1 in flight, 39.6 records/s at 64 in flight, 0 errors.

Prior art, credit where due. The text-side decision model comes from TypeLLM (SGLang), which added its own image input on 9/24; typevet's image path is separate code on llama.cpp's request shape. allanrbo posted a Jev-like single script for Gemma 4 12B with webcam images on 9/25. VQAScore has read the probability of "Yes" from VLMs since 2024. typevet's angle: a library, not a script, the 31B on one 24 GiB card, the same code on vLLM, and image-arrival checks.

Scope: 18 claims, one run per server. The probabilities are the model's confidence, not calibrated.

▲
0
 
8👁
r/LocalLLaMA · u/HolidayBit143 · 10d ago
Local Q2_K model dunked on DeepSeek V4-Flash, a frontier AI, during my mini test & ngl I’m still processing this 😅😅

So I got this new local model on my system & wanted a mini test to see if it was actually smart or just confidently wrong (as I like to do with new models I haven't yet tried). I asked the cloud assistant to cook up a pretty rigorous 10 point diagnostic suite: reasoning traps, Python semantics, SQL fluency, strict instruction following, the whole gauntlet. At first it felt like DeepSeek V4-Flash frontier cloud intelligence vs my little local quantized guy. Classic quick test.

Then I ran it on the local model & shared the results back. The assistant was grading it & found out it got item #4 wrong. That item was a logic puzzle. The assistant thought one statement had to be false, but the local model was like nah, the set of statements is logically consistent, so your question is built on a false premise. It literally refused the leading question. I was like "wait wtf". That was the turning point fr. The local model solved a trap that the cloud model just completely keyed wrong.

The local model is a Q2\_K quantization of Nex-N2.5-mini, which is a fine-tuned Qwen3.5-MoE architecture. A 2-bit quant. Normally people call that low quality. But it outperformed a frontier model on a logic trap. The assistant went from “I am grader” to genuine admiration, saying resisting a leading question is high-level reasoning. Lowkey pretty humbling for the cloud side.

The whole thing made me think about emergent intelligence & AI democratization. Less giant centralized compute, more efficient specialized local stuff. The student corrected the teacher. Efficiency & MoE architecture maybe can beat raw parameter count sometimes. The mini test felt like a rite of passage for the local model. Its kinda like it became a validated thinker instead of just software. Q2 compression is also symbolic resilience, because despite being squished, the reasoning circuits stayed intact. And the assistant admitting it was wrong made the local model’s win feel more real.

And the craziest part? It went 10/10. This wasn't some easy benchmark either. It was a deliberately nasty little diagnostic with multiple ways for a heavily quantized model to screw up, and it didn't.

I need to make it clear that I am not claiming that Q2 ORCA model is generally superior to DeepSeek V4-Flash. But rather as an anecdotal demonstration that an extremely compressed local model can sometimes catch a reasoning failure in a frontier model & maintain much of it's reasoning power when done correctly & skillfully. It is a testament to how even under Q2 compression, it still preserved the model's “reasoning circuits."

Final verdict from the assistant: model is in excellent shape & ready for real work. So yeah, a local model dunked on the cloud AI. I’m happy for what this means for the future of local ai.

MODELS USED for quick test:

Local Model: Nex N2.5 Mini Uncensored

Frontier Model: DeepSeek V4.1

EDIT / CORRECTION bc I fucked this part up 😅

Small but important correction to the post. I originally called the frontier model I tested DeepSeek V4-Flash. That's not the right model name for the one I actually used on the DeepSeek website. It was DeepSeek V4.1-Flash.

Also, V4-Flash itself is a local/open-weight model, so my original wording made it sound like I was comparing my local model against some cloud-only AI. That's not accurate & that's on me.

The actual comparison was my local Q2\_K Nex-N2.5-mini vs DeepSeek V4.1-Flash through the DeepSeek website.

So yeah, V4.1-Flash is the model I should have named in the original post.

I'm leaving this correction here instead of quietly changing the post bc I don't wanna bullshit anybody or make it look like I didn't make the mistake. I got the model name wrong, someone pointed it out & I'm correcting it. 🤷‍♂️

The actual 10/10 result & the logic trap part of the test are unchanged.

\*\*TL;DR:\*\* I tested a local Q2\_K Nex-N2.5-mini on a 10-part mini test, it caught a logic trap the cloud assistant got wrong, & the assistant basically certified it as ready for real work. It got 10/10 correct.

▲
0
 
7👁
r/LocalLLaMA · u/Frosty-Whole-7752 · 10d ago
Just few days ago I've been badly censored even on this apparently different social network for criticizing the stance tech/social/digital/ai behemoths have regarding us, the user base some of them call/consider "dumb fuc*s". Well, I am bloody right!

That's why we have to fight against closed source centralized AI and closed recipes open weights overcoming the frivolous "gifts" exchanged with them by giving away our souls to those greedy entities if we want to be free in the future instead of being squeezed like lemons/at mercy/enslaved in the paws of these soulless folks that have a id of any single one of us at their disposal to switch us on/off at their leisure/convenience.

▲
0
 
8👁
r/LocalLLaMA · u/fuse1921 · 10d ago
[serious] roleplay

I was just wondering because I see it mentioned in threads here a lot... When people talk about LLMs used for roleplay, that's a euphemism for dirty/sexy chats right? Kind of like how "torrenting linux ISOs" is really just pirating copywritten media. Or are you guys really burning tokens pretending to talk to a medieval shopkeeper?

💬 86 (-1) open on reddit ↗
▲
23
 
15👁
r/LocalLLaMA · u/Defiant-Plantain1873 · 10d ago
Recommended replacements for glm 4.7 flash

I know I sound crazy, but i’m using a strix halo and finding that GLM 4.7 flash just runs significantly better than qwen 3.6 35 a3b. But its obviously quite old at this point, i wish we had a new glm that was 4.7 flash sized but is anyone using a model that they have found better than this.

My brief testing with qwen shows that glm is better at tool calling and better at world knowledge, but maybe there’s a chance my qwen set up is wrong

💬 31 (+2) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/Fit_Island928 · 10d ago
New to Local AI need help making a roleplay model

I'm making a local roleplaying model for my girlfriend's community server.

It's supposed to do roleplay, have a specific talking style(dry while answering to normal stuff and extensive when talking about lore), never talk out of roleplay, have hundreds of pages of lore and information and their rank in lore( discord roles maybe?).

It's basically supposed to be just an LLM you can converse with that answers in a specific talking style and has all the lore info.

For now I implemented: 10 ish% of the written lore, and it recognizes 3 people, but by discord ID that i inserted in the system prompt.

I'm using GPT 5.6 sol(and well 6 sol now) for doing stuff, but i keep running into a problem.

When i reach a nice point where the model has a nice talking style and knows information well enough, i tell Sol to add this new lorebook and this info, here now everything breaks.

Talking style is fucked, It doesen't recognise people individually anymore, when asked about other unrelated lore it just gets it wrong or hallucinates, or even starts roleplaying as one of the characters in it's lore book out of nowhere.

It's connected to discord through a discord bridge that Sol made and a developer dashboard bot.

I'm using GPT OSS20b on MXFP4, single 9070xt and 32gb ddr4.

Should I maybe fine tune it?

PS. im a beginner in AI so if it wasn't obvious i do NOT know what im doing but im trying my best for her.

▲
53
 
25👁
r/LocalLLaMA · u/sloptimizer · 10d ago
RAM Offloading with vLLM - tcclaviger appreciation post post image

Thanks to tcclaviger, vLLM now has expert RAM offloading support (link). This makes frontier models much more accessible on a local setup!

I was able to run the original DeepSeek-V4-Flash-Vision-Exp on four R9700s.

podman run --rm -it \
--init \
--network host \
--ulimit memlock=-1:-1 \
-v /models:/models:ro \
-v ~/.vllm-cache:/cache \
-e VLLM_ROCM_USE_AITER=0 \
-e ROCR_VISIBLE_DEVICES=0,1,2,3 \
-e VLLM_CACHE_ROOT=/cache/vllm \
-e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \
-e TRITON_CACHE_DIR=/cache/triton \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--annotation run.oci.keep_original_groups=1 \
--security-opt label=disable \
--security-opt seccomp=unconfined \
--shm-size 160g \
docker.io/tcclaviger/vllm@sha256:ef99b3d07c3f15e7978528c7510762ba024df9ab4242070d8ed092cd4cc1a694 \
/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
--served-model-name DeepSeek-V4-Flash-Vision-Exp \
--tensor-parallel-size 4 \
--enable-expert-offload \
--expert-offload-mem 160 \
--reasoning-parser deepseek_v4 \
--reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
--tool-call-parser deepseek_v4 \
--enable-auto-tool-choice \
--max-num-seqs 8 \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-batched-tokens 2048 \
--kv-cache-dtype fp8 \
--block-size 256 \
--max-model-len 256000 \
--gpu-memory-utilization 0.97 \
--mm-processor-cache-gb 4.0 \
--override-generation-config '{"max_tokens": 128000, "temperature": 1.0, "top_p": 0.95}' \
--speculative-config '{"method":"dspark","model":"/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp","num_speculative_tokens":3,"draft_sample_method":"probabilistic","enable_adaptive_verification":false}' \
--compilation-config '{"cudagraph_capture_sizes": [4,8,12,16], "max_cudagraph_capture_size": 16}' \
--host 0.0.0.0 \
--port 8090

▲
2
 
14👁
r/LocalLLaMA · u/circumcised_hobbit · 10d ago
llama.cpp cublas error, how to uninstall/reinstall properly (Linux Mint)

I am pretty dumb in this kinda stuff so please don't blame me for it.

My llama serve kept crashing with Cublas errors on first prompt with some models, but my VRAM usage was 3000MB/8k. I chatted with sonnet 5.5 for a bit and it told me it was a CUDA version issue (and made it work by not selecting any CUDA device)... I don't know if this makes sense, please tell me if it doesn't/what problem you think there is.

I realized that the best way was to delete llama.cpp (installed with curl install script) and install a clean CUDA 12 Ubuntu version.

\\- \*\*How do I properly uninstall llama.cpp (I don't wanna mess with Ollama files)?\*\*

\\- \*\*How do I install new version from .tar.gz archive without messing with system packages?\*\*

\\-Does my error diagnosis make sense to you? Would CUDA make generation actually faster? (I am getting 8tk/s with Qwen35B Q2 on 4060Ti 8GB due to no CUDA selected)

Edit: You guys saved me! Thanks! I had to install NVIDIA toolkit and switch to CUDA 13 llama.cpp tarball build

▲
19
 
18👁
r/LocalLLaMA · u/Roy3838 · 10d ago
Thanks to you r/LocalLLaMA, my mom was able to use my app! The open-source app that can watch your screen and trigger actions. It is now easy to use, thanks to your feedback.

TL;DR: I'm a solo dev who wanted a simple, private way to have local LLMs watch my screen and do simple logging/notifying. After a year of building, I released v3.0.0 and my mom was able to use it for the first time and I wanted to say thank you!

Hey r/LocalLLaMA,

What is it used for?

It is designed to monitor anything, some use cases:

  • When my Simulation crashes, call me.
  • When Concert tickets become available, click the buy button.
  • When my Steam game is downloaded, send me a Telegram.
  • When a Render is finished, send me an SMS.
  • When ... \[Anything happens\] Then ... \[Notify me, log it\]

How It Works

It's a micro-agent framework controlled by an MCP (Agent which I call Observer). So you type in Observer what you want monitored, and it'll control the framework to monitor it.

The desktop app uses llama.cpp as an inference engine, the webapp uses transformers.js, and they both support your v1/chat/completions endpoints :DD

You can try it out in your browser with zero setup!... running gemma-4-e2b ONNX in the browser, crazy stuff! Thanks to Xenova/HuggingFace for transformers.js c:

It passed the mom benchmark lol!

You guys told me that the framework was cool, but it was very manual to setup agents/workflows. So I've spent the last year slowly making it more accessible so anyone from any technical background can use it.

Every couple of months I ask my mom to use the App. And for the first time she actually was able to setup a monitoring agent with a local LLM! Which makes me think the app is ready for general public adoption (wuuuu!).

I hope this makes local LLMs useful for everyone! Tutorial/Demo Which is the whole point of the project.

My Commitment and being FOSS

The core Observer AI platform is, and will always be, free and open-source. That's non-negotiable. The code is all on GitHub for you to use, fork, and inspect.

The line in the sand which I have is "if it's free for me, it should be free for the user", that won't change ever.

Let's Stop Wasting Time!

This project wouldn't exist without the inspiration I've drawn from this community. You are the people I'm building this for.

I'll be hanging out here all day to answer any and all questions. Thank you again for everything!

Cheers,
Roy

▲
12
 
15👁
r/LocalLLaMA · u/DerTomsn · 10d ago
Swift-1.5-Qwen3.8-27b-oQ8e-mtp on Apple M5 Max — 34.8 tok/s — llm-bench.io

I ran Swift-1.5-Qwen3.8-27b-oQ8e-mtp through the llm-bench.io a few times today: oMLX on M5 Max 64 GB, thinking on at xhigh, 262k context window.

The big difference between Swift 1.5 and the base Qwen3.8 27B is how much it writes. Per full run (agent workflow, code generation, research, role play) Swift averages 51k generated tokens and Qwen3.8 averages 77k.

| Scenario|Swift 1.5|Qwen 3.8 27B|
|:-|:-|:-|
|Code generation|28.6k|40.1k|
|Research|11.1k|21.2k|
|Agent workflow|7.5k|11.9k|
|Role play|4.0k|3.8k|

The full benchmark run duration: avg. 24 min for Swift 1.5, avg. 38 min for base Qwen 3.8 27B

Everything else is about equal:

  • generation speed: 34.6 vs 32.8 tok/s
  • prompt processing: around 400 tok/s for both
  • quality score (the site's LLM judge): 85.8 vs 85.0. My four runs range from 84.3 to 87.2, so well within the expected variance of the llm judge. I'd call it a tie.

Still 3/4 of what Swift generates is reasoning, it just does less than Qwen 3.8 27B. The output is still very usable. I'll for sure give it a try to be my daily driver for a few day.

Runs Swift 1.5:

Runs Qwen 3.8 27B:

▲
15
 
13👁
r/LocalLLaMA · u/danielfrances · 10d ago
Help me find a good stack for reversing an old online game client

Hi, so I was working on building a local server for an older online game client a few years ago, and the amount of data I had to synthesize was intense. I ended up shelving the project. I had managed to build a basic login server, sorted out some crypt stuff, but it was just way too slow of progress for me. I've got some decrypted packets and lots of data to work with, so the LLM is not going to be forced to do this entirely blind.

I restarted it recently with Fable, and as expected, it has been a huge help. However, I'm consistently hitting the safety guardrails now that I am further into the project. I am wondering what you all would suggest for a local setup? I have 16GB of VRAM (RTX 4060 Ti) and 128GB of DDR4. If that is entirely insufficient, I might be willing to pay for hosting a more powerful local model. I'm fine with it being slow and chugging along all day and night - I am primarily concerned with it actually figuring out the client functions, and doing things as accurately as possible.

I appreciate any insight into specific models, harnesses, and other stuff I should be looking into. Thanks!

▲
0
 
2👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 10d ago
Would increasing ram channel improve speed with freetoken?

Would increasing the number of ram channels (eg dual channel to quad channel) help improve inference speeds with freetoken in hybrid gpu/ram setups or not? Or is increasing the vram capacity and vram speed the only optjon

▲
0
 
11👁
r/LocalLLaMA · u/opUserZero · 10d ago
Jev mode for images! post image

So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17

Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV

▲
10
 
10👁
r/LocalLLaMA · u/SignificantZebra5883 · 10d ago
50B+ MoEs with few active parameters, what's the sweet spot for intelligence, agent speed, and affordable fine-tuning?

I’m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:

Question → many sequential tool calls → final answer/document

Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.

I’ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. I’m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.

The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.

Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?

I understand that small active parameter counts don’t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also don’t assume that more total parameters automatically means a better model. I’m interested in where that tradeoff works in practice.

There are three things I’m trying to pin down:

  • Inference hardware: Ideally inference runs rented with parallel agentic loops (this is for a B2C project, not single person use, we scale based on demand)
  • Fine-tuning hardware: Obviously FT LoRA will take more memory than inference, max like 4GPUs on vastai fits the budget.
  • Agent performance: After it gets the prompt the tool calls and everything will be local, so imo it has no problems being blazing fast, as soon as the model calls a tool call it will be back very fast, so for this agentic use case, quick TTFT and t/s and adaptive dynamic reasoning are prefer right?

For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.

I’m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).

I’m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like “model + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,” along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.

Has moving from a \~30B model to a 50B+ low-active-parameter MoE actually improved your agent’s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading

💬 22 (+1) open on reddit ↗
▲
8
 
15👁
r/LocalLLaMA · u/sToeTer · 10d ago
Is there even an easy, seamless vision assistant program?

I read textbooks on my PC and ideally i want a program with a normal chat environment where i can just hammer in questions about what's currently on my screen. Example: I'm working on a PDF, underline or circle things... and then just type "what does this sentence mean?", you get it.

I do NOT want to manually screenshot, navigate to the folder, drag the picture into the environment and then also have to type the question. It should also naturally be aware that the conversation is about what's on the screen, so i don't have to steer it with "make a screenshot; use your vision capabilites" etc.

I tried multiple different MCP in LM Studio, none of them were great...or worked :/

Someone said AnythingLLM has this function but i couldn't find it.

Is there a good solution?

Thank you in advance! :)

▲
0
 
17👁
r/LocalLLaMA · u/ChopSticksPlease · 10d ago
What would you buy for $5k...$10k USD? post image

What (and if) would you buy if you had $5k ... $10k ... $20k to spend on local AI?

So, I'm a contractor and a solo dev working on some products/saas/apps. Basically I usually run up to three cline/opencode sessions in the same time, long running software engineering tasks, often run out of 128k context, so 256k ctx is prefferable. Pretty much every day for multiple hours so I could burn quite a lot of $ daily on OpenRouter. Fortunately, since Qwen3.8 i rarely need to delegate to larger models like Kimi K3 or MiniMax M3.

Apart from code I often work on confidential documents so a local AI or an approved remote AI is a must.

My current AI setup is:
\- dev server with RTX3090 running Qwen3.8 UD Q4\_K\_XL with 128k ctx q8
\- lab server with 2x RTX3090 + 128gb ram running Qwen3.8 Flash Next with 256k ctx

Both machines are fine to run up to three sessions, one on dev and 2 concurrent 128k ctx tasks on the lab server. The performance i get from the 2x RTX3090 with Qwen3.8 Flash Next is close to a single DGX Spark GB10 (according to numbers).

Soon I may need to run more agents and work with other people so started thinking of an upgrade.

Does it make sense to invest in either a single GB10 machine or two and cluster them to get more space for more context and therefore more concurrent sessions? Would you consider other options?

Any feedback appreciated.

💬 93 (+2) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/Dany0 · 10d ago
Where are the Opus 5.5 datasets?

Another day, another refusal. Apparently asking opus "what are your thoughts on this?" is an attempt at a 'distillation attack'

Our precinct is hugging face, we work at breakneck speed, we're up against art thieves, code thieves, extortionists, we're on call around the clock. The people of LocalLLaMA -- our finetunes is our job (and we write our own emdashes thank you)

WHERE ARE THE DATASETS PEOPLE. What happened to us? We used to throw pies at Dario Altman and now, what, we're penniless, downtrodden, what happened?

▲
40
 
18👁
r/LocalLLaMA · u/jacek2023 · 10d ago
Ornith-1.5 DFlash

Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash

Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-397B-DFlash

Ornith-1.5-35B-A3B-DFlash pairs the Ornith-1.5-35B-A3B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-DFlash

▲
0
 
8👁
r/LocalLLaMA · u/emperorofrome13 · 10d ago
Using unsloth I created the worlds best 9B model post image

#

https://huggingface.co/emperorofrome/Gmcoder

Beats Ornith 1.5 and Oxcoder on HumanEval+ Mini — and does it without the overthinking. It gets to the answer using 40–68% fewer tokens. Built as a finetuned merge.

Edit: Best in the world is just hype. The coder is comparable to Ornith 1.5 but more token efficient by about 50% on average up to 78% at times and can be faster.

▲
2
 
12👁
r/LocalLLaMA · u/Forward_Compute001 · 11d ago
Cheapest Epyc 7003 (Milan) Bundle (ddr4)?

I'm building a new rig to host the mission control application that should sit on its own node and I immediatly thought of a cheap single socket ddr4 solution,

does anyone have some suggestions which bundle is cheapest or gives best value for price...?

\-no need for gpus

\-no need for much ram (8gb ram sticks)

\-many threads and max core speed would be important (maybe if it doesnt spike the price)

▲
9
-1
24👁
r/LocalLLaMA · u/stevyhacker · 9d ago
Five local models, 6.8 GB of weights: my open-source Mac meeting notetaker

https://preview.redd.it/xs01ds930nsh1.png?width=4800&format=png&auto=…

Back in July I shared LokalBot here. It's a free, open-source Mac app that records your meetings and keeps a daily summary of your activity, all on-device.

0.9.2 came out today. Since July I've benchmarked every model in it and swapped most of the defaults for smaller ones. The whole stack is now 6.8 GB:

  • Qwen3-ASR 1.7B (MLX, 8-bit): transcription
  • Nemotron 3 (Core ML): who spoke when
  • Qwen3.5 4B Q4\_K\_M (llama.cpp): notes and action items
  • Harrier 0.6B Q8\_0: search embeddings
  • LFM2.5 1.2B Q4\_K\_M: autocomplete in any app
  • Apple Vision: screen OCR (opt-in)

A few numbers from my M4 Max (48 GB):

  • 26-min meeting to finished notes in 33 s warm, \~85 tok/s decode
  • Speaker error went from 43.4% to 14.6% DER on AMI. That's against my old pyannote setup, so it says more about my config than about pyannote.
  • Autocomplete p95 went from 1.83 s (Gemma 4 E4B) to 0.49 s

There's also a read-only MCP server and CLI, off by default, so Claude Code or any other MCP client can pull context from your meetings.

I don't have any 16 GB or other M series numbers yet. If you've got one of those, especially M5 or M6 I'd love to see what you get.

I also tried MiniCPM5 2B for notes. It was smaller and faster, but it got stuck repeating itself on one summary and assigned action items to the wrong person. I kept Qwen3.5 4B as the default as saving a few seconds wasn’t worth getting who agreed to do what wrong.

💬 4 (+1) open on reddit ↗
▲
6
-1
16👁
r/LocalLLaMA · u/lucasbennett_1 · 9d ago
on prem LLM stack for data that cant leave the building

Running the model locally is not a problem thats easy part but the leaks are the third party integrations along with it, like you designed everything perfect and then just added a cloud api along with it maybe a hosted judge for evals or a tracing saas or embedding point. one http call and the on prem things over

parts we already keep local are

  1. runtime: llama.cpp/ vllm /ollama
  1. models: qwen or llama family depending on rig
  1. vector db: pgvector or qdrant

some that leak but remain unnoticed:

  1. ingestion: pdfs and scans for some projects need a parse and the ocr step before chunking them and its where we often reach for a cloud parser and break the rule, although we can keep it local with liteparse sort of inbound parsers or other open source options on huggingface
  1. Eval: plenty of local setups still need prompts and outputs to a hosted judge or a tracing dashboard to see quality which is the same leak but seems different. instead a  local score set or a local judge model and keeping it self hosted if possible handles the tracing part

I am curious to know about others end to end stack who keep it 100% local, eager to learn more

💬 27 (+3) open on reddit ↗
▲
11
-1
8👁
r/LocalLLaMA · u/Brilliant-Hall1387 · 10d ago
Sherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version

Not an LLM, but the ternary findings should carry over, and we hadn't seen Sherry-style 3:4 weights run in a browser before. Disclosure: this is our work at Precisit, everything is MIT.

What it is

  • A 7.4M-parameter one-pass scorer (the jevlike family): the board goes in, one score per legal column comes out. No search.
  • Weights in T34, Sherry's 3:4 format: in every four weights one is zero and three are ±1, so four weights fit in 5 bits. One fp16 scale per 128 weights gives 1.375 bits per weight. The embedding is int8; norms and biases are fp16.
  • It runs in the browser on a small WebGPU runtime: 1.1 ms per move (idle M5 Pro, Chrome).

|Model|File size|vs depth-4 bot|vs depth-6 bot|
|:-|:-|:-|:-|
|dense (fp32)|29.7 MB|0.92|0.89|
|T34, trained ternary|1.59 MB|0.93|0.91|
|T34, fine-tuned from dense|1.59 MB|0.89|0.9|
|T34, converted after training|1.59 MB|0.13|0.11|
|Base243 (TQ1\_0 style), trained|1.93 MB|0.89|0.88|

200 games each, both sides play a random move 5% of the time, a win counts 1 and a draw ½.

What we learned

  1. Converting the finished model to 3:4 collapsed it (0.13 against the depth-4 bot). Training with the format in the forward pass fixed it completely, whether from scratch or fine-tuning.
  2. Attention's q/k/v matrices are the sensitive ones. Group size (64/128/256) barely mattered.
  3. Seeds matter: two runs of the same T34 recipe scored 0.945 and 0.882.

Play it:
https://precisit.github.io/onepass-web/demo/c4-size/

Code, models, every result:
https://github.com/precisit/onepass-webgpu-ternary

The write-up:
https://precisit.com/en/blog/onepass-c4-size/

Has anyone gotten post-training 3:4 conversion to work on models, or does it need training?

▲
17
-1
8👁
r/LocalLLaMA · u/giveen · 10d ago
exl3 now in ninfer-ext

My ninfer-ext fork of the famous ninfer now supports exl3 , and have released two models as well. https://huggingface.co/jabbatheduck/ninfer-ext-models

▲
6
-1
17👁
r/LocalLLaMA · u/HitarthSurana · 10d ago
Any good harness or tools for helping students?

What FOSS/self-hosted tools or AI tools have you found genuinely helpful for students? Also interested in anything non-AI that has helped with studying, notes, organization, research, etc. Curious what you guys actually use or found useful. (written fully by human)

💬 17 (+1) open on reddit ↗
▲
14
-2
11👁
r/LocalLLaMA · u/norenEnmotalen · 9d ago
peculiar-ragdoll's Dirk-Qwen 3.8-27B vs. UkisAI Swift-1.5 Qwen3.8-27B

EDIT: Post 2 with more model fint-tunes here https://www.reddit.com/r/LocalLLaMA/comments/1wv1ico/unsloth\_swift15\_peculiarragdoll\_thinkingcap/

I have a long list of my own domain specific eval questions that I run to validate which models I can rely on: coding, coding (numpy/pandas), data analytics decision making, local RAG, and voice assistant. It's made up of the types of things I'm likely to deal with on the daily. The test questions vary in dififculty and composition: easy, medium, hard.

System: M1 Max 32c 32GB with context 128K for Dirk and 110K for Swift.

Swift doesn't have XL. So I had to test with the L quant to stay as close as possible.

I ran the eval (using my tuieval tool) on peculiar-ragdoll's Dirk-Qwen3.8-27B-UD-Q4\_K\_XL and Swift-1.5-Qwen3.8-27B-Q4\_K\_L loaded with a modified version of Splash. The "amalgam" is a local I made out of incoai/Splash 1.1 and paperniuk's apple7-m1-kernels. It is modififed a little but not in ways that would alter model performance. I only merged and tweaked for some memory features I like from llama.cpp such as fit context check at the start of a load and personal QoL updates re auto-context manipulations that I don't want to think about, etc.

To say this result surprised me is quite an understatement. It's blown my mind.

When I did the first test a couple of days ago with only 44 questions, I thought it must be a prompt caching issue I missed that Dirk was benefiting from. I validated it is not and ran it against a lot more questions to certify it. It's a legit test outcome.

Dirk-Qwen is much sharper at getting to decisions and responses. The "be brief" instruction that gets passed each time in the chat templste is doing more magic than I had anticipated. It also gets more answers correctly with way less time consumed.

What trips up Swift-1.5 are mostly hard questions. It tries and tries until the 16,384 max token limit per question is reached and it fails with truncation.

Even when you ignore the 16,384 truncation failures and compare the other questions, Dirk token usage comes out on top.

Snipped view... this basically goes on pattern for another 191 unique questions.

https://preview.redd.it/otg3d1rkbqsh1.png?width=1420&format=png&auto=…

More importantly, this behavior is not just in question answering. You can see it in actual code refactor tasks.

On an unrelated note: tne model that has been able to pass a 100% of my eval packs is Opus 5.5. Deepseek Flash 4.1 fp32 got them all right except three.

💬 22 (+1) open on reddit ↗
▲
30
-2
19👁
r/LocalLLaMA · u/rorowhat · 10d ago
Best model for blender?

Trying to see if I can use a local model to generate game assets, or even 3D printer models. Any suggestions? Something that would fit in 64GB of ram, speed is not an issue. Just need it to work well.

💬 50 (-1) open on reddit ↗
▲
43
-6
20👁
r/LocalLLaMA · u/Haunting-Stretch8069 · 10d ago
Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide

I'm running Qwen 3.8 27B Q4 XS with \~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.

Build llama.cpp with Vulkan:

cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j

Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then:

llama-server \
--model Qwen3.8-27B-UD-IQ4_XS.gguf \
--mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \
--n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \
--batch-size 2048 --ubatch-size 512 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \
--load-mode none --fit off \
--cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \
--no-context-shift --jinja --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080

---

Edit: The process is documented here: https://zenodo.org/records/23088880. Feedback will be integrated into the upcoming Qwen 4 setup.