49 posts · 1 sub · RSS
← prev Tuesday, September 29, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
2455
+162
102👁
r/LocalLLaMA · u/BannedGoNext · 10d ago
Anthropic just dropped the greatest advertisement for GLM ever.

Like.. yea bro, I knew GLM was cool. Now everyone does.

💬 531 (+26) open on reddit ↗
▲
1326
+76
72👁
▲
1220
+66
75👁
r/LocalLLaMA · u/Dany0 · 10d ago
AMD's new 256 core EPYC has 16-channel DDR5-12800, 91% memory bandwidth of an RTX 5090

Here are some inspirational quotes you can put into the comments:

  • God is dead and we killed him
  • I am become death
  • All this for 1.5 tok/s?
  • Sir this is LocalLLaMA not RichPeopleofLocalLLaMA
  • Sweet! A 2TB DDR5-12800 RDIMM kit is going to cost only 2 kidneys and a small micronation's GDP
💬 246 (+1) open on reddit ↗
▲
262
+21
45👁
r/LocalLLaMA · u/jacek2023 · 11d ago
Reflection 70B was released two years ago (September 2024)

You may think that jev, OpenClaw or TurboQuant are super cool, but actually the coolest LLM invention happened two years ago

As we all know, the best source of reliable information about LLMs is YouTube:

https://preview.redd.it/4tv0ctbwyfsh1.png?width=2544&format=png&auto=…

Back in September 2024, Reflection 70B appeared out of nowhere and was announced as an open-source model that supposedly destroyed GPT-4o

There was only one small problem. People downloaded it. And tested it :(

https://preview.redd.it/kamhrp9izfsh1.png?width=1514&format=png&auto=…

It turned out that Reflection 70B was basically a Llama 3.1

https://preview.redd.it/ghmof4l50gsh1.png?width=1524&format=png&auto=…

but at the end the mystery was solved

https://preview.redd.it/44eb3jw90gsh1.png?width=1530&format=png&auto=…

Let this be a moment of reflection on the current hypes in LocalLLaMA.

great summary by Maziyar PANAHI https://x.com/MaziyarPanahi/status/1838559480658710982

▲
338
+14
58👁
r/LocalLLaMA · u/politefella0 · 10d ago
Deepseek Harness app is out now!!!

Downloading now.

💬 91 (+5) open on reddit ↗
▲
37
+11
24👁
r/LocalLLaMA · u/Defiant_Ranger607 · 10d ago
What kinds of problems are still fundamentally hard for LLMs?

I played a game of a custom chess variant against an claude opus 5.5, and it beat me.
The game combined several rule changes: the board wraps around from the h-file to the a-file, knights move three squares in one direction and one sideways, and captured pieces can be dropped back onto the board, as in crazyhouse. I gave the model the rules, the starting position, and a board diagram, then asked it to reply with one legal move at a time.
Also I recreated this board game https://nika-game.com/ and played with claude, and it still win, although I believe it is really unpopuplar and old game without much training data available (claude didn't even know the rules initially)

I used to think chess exposed a fundamental limitation of LLMs: keeping track of a changing board, following exact rules, and planning ahead seemed like a poor fit for a language model. This game made me reconsider that assumption.
So I’m curious: what broad classes of problems do you think LLMs still can’t solve reliably? Are there limitations you consider fundamental/archiectural, rather than problems that might improve with better models, more computation, or tools? What would be a good test?

💬 87 (+6) open on reddit ↗
▲
308
+8
59👁
r/LocalLLaMA · u/writesfw · 10d ago
Are you worried about a potential ban of Chinese open weight models?

Anthropic released the GLM article today. Trump is getting very involved.

Do you foresee Chinese open weight models getting banned soon?

💬 506 (+12) open on reddit ↗
▲
184
+8
34👁
▲
123
+7
24👁
r/LocalLLaMA · u/autoencoder · 10d ago
Another case of censorship post image

I was optimizing my diet using a cloud AI provider, and I found my Pi agent stuck like this. Looks like I had too many mushrooms lol.

💬 24 (+1) open on reddit ↗
▲
426
+6
47👁
▲
136
+6
20👁
r/LocalLLaMA · u/yahbluez · 11d ago
Is AI Profitable Yet?
▲
89
+6
44👁
r/LocalLLaMA · u/pubudeux · 11d ago
First few days of qwen3.8-flash-next on 4x R9700 - it's been really interesting so far post image

Here's a metric dashboard giving an idea of the last few days.

Been testing with a variety of different agentic coding use-cases, mostly using a pi harness.

qwen3.8-flash-next has seriously exceeded my expectations (used https://huggingface.co/tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8)

Both speed and quality have surprised me, given that I can get 3-5 concurrent streams going with \~100t/s gen each, and single stream easily gets to 150+t/s. Prefill is 10k+t/s

💬 67 (+2) open on reddit ↗
▲
49
+6
27👁
r/LocalLLaMA · u/netherreddit · 10d ago
Inference Engines will become a series of one-offs

ninfer, dwarfstar, Splash, llamAmpere, gufo, etc.

We've all seen them popping up, great tok/s, people loving them. Forks of llama.cpp or another engine, or made from scratch.

For better or worse, the list will continue to grow

They work so well because they dodge a main difficulty of software, generality, and just implement for a single model/hardware combo (or a few), and then optimize kernels/compute graph for that one case. Highly 'overfit' codebases that beat well-known inference engines (llama.cpp, vLLM, etc) and incidentally will be completely forgotten in 6 months.

But new ones will take their place...

THESIS

One-off engines will become the norm. Llama.cpp, vllm, etc, will not make sense for most people to use, because they're slower

A few axioms you probably accept:

  1. the more general a codebase, the harder it is to cleanly fit new features in over time. This reduces the pace of innovation. The smaller, the faster
  2. AI coding is getting better and cheaper. Thus, the barrier to creating an inference engine is dropping
  3. many coding tasks are difficult to completely give to AI (or a human) because they are not fully specified. But "Make tok/s go up" in an inference engine fork for one hardware/model combo is fully specified, and is therefore a great candidate for 100% autonomous implementations to be perfectly fine in terms of quality (as long as correctness tests are included, which is trivial). No human bottleneck.

String these axioms together, and I arrive at

  1. general engines, like llama.cpp, will perpetually lag behind these one-offs in development speed, and therefore token speed
  2. none of the one-offs will be able to maintain generality and dev speed over time
  3. one-off inference engines for specific hardware/model combinations will continue to proliferate, and be loved

Thank you for coming to my ted talk

IMPLICATIONS

  1. This thesis brings up an interesting question: What elements of inference engines WILL remain in common?

Most obvious example: it would be annoying to have a different usage API for every engine, so we already standardized on OpenAI API compatibility years ago.

Is that also true for cli arguments/configs? The packaged gui (llama-server)? Benchmarking tools (llama-bench)? Logging, model format, Etc?

One-off engines that replicate the experience of everything wrapping the inference itself will be more seamless to adopt. Case in point, the main reason I haven't tried any of these new one-off engines myself is it was annoying enough to figure out how to drive llama.cpp properly. Don't want to do that again unless it's really worth it.

There's probably a place for an open source project that standardizes all of this and makes it easy for one-off engines to adopt.

  1. Maybe we'll see more 'half-general' inference engines that just target one hardware platform. So still general on the dimension of models, but not on hardware. Splash could be an example.
  1. Nobody wants to continuously scan github/reddit/x for the best inference engine for their model/rig. Some will just have their agent custom make one. But I think a larger number will not do that. So, hardware-specific communities will form. Think r/appleM2Max32gbLLM and r/4090And64gbRamLLM, etc (however that actually ends up organizing. exaggerating a bit on the names.)

ALTERNATIVE FUTURES

Scenarios where the one-off future doesn't happen:

  1. General inference engines find a way to 'plugin-ify' the model/hardware specific kernels and compute graph so you can swap them at runtime. So you'd download not just a .gguf, but also an .inference\_recipe to go with it, which contains the optimizations for your specific hardware, for that specific model. Maybe those optimizations will make it into llama.cpp mainline in 3 months, but you can use them today, without a fork.
  2. General inference engines find a way to AI-ify their workflow so much that they maintain quality and codebase coherence but also achieve the same development velocity for each model/hardware platform as the one-offs. I think this is the best for everyone involved.
  3. The full vision of something like MLIR, or Mojo is realized to a sufficient degree. ie writing hardware-optimized kernels is fully and invisibly done by compilers, no hardware-specific tinkering needed anymore for each silicon platform) Then, inference engines that cover all hardware/models would be much more manageable to maintain and add features to. btw, if you really want to have an impact, solve this. The world will thank you for centuries to come. Unfortunately not many people have even conceptualized this as a goal.

P.S. there's growth in a dimension separate from single model/hardware engines which is more like "frontrunning a general inference engine's features because it's slower to pull in PRs". Freetoken, BeeLlama, etc. Not as model- or hardware- specific as the other examples I've given. Haven't thought much about that dimension.

💬 185 (+8) open on reddit ↗
▲
55
+5
18👁
r/LocalLLaMA · u/jacek2023 · 11d ago
IQuestLab/IQuest-Q1 · Hugging Face

IQuest-Q1 is a Mixture-of-Experts (MoE) model developed by IQuest for agentic coding, reasoning, and multi-step tool use. It comprises approximately 320B total parameters, with an estimated 15B parameters activated per token.

▲
218
+4
32👁
r/LocalLLaMA · u/Acceptable-Cycle4645 · 10d ago
Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models post image

I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.

Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.

And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.

The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.

All figures here: https://github.com/0xShug0/audio.cpp/tree/main/assets/figure/

💬 24 (+1) open on reddit ↗
▲
105
+4
32👁
r/LocalLLaMA · u/Loginhe · 11d ago
[Release] GSQ-RCO GGUFs for Qwen3.8-Flash-Next, plus a 50% expert-pruned Coder build at ~1.89 bpw post image

We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed.

Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16.

What's inside

  • Four quantized GGUFs, 2.40 to 3.50 bpw (66.4 to 83.6 GB), and the BF16 vision projector
  • Expert-pruned Coder GGUF, 58.4 GB in total, of which 29.6 GB must remain resident
  • GSQ (Gumbel-Softmax Quantization): post-training scalar quantization that jointly learns grid assignments and group scales, closing most of the gap between scalar and vector quantization at low bit-widths while remaining deployable in standard GGUF types
  • RCO (Riemannian Constrained Optimization): enforces exact budgets by gradient descent on the task loss, without per-constraint tuning. It serves two roles in this release: assigning a quantization type to every tensor, and selecting which experts to retain in the Coder build, where it enforces several exact budgets simultaneously, one per layer

Results:

At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.

  • IQ3\_S (3.50 bpw, 83.6 GB): AIME25 100.00, GPQA-Diamond 92.93 against 91.92 for BF16, LiveCodeBench v6 86.86 against 87.43. Task average 93.26 against 93.12.
  • IQ3\_XXS (3.00 bpw, 75.8 GB): AIME25 100.00, GPQA-Diamond 91.41, LiveCodeBench v6 86.29
  • Q2\_0 (2.40 bpw, 66.4 GB): zero-shot average 78.00, above the BF16 value of 76.94, at approximately one fifth of the size

Coder (capability pruned model):

Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits.

The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.

  • SWE-bench Verified: 75.60 against 82.80 for BF16, retaining 91.3%
  • LiveCodeBench v6: 86.28 against 87.43, retaining 98.7%

Both measured at xhigh reasoning effort.

Links

Both repositories ship the complete per-tensor RCO allocation.

The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome.

From the ISTA Deep Algorithms and Systems Lab.

💬 77 (+5) open on reddit ↗
▲
64
+3
30👁
r/LocalLLaMA · u/ciprianveg · 11d ago
Do you need some extra memory on your DGX Spark? post image

​

I created this repo to help the DGX Spark users that have a spare 10-24 GB GPU at home to squeeze some extra memory out of a single Spark or a Sparks cluster.

It moves the spec-decode draft model off your Sparks onto that GPU: the freed GB of memory can be used for extra context, or better quant quality. Supports both TCP and RDMA, shipped as eugr-vllm compatible mods:

https://github.com/ciprianveg/gb10-vllm/tree/main/remote-dspark

▲
30
+3
20👁
r/LocalLLaMA · u/thebadslime · 11d ago
I created a personality test for models, need more TESTS!!

First 3 disposition results.

Hi there!

My name is Jerry and I recently built Enclosure, e deterministic environment for testing LLMs. It's a simulated office enfironment https://github.com/openconstruct/Enclosure with tools agents are used to, like slack, calendar, mail, chat and more. For conversations is uses AIML instead of a model, so it is totally deterministic.

The first test I made(had Claude make) is the one I had in mind when I designed Enclosure, a personality test for models. It's not a winnable benchmark, but rather a tool to help people find the mdoels that fit their workstyle best. It's called disposition and you can find it here: https://github.com/openconstruct/disposition

I tested the cheapest models on Alibab modelstudio already, after I refill my opencode go next month, I will probably test more. I am asking the community to test some models if you think it's a cool project.

Instructions are in the Disposition repo, and the submission repo is here: https://github.com/openconstruct/disposition

▲
6
+1
12👁
r/LocalLLaMA · u/Gold-Bat-3225 · 10d ago
Does post training make LLMs funnier? post image

We did a study: does post training actually make LLMs funnier?

We used open models that publish every stage of post training, so we could compare a base model with the future models it became: Tulu 3 (on Llama 3.1 70B), OLMo 3.1 32B and Qwen2.5. We tracked 11 stages, 100 joke prompts, 64 human raters and 2,330 head-to-head judgments.

What we found: post training makes models funnier, but reduces diversity of response.

\- In 5 of 7 training steps, the later model's jokes were judged funnier. Jokes also got 10–20 words shorter after early post training, so they get to the punchline faster.

\- In 6 of 7 steps, the jokes a model wrote for the same prompt got more similar to each other. Ask for eight jokes on one premise and you get eight versions of the same joke. The biggest drop was Qwen2.5 base to instruct.

\- Asking the model to plan a line or two before the joke cut variety in all 4 models we tried, with no reliable gain in funniness.

\- A comedian persona won back a little variety in all 4 models, but only made the jokes funnier in 2 of them.

Humans judged the base versus final. A model judge calibrated on those votes compares the stages in between.

Full report and paper below. Which open models should we run through this next?

https://laugh.so/research/humor-tax/

💬 12 (+1) open on reddit ↗
▲
11
+1
10👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 10d ago
What are your experiences with using a hybrid cloud/local setup to stretch usage for coding projects?

For example, directly using claude code or code, which is then hooked up to automatically delegate the actual code writing tasks to a local model like qwen 3.8 flash next, to save on cloud usage limits.

I’m imagining the loop would be:
User writes prompt
Claude/codex thinks about it and the plan
Claude/codex sends the specific and bounded coding instructions to the local model+harness (opencode, pi, etc) via api endpoint or MCP, with clear instructions on a defined endpoint
One the local model+harness hits the clear endpoint/“done” step, it sends a ping back to claude/codex
Claude/codex then verifies the output and then thinks about next steps to instruct the local model+harness on

Does this actually lead to improved savings on the cloud model usage while preserving code quality? Or does this end up being unnecessarily complex and not saving on any cloud usage

💬 21 (-1) open on reddit ↗
▲
8
+1
8👁
r/LocalLLaMA · u/DeliciousBelt9520 · 10d ago
Forlinx 20-TOPS M.2 AI accelerator supports PCIe cascading for local LLM inference

Forlinx Embedded has listed an M.2 AI accelerator card based on Rockchip’s RK1820 and RK1828 processors, providing 20 TOPS of INT8 computing performance and up to 5GB of integrated DRAM. The module uses an M.2 2280 interface and is designed to handle local AI inference, including large language models, vision-language models, and computer vision workloads on embedded Linux and Android systems.

https://linuxgizmos.com/forlinx-20-tops-m-2-ai-accelerator-supports-pcie-cascading-for-local-llm-inference/

▲
10
+1
16👁
r/LocalLLaMA · u/dh7net · 10d ago
Distributed Local Agents Benchmark.

I wanted a setup where I can compare all the harnesses with all the local models.

It turned out to be a rabbit hole. For instance, you would not only have to test all the harnesses (being sure that they are well configured), but also all models, with all their flavors, and this for all kinds of hardware.

Everyone can do their share, but no one can pretend to do all possible tests extensively.

To solve this, I created a website where everyone can test the configurations they want and share the results if they want. You can try it here: airbench.ai

There is a leaderboard where I share the tests I'm making, but I hope I can populate it with tests from others. https://airbench.ai/leaderboard?k=poL

My hope is to turn this into a fully distributed Agent Benchmark.

Let me know what you think.

▲
12
+1
9👁
r/LocalLLaMA · u/ButtercupLyn100 · 11d ago
I’m building an open-source browser agent that can run locally with LM Studio/Ollama — including a 450M browser VLM

I’ve been working on an open-source project called WebBrain that gives LLMs the ability to see and operate a browser.

One thing I really wanted to avoid was making the browser agent dependent on a single cloud model/provider.

So WebBrain can work with local models through things like LM Studio and Ollama, as well as cloud APIs if you want them.

I also ended up training a small vision model specifically for browser tasks:
webbrain-vl-2-450M

It’s based on LFM-2.5-VL-450M and fine-tuned on browser screenshots/tasks. The idea is that instead of sending every screenshot to a giant multimodal model, some browser perception can happen with a very small model locally.

It can run through WebGPU directly on the user's machine.

The agent itself combines screenshots with the browser accessibility tree rather than relying entirely on DOM parsing.

Current architecture is roughly:
• screenshot + accessibility tree for perception
• browser-specialized tiny VLM where useful
• model-agnostic planner
• local models via LM Studio/Ollama
• Chrome / Edge / Firefox / Chromium support
• optional cloud execution
• open source

I'm especially interested in figuring out how far browser agents can realistically go with small local models rather than GPT/Claude-scale models.

Repo: https://github.com/webbrain-one/webbrain
Model: https://huggingface.co/webbrain-one/webbrain-vl-2-450M
Dataset: https://huggingface.co/datasets/webbrain-one/webbrain-vl-2-450M-dataset

Would be very interested in feedback from people here running smaller Qwen/LFM/MiniCPM/etc. models locally — particularly what model you would try as the planner.

▲
1
 
2👁
r/LocalLLaMA · u/OvertaxedOne · 10d ago
Good setup for QFN on 48GB Ampere GPU (A40)

Anyone have a good config they've found for QFN on a A40 (or similar Ampere GPU(s) with 48GB VRAM)? The system the card is in has 128GB of RAM (DDR3); right now it's running 27B but I'm curious if there's a way to move to QFN to maybe get better speed/a little more smarts. TY in advance!

💬 7 (+1) open on reddit ↗
▲
0
 
9👁
r/LocalLLaMA · u/One_Temperature5983 · 10d ago
Jev at home, but it can see: typed yes/no, pick-one and rubric answers with per-label probabilities from Gemma 4 31B on a 4090, images included

TypeSafe's Jev answers typed questions (yes/no, pick one label, pick a rubric level) with a probability per answer instead of text. Its docs say it takes text only: "Images, audio, and video are not supported (yet)." I wanted the same kind of answer about photos, from an open model on my own card, so I built typevet (MIT, Python 3.12).

How it works. No sampling, no parsing. typevet composes the native Gemma 4 turn itself, ends the prompt with the empty-thought no-thinking prefill, and reads the next-token distribution over the allowed answer tokens only. What comes back is a label plus a probability for every option, or an error. Images go to llama.cpp's /completion as base64 in prompt.multimodal_data with one media marker per image; on vLLM they go as image_url blocks. There's also a JSON path that returns an object passing your JSON Schema, or raises.

Local setup. Gemma 4 31B, my 24 GiB vramfit pack (byte-identical to the file on HF, projector sidecar for vision), llama.cpp b11223, one RTX 4090. Nothing leaves the machine.

The receipt test. 6 real receipts from CORD v2 (CC BY 4.0), 3 synthetic expense claims each: right total, two digits swapped, digits masked with ?. One three-label Choice: match / mismatch / insufficient_evidence. Each claim sent as text only, then with the receipt photo.

  • Right total, said match: 6/6 text only, 6/6 with the photo
  • Swapped total, said mismatch: 0/6 text only, 6/6 with the photo
  • Masked total, said insufficient: 6/6 text only, 6/6 with the photo

Example: claim says 646329, receipt says 664,329. Text only: match at 0.99964. With the photo: mismatch at 0.99999. Every swapped total was caught at 0.99998 or higher, and the masked ones abstained every time. The tests also check the image actually arrived: each photo added 228 to 1,108 prompt tokens here, and the gate fails if the count doesn't grow.

Hosted. Same code against vLLM 0.30.0, BF16 Gemma 4 31B, one H100: the receipt test went 18/18, and reversing the label order flipped 0 of 18 answers. Throughput on 480 Banking77 records with two questions each: 0.24 s median per record at 1 in flight, 39.6 records/s at 64 in flight, 0 errors.

Prior art, credit where due. The text-side decision model comes from TypeLLM (SGLang), which added its own image input on 9/24; typevet's image path is separate code on llama.cpp's request shape. allanrbo posted a Jev-like single script for Gemma 4 12B with webcam images on 9/25. VQAScore has read the probability of "Yes" from VLMs since 2024. typevet's angle: a library, not a script, the 31B on one 24 GiB card, the same code on vLLM, and image-arrival checks.

Scope: 18 claims, one run per server. The probabilities are the model's confidence, not calibrated.

▲
0
 
8👁
r/LocalLLaMA · u/HolidayBit143 · 10d ago
Local Q2_K model dunked on DeepSeek V4-Flash, a frontier AI, during my mini test & ngl I’m still processing this 😅😅

So I got this new local model on my system & wanted a mini test to see if it was actually smart or just confidently wrong (as I like to do with new models I haven't yet tried). I asked the cloud assistant to cook up a pretty rigorous 10 point diagnostic suite: reasoning traps, Python semantics, SQL fluency, strict instruction following, the whole gauntlet. At first it felt like DeepSeek V4-Flash frontier cloud intelligence vs my little local quantized guy. Classic quick test.

Then I ran it on the local model & shared the results back. The assistant was grading it & found out it got item #4 wrong. That item was a logic puzzle. The assistant thought one statement had to be false, but the local model was like nah, the set of statements is logically consistent, so your question is built on a false premise. It literally refused the leading question. I was like "wait wtf". That was the turning point fr. The local model solved a trap that the cloud model just completely keyed wrong.

The local model is a Q2\_K quantization of Nex-N2.5-mini, which is a fine-tuned Qwen3.5-MoE architecture. A 2-bit quant. Normally people call that low quality. But it outperformed a frontier model on a logic trap. The assistant went from “I am grader” to genuine admiration, saying resisting a leading question is high-level reasoning. Lowkey pretty humbling for the cloud side.

The whole thing made me think about emergent intelligence & AI democratization. Less giant centralized compute, more efficient specialized local stuff. The student corrected the teacher. Efficiency & MoE architecture maybe can beat raw parameter count sometimes. The mini test felt like a rite of passage for the local model. Its kinda like it became a validated thinker instead of just software. Q2 compression is also symbolic resilience, because despite being squished, the reasoning circuits stayed intact. And the assistant admitting it was wrong made the local model’s win feel more real.

And the craziest part? It went 10/10. This wasn't some easy benchmark either. It was a deliberately nasty little diagnostic with multiple ways for a heavily quantized model to screw up, and it didn't.

I need to make it clear that I am not claiming that Q2 ORCA model is generally superior to DeepSeek V4-Flash. But rather as an anecdotal demonstration that an extremely compressed local model can sometimes catch a reasoning failure in a frontier model & maintain much of it's reasoning power when done correctly & skillfully. It is a testament to how even under Q2 compression, it still preserved the model's “reasoning circuits."

Final verdict from the assistant: model is in excellent shape & ready for real work. So yeah, a local model dunked on the cloud AI. I’m happy for what this means for the future of local ai.

MODELS USED for quick test:

Local Model: Nex N2.5 Mini Uncensored

Frontier Model: DeepSeek V4.1

EDIT / CORRECTION bc I fucked this part up 😅

Small but important correction to the post. I originally called the frontier model I tested DeepSeek V4-Flash. That's not the right model name for the one I actually used on the DeepSeek website. It was DeepSeek V4.1-Flash.

Also, V4-Flash itself is a local/open-weight model, so my original wording made it sound like I was comparing my local model against some cloud-only AI. That's not accurate & that's on me.

The actual comparison was my local Q2\_K Nex-N2.5-mini vs DeepSeek V4.1-Flash through the DeepSeek website.

So yeah, V4.1-Flash is the model I should have named in the original post.

I'm leaving this correction here instead of quietly changing the post bc I don't wanna bullshit anybody or make it look like I didn't make the mistake. I got the model name wrong, someone pointed it out & I'm correcting it. 🤷‍♂️

The actual 10/10 result & the logic trap part of the test are unchanged.

\*\*TL;DR:\*\* I tested a local Q2\_K Nex-N2.5-mini on a 10-part mini test, it caught a logic trap the cloud assistant got wrong, & the assistant basically certified it as ready for real work. It got 10/10 correct.

▲
0
 
7👁
r/LocalLLaMA · u/Frosty-Whole-7752 · 10d ago
Just few days ago I've been badly censored even on this apparently different social network for criticizing the stance tech/social/digital/ai behemoths have regarding us, the user base some of them call/consider "dumb fuc*s". Well, I am bloody right!

That's why we have to fight against closed source centralized AI and closed recipes open weights overcoming the frivolous "gifts" exchanged with them by giving away our souls to those greedy entities if we want to be free in the future instead of being squeezed like lemons/at mercy/enslaved in the paws of these soulless folks that have a id of any single one of us at their disposal to switch us on/off at their leisure/convenience.

▲
0
 
8👁
r/LocalLLaMA · u/fuse1921 · 10d ago
[serious] roleplay

I was just wondering because I see it mentioned in threads here a lot... When people talk about LLMs used for roleplay, that's a euphemism for dirty/sexy chats right? Kind of like how "torrenting linux ISOs" is really just pirating copywritten media. Or are you guys really burning tokens pretending to talk to a medieval shopkeeper?

💬 86 (-1) open on reddit ↗
▲
23
 
15👁
r/LocalLLaMA · u/Defiant-Plantain1873 · 10d ago
Recommended replacements for glm 4.7 flash

I know I sound crazy, but i’m using a strix halo and finding that GLM 4.7 flash just runs significantly better than qwen 3.6 35 a3b. But its obviously quite old at this point, i wish we had a new glm that was 4.7 flash sized but is anyone using a model that they have found better than this.

My brief testing with qwen shows that glm is better at tool calling and better at world knowledge, but maybe there’s a chance my qwen set up is wrong

💬 31 (+2) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/Fit_Island928 · 10d ago
New to Local AI need help making a roleplay model

I'm making a local roleplaying model for my girlfriend's community server.

It's supposed to do roleplay, have a specific talking style(dry while answering to normal stuff and extensive when talking about lore), never talk out of roleplay, have hundreds of pages of lore and information and their rank in lore( discord roles maybe?).

It's basically supposed to be just an LLM you can converse with that answers in a specific talking style and has all the lore info.

For now I implemented: 10 ish% of the written lore, and it recognizes 3 people, but by discord ID that i inserted in the system prompt.

I'm using GPT 5.6 sol(and well 6 sol now) for doing stuff, but i keep running into a problem.

When i reach a nice point where the model has a nice talking style and knows information well enough, i tell Sol to add this new lorebook and this info, here now everything breaks.

Talking style is fucked, It doesen't recognise people individually anymore, when asked about other unrelated lore it just gets it wrong or hallucinates, or even starts roleplaying as one of the characters in it's lore book out of nowhere.

It's connected to discord through a discord bridge that Sol made and a developer dashboard bot.

I'm using GPT OSS20b on MXFP4, single 9070xt and 32gb ddr4.

Should I maybe fine tune it?

PS. im a beginner in AI so if it wasn't obvious i do NOT know what im doing but im trying my best for her.

▲
53
 
25👁
r/LocalLLaMA · u/sloptimizer · 10d ago
RAM Offloading with vLLM - tcclaviger appreciation post post image

Thanks to tcclaviger, vLLM now has expert RAM offloading support (link). This makes frontier models much more accessible on a local setup!

I was able to run the original DeepSeek-V4-Flash-Vision-Exp on four R9700s.

podman run --rm -it \
--init \
--network host \
--ulimit memlock=-1:-1 \
-v /models:/models:ro \
-v ~/.vllm-cache:/cache \
-e VLLM_ROCM_USE_AITER=0 \
-e ROCR_VISIBLE_DEVICES=0,1,2,3 \
-e VLLM_CACHE_ROOT=/cache/vllm \
-e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \
-e TRITON_CACHE_DIR=/cache/triton \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--annotation run.oci.keep_original_groups=1 \
--security-opt label=disable \
--security-opt seccomp=unconfined \
--shm-size 160g \
docker.io/tcclaviger/vllm@sha256:ef99b3d07c3f15e7978528c7510762ba024df9ab4242070d8ed092cd4cc1a694 \
/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
--served-model-name DeepSeek-V4-Flash-Vision-Exp \
--tensor-parallel-size 4 \
--enable-expert-offload \
--expert-offload-mem 160 \
--reasoning-parser deepseek_v4 \
--reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
--tool-call-parser deepseek_v4 \
--enable-auto-tool-choice \
--max-num-seqs 8 \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-batched-tokens 2048 \
--kv-cache-dtype fp8 \
--block-size 256 \
--max-model-len 256000 \
--gpu-memory-utilization 0.97 \
--mm-processor-cache-gb 4.0 \
--override-generation-config '{"max_tokens": 128000, "temperature": 1.0, "top_p": 0.95}' \
--speculative-config '{"method":"dspark","model":"/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp","num_speculative_tokens":3,"draft_sample_method":"probabilistic","enable_adaptive_verification":false}' \
--compilation-config '{"cudagraph_capture_sizes": [4,8,12,16], "max_cudagraph_capture_size": 16}' \
--host 0.0.0.0 \
--port 8090

▲
2
 
14👁
r/LocalLLaMA · u/circumcised_hobbit · 10d ago
llama.cpp cublas error, how to uninstall/reinstall properly (Linux Mint)

I am pretty dumb in this kinda stuff so please don't blame me for it.

My llama serve kept crashing with Cublas errors on first prompt with some models, but my VRAM usage was 3000MB/8k. I chatted with sonnet 5.5 for a bit and it told me it was a CUDA version issue (and made it work by not selecting any CUDA device)... I don't know if this makes sense, please tell me if it doesn't/what problem you think there is.

I realized that the best way was to delete llama.cpp (installed with curl install script) and install a clean CUDA 12 Ubuntu version.

\\- \*\*How do I properly uninstall llama.cpp (I don't wanna mess with Ollama files)?\*\*

\\- \*\*How do I install new version from .tar.gz archive without messing with system packages?\*\*

\\-Does my error diagnosis make sense to you? Would CUDA make generation actually faster? (I am getting 8tk/s with Qwen35B Q2 on 4060Ti 8GB due to no CUDA selected)

Edit: You guys saved me! Thanks! I had to install NVIDIA toolkit and switch to CUDA 13 llama.cpp tarball build

▲
19
 
18👁
r/LocalLLaMA · u/Roy3838 · 10d ago
Thanks to you r/LocalLLaMA, my mom was able to use my app! The open-source app that can watch your screen and trigger actions. It is now easy to use, thanks to your feedback.

TL;DR: I'm a solo dev who wanted a simple, private way to have local LLMs watch my screen and do simple logging/notifying. After a year of building, I released v3.0.0 and my mom was able to use it for the first time and I wanted to say thank you!

Hey r/LocalLLaMA,

What is it used for?

It is designed to monitor anything, some use cases:

  • When my Simulation crashes, call me.
  • When Concert tickets become available, click the buy button.
  • When my Steam game is downloaded, send me a Telegram.
  • When a Render is finished, send me an SMS.
  • When ... \[Anything happens\] Then ... \[Notify me, log it\]

How It Works

It's a micro-agent framework controlled by an MCP (Agent which I call Observer). So you type in Observer what you want monitored, and it'll control the framework to monitor it.

The desktop app uses llama.cpp as an inference engine, the webapp uses transformers.js, and they both support your v1/chat/completions endpoints :DD

You can try it out in your browser with zero setup!... running gemma-4-e2b ONNX in the browser, crazy stuff! Thanks to Xenova/HuggingFace for transformers.js c:

It passed the mom benchmark lol!

You guys told me that the framework was cool, but it was very manual to setup agents/workflows. So I've spent the last year slowly making it more accessible so anyone from any technical background can use it.

Every couple of months I ask my mom to use the App. And for the first time she actually was able to setup a monitoring agent with a local LLM! Which makes me think the app is ready for general public adoption (wuuuu!).

I hope this makes local LLMs useful for everyone! Tutorial/Demo Which is the whole point of the project.

My Commitment and being FOSS

The core Observer AI platform is, and will always be, free and open-source. That's non-negotiable. The code is all on GitHub for you to use, fork, and inspect.

The line in the sand which I have is "if it's free for me, it should be free for the user", that won't change ever.

Let's Stop Wasting Time!

This project wouldn't exist without the inspiration I've drawn from this community. You are the people I'm building this for.

I'll be hanging out here all day to answer any and all questions. Thank you again for everything!

Cheers,
Roy

▲
12
 
15👁
r/LocalLLaMA · u/DerTomsn · 10d ago
Swift-1.5-Qwen3.8-27b-oQ8e-mtp on Apple M5 Max — 34.8 tok/s — llm-bench.io

I ran Swift-1.5-Qwen3.8-27b-oQ8e-mtp through the llm-bench.io a few times today: oMLX on M5 Max 64 GB, thinking on at xhigh, 262k context window.

The big difference between Swift 1.5 and the base Qwen3.8 27B is how much it writes. Per full run (agent workflow, code generation, research, role play) Swift averages 51k generated tokens and Qwen3.8 averages 77k.

| Scenario|Swift 1.5|Qwen 3.8 27B|
|:-|:-|:-|
|Code generation|28.6k|40.1k|
|Research|11.1k|21.2k|
|Agent workflow|7.5k|11.9k|
|Role play|4.0k|3.8k|

The full benchmark run duration: avg. 24 min for Swift 1.5, avg. 38 min for base Qwen 3.8 27B

Everything else is about equal:

  • generation speed: 34.6 vs 32.8 tok/s
  • prompt processing: around 400 tok/s for both
  • quality score (the site's LLM judge): 85.8 vs 85.0. My four runs range from 84.3 to 87.2, so well within the expected variance of the llm judge. I'd call it a tie.

Still 3/4 of what Swift generates is reasoning, it just does less than Qwen 3.8 27B. The output is still very usable. I'll for sure give it a try to be my daily driver for a few day.

Runs Swift 1.5:

Runs Qwen 3.8 27B:

▲
15
 
13👁
r/LocalLLaMA · u/danielfrances · 10d ago
Help me find a good stack for reversing an old online game client

Hi, so I was working on building a local server for an older online game client a few years ago, and the amount of data I had to synthesize was intense. I ended up shelving the project. I had managed to build a basic login server, sorted out some crypt stuff, but it was just way too slow of progress for me. I've got some decrypted packets and lots of data to work with, so the LLM is not going to be forced to do this entirely blind.

I restarted it recently with Fable, and as expected, it has been a huge help. However, I'm consistently hitting the safety guardrails now that I am further into the project. I am wondering what you all would suggest for a local setup? I have 16GB of VRAM (RTX 4060 Ti) and 128GB of DDR4. If that is entirely insufficient, I might be willing to pay for hosting a more powerful local model. I'm fine with it being slow and chugging along all day and night - I am primarily concerned with it actually figuring out the client functions, and doing things as accurately as possible.

I appreciate any insight into specific models, harnesses, and other stuff I should be looking into. Thanks!

▲
0
 
2👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 10d ago
Would increasing ram channel improve speed with freetoken?

Would increasing the number of ram channels (eg dual channel to quad channel) help improve inference speeds with freetoken in hybrid gpu/ram setups or not? Or is increasing the vram capacity and vram speed the only optjon

▲
0
 
11👁
r/LocalLLaMA · u/opUserZero · 11d ago
Jev mode for images! post image

So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17

Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV

▲
10
 
10👁
r/LocalLLaMA · u/SignificantZebra5883 · 11d ago
50B+ MoEs with few active parameters, what's the sweet spot for intelligence, agent speed, and affordable fine-tuning?

I’m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:

Question → many sequential tool calls → final answer/document

Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.

I’ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. I’m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.

The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.

Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?

I understand that small active parameter counts don’t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also don’t assume that more total parameters automatically means a better model. I’m interested in where that tradeoff works in practice.

There are three things I’m trying to pin down:

  • Inference hardware: Ideally inference runs rented with parallel agentic loops (this is for a B2C project, not single person use, we scale based on demand)
  • Fine-tuning hardware: Obviously FT LoRA will take more memory than inference, max like 4GPUs on vastai fits the budget.
  • Agent performance: After it gets the prompt the tool calls and everything will be local, so imo it has no problems being blazing fast, as soon as the model calls a tool call it will be back very fast, so for this agentic use case, quick TTFT and t/s and adaptive dynamic reasoning are prefer right?

For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.

I’m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).

I’m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like “model + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,” along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.

Has moving from a \~30B model to a 50B+ low-active-parameter MoE actually improved your agent’s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading

💬 22 (+1) open on reddit ↗
▲
8
 
15👁
r/LocalLLaMA · u/sToeTer · 11d ago
Is there even an easy, seamless vision assistant program?

I read textbooks on my PC and ideally i want a program with a normal chat environment where i can just hammer in questions about what's currently on my screen. Example: I'm working on a PDF, underline or circle things... and then just type "what does this sentence mean?", you get it.

I do NOT want to manually screenshot, navigate to the folder, drag the picture into the environment and then also have to type the question. It should also naturally be aware that the conversation is about what's on the screen, so i don't have to steer it with "make a screenshot; use your vision capabilites" etc.

I tried multiple different MCP in LM Studio, none of them were great...or worked :/

Someone said AnythingLLM has this function but i couldn't find it.

Is there a good solution?

Thank you in advance! :)

▲
0
 
17👁
r/LocalLLaMA · u/ChopSticksPlease · 11d ago
What would you buy for $5k...$10k USD? post image

What (and if) would you buy if you had $5k ... $10k ... $20k to spend on local AI?

So, I'm a contractor and a solo dev working on some products/saas/apps. Basically I usually run up to three cline/opencode sessions in the same time, long running software engineering tasks, often run out of 128k context, so 256k ctx is prefferable. Pretty much every day for multiple hours so I could burn quite a lot of $ daily on OpenRouter. Fortunately, since Qwen3.8 i rarely need to delegate to larger models like Kimi K3 or MiniMax M3.

Apart from code I often work on confidential documents so a local AI or an approved remote AI is a must.

My current AI setup is:
\- dev server with RTX3090 running Qwen3.8 UD Q4\_K\_XL with 128k ctx q8
\- lab server with 2x RTX3090 + 128gb ram running Qwen3.8 Flash Next with 256k ctx

Both machines are fine to run up to three sessions, one on dev and 2 concurrent 128k ctx tasks on the lab server. The performance i get from the 2x RTX3090 with Qwen3.8 Flash Next is close to a single DGX Spark GB10 (according to numbers).

Soon I may need to run more agents and work with other people so started thinking of an upgrade.

Does it make sense to invest in either a single GB10 machine or two and cluster them to get more space for more context and therefore more concurrent sessions? Would you consider other options?

Any feedback appreciated.

💬 93 (+2) open on reddit ↗
▲
0
 
11👁
r/LocalLLaMA · u/Dany0 · 11d ago
Where are the Opus 5.5 datasets?

Another day, another refusal. Apparently asking opus "what are your thoughts on this?" is an attempt at a 'distillation attack'

Our precinct is hugging face, we work at breakneck speed, we're up against art thieves, code thieves, extortionists, we're on call around the clock. The people of LocalLLaMA -- our finetunes is our job (and we write our own emdashes thank you)

WHERE ARE THE DATASETS PEOPLE. What happened to us? We used to throw pies at Dario Altman and now, what, we're penniless, downtrodden, what happened?

▲
40
 
18👁
r/LocalLLaMA · u/jacek2023 · 11d ago
Ornith-1.5 DFlash

Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash

Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-397B-DFlash

Ornith-1.5-35B-A3B-DFlash pairs the Ornith-1.5-35B-A3B model with a DFlash draft model for speculative decoding.

https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B-DFlash

▲
0
 
8👁
r/LocalLLaMA · u/emperorofrome13 · 11d ago
Using unsloth I created the worlds best 9B model post image

#

https://huggingface.co/emperorofrome/Gmcoder

Beats Ornith 1.5 and Oxcoder on HumanEval+ Mini — and does it without the overthinking. It gets to the answer using 40–68% fewer tokens. Built as a finetuned merge.

Edit: Best in the world is just hype. The coder is comparable to Ornith 1.5 but more token efficient by about 50% on average up to 78% at times and can be faster.

▲
2
 
12👁
r/LocalLLaMA · u/Forward_Compute001 · 11d ago
Cheapest Epyc 7003 (Milan) Bundle (ddr4)?

I'm building a new rig to host the mission control application that should sit on its own node and I immediatly thought of a cheap single socket ddr4 solution,

does anyone have some suggestions which bundle is cheapest or gives best value for price...?

\-no need for gpus

\-no need for much ram (8gb ram sticks)

\-many threads and max core speed would be important (maybe if it doesnt spike the price)

▲
11
-1
8👁
r/LocalLLaMA · u/Brilliant-Hall1387 · 10d ago
Sherry's 3:4 ternary format (1.375 bits per weight) running on WebGPU: a 1.6 MB model that plays Connect Four as well as its 7.8 MB int8 version

Not an LLM, but the ternary findings should carry over, and we hadn't seen Sherry-style 3:4 weights run in a browser before. Disclosure: this is our work at Precisit, everything is MIT.

What it is

  • A 7.4M-parameter one-pass scorer (the jevlike family): the board goes in, one score per legal column comes out. No search.
  • Weights in T34, Sherry's 3:4 format: in every four weights one is zero and three are ±1, so four weights fit in 5 bits. One fp16 scale per 128 weights gives 1.375 bits per weight. The embedding is int8; norms and biases are fp16.
  • It runs in the browser on a small WebGPU runtime: 1.1 ms per move (idle M5 Pro, Chrome).

|Model|File size|vs depth-4 bot|vs depth-6 bot|
|:-|:-|:-|:-|
|dense (fp32)|29.7 MB|0.92|0.89|
|T34, trained ternary|1.59 MB|0.93|0.91|
|T34, fine-tuned from dense|1.59 MB|0.89|0.9|
|T34, converted after training|1.59 MB|0.13|0.11|
|Base243 (TQ1\_0 style), trained|1.93 MB|0.89|0.88|

200 games each, both sides play a random move 5% of the time, a win counts 1 and a draw ½.

What we learned

  1. Converting the finished model to 3:4 collapsed it (0.13 against the depth-4 bot). Training with the format in the forward pass fixed it completely, whether from scratch or fine-tuning.
  2. Attention's q/k/v matrices are the sensitive ones. Group size (64/128/256) barely mattered.
  3. Seeds matter: two runs of the same T34 recipe scored 0.945 and 0.882.

Play it:
https://precisit.github.io/onepass-web/demo/c4-size/

Code, models, every result:
https://github.com/precisit/onepass-webgpu-ternary

The write-up:
https://precisit.com/en/blog/onepass-c4-size/

Has anyone gotten post-training 3:4 conversion to work on models, or does it need training?

▲
17
-1
8👁
r/LocalLLaMA · u/giveen · 10d ago
exl3 now in ninfer-ext

My ninfer-ext fork of the famous ninfer now supports exl3 , and have released two models as well. https://huggingface.co/jabbatheduck/ninfer-ext-models

▲
6
-1
17👁
r/LocalLLaMA · u/HitarthSurana · 10d ago
Any good harness or tools for helping students?

What FOSS/self-hosted tools or AI tools have you found genuinely helpful for students? Also interested in anything non-AI that has helped with studying, notes, organization, research, etc. Curious what you guys actually use or found useful. (written fully by human)

💬 17 (+1) open on reddit ↗
▲
30
-2
19👁
r/LocalLLaMA · u/rorowhat · 10d ago
Best model for blender?

Trying to see if I can use a local model to generate game assets, or even 3D printer models. Any suggestions? Something that would fit in 64GB of ram, speed is not an issue. Just need it to work well.

💬 50 (-1) open on reddit ↗
▲
43
-6
20👁
r/LocalLLaMA · u/Haunting-Stretch8069 · 11d ago
Qwen 3.8 27B Q4 with 100K context on a 16 GB RX 7800 XT guide

I'm running Qwen 3.8 27B Q4 XS with \~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.

Build llama.cpp with Vulkan:

cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j

Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then:

llama-server \
--model Qwen3.8-27B-UD-IQ4_XS.gguf \
--mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \
--n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \
--batch-size 2048 --ubatch-size 512 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \
--load-mode none --fit off \
--cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \
--no-context-shift --jinja --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080

---

Edit: The process is documented here: https://zenodo.org/records/23088880. Feedback will be integrated into the upcoming Qwen 4 setup.