15 posts · 1 sub · RSS
← prev Friday, September 4, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
94
-3
16👁
r/LocalLLaMA · u/sugarfreecaffeine · 36d ago
If you had ~15k would you build a home server today or wait

Title help me decide and avoid making impulse purchases 😩 I already have dual 3090 which I can sell to help EDIT: ty all I’ll just wait it out, doesn’t seem worth it right now

▲
334
+5
15👁
r/LocalLLaMA · u/Storterald · 36d ago
I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM

After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code (C code), the results were not completely unexpected but some quants were definitely underwhelming.

TLDR: Best overall: bartowski/Qwen3.8-27B-IQ4_XS. Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS. For a bit more context: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S or uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp

  • edit1: added TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS, bartowski/Qwen3.8-27B-IQ3_XS, prism-ml/Ternary-Bonsai-27B-Q2_g64 and magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unsloth
  • edit2: added unsloth/Qwen3.8-27B-UD-IQ3_S and ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S
  • edit3: added magiccodingman/Qwen3.8-27B-MQ-IQ2_M_1 and huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3_S
  • edit4: added AtomicChat/Qwen3.8-27B-AD-IQ4_XS-IQ3_S
  • edit5: added IvanKrastevAdventics/qwen3.8-27b-awq-int4-q4_0 (gguf of cyankiwi/qwen3.8-27b-awq-int4)
  • edit6: added the updated bartowski/Qwen3.8-27B-IQ4_XS and bartowski/Qwen3.8-27B-IQ3_XXS
  • edit7: added bartowski/Qwen3.8-27B-Q3_K_M, Thireus/09ae8ba_22b6bb2 and Thireus/09ae8ba_248b31b
  • edit8: removed all MTP heads from the GGUF size for a more fair comparison. added bartowski/Qwen3.8-27B-IQ2_S
  • edit9: added huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp
  • edit10: added mradermacher/Eintopf-Qwen3.8-27B.i1-IQ3_M and hitsfmdj/Qwen3.8-27B-4.2BPW-16GB
  • edit11: added Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3_S-recovered
  • edit12: added Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw and turboderp/SC_3.00bpw_H4_V4
  • edi13: added prism-ml/Ternary-Bonsai-2-27B-PQ2_0 and prism-ml/Ternary-Bonsai-2-27B-PTQ1_0, replaced prism-ml/Ternary-Bonsai-27B-Q2_g64 with prism-ml/Ternary-Bonsai-27B-PQ2_0
  • edit14: added Bucoid/Qwen3.8-27B-Heretic-Ara-iq4_xs-3.0
  • edit15: added agentionai/Qwen3.8-27B-AP-IQ3_S, agentionai/Qwen3.8-27B-AP-IQ4_XS and RonnieOps/Qwen3.8-27B-IQ4_XS-fullvocab-E3. This is probably the final edit as Qwen4 27B will likely come out soon.
  • edit16: added byteshape/Qwen3.8-27B-IQ4_XS-3.84bpw and byteshape/Qwen3.8-27B-IQ3_S-3.23bpw. Again, probably the last edit.

(sorted by Mean KLD)

|Model|Mean KLD|Same top p|GGUF size (without MTP)|
|:-|:-|:-|:-|
|prism-ml/Ternary-Bonsai-27B-PQ2\_0|1.289582 ± 0.008684|82.849 ± 0.118 %|6.7GiB|
|prism-ml/Ternary-Bonsai-2-27B-PQ2\_0|1.096134 ± 0.007705|84.596 ± 0.113 %|6.7GiB|
|prism-ml/Ternary-Bonsai-2-27B-PTQ1\_0|1.095914 ± 0.007703|84.582 ± 0.113 %|5.5GiB|
|sdkyuan/qwen38-27b-qat-q2\_0|0.893177 ± 0.006948|85.727 ± 0.110 %|8.2GiB|
|bartowski/Qwen3.8-27B-IQ2\_Sbartowski/Qwen3.8-27B-IQ2\_S (NEW)|0.784060 ± 0.006457|87.016 ± 0.105 %|8.7GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2\_XS|0.767174 ± 0.006291|86.166 ± 0.108 %|7.8GiB|
|TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS|0.514311 ± 0.004864|89.023 ± 0.098 %|8.9GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2\_S|0.512614 ± 0.004909|88.802 ± 0.099 %|8.6GiB|
|empero-ai/Qwen3.8-27B-Ridge-3.7bpw|0.475767 ± 0.004483|89.612 ± 0.096 %|11.4GiB|
|magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4\_K\_S-Unsloth|0.419585 ± 0.004076|89.661 ± 0.095 %|13.1GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3\_XXS|0.379222 ± 0.003992|90.270 ± 0.093 %|9.4GiB|
|unsloth/Qwen3.8-27B-UD-Q2\_K\_XL (UD2)|0.350861 ± 0.003745|90.626 ± 0.091 %|9.6GiB|
|byteshape/Qwen3.8-27B-IQ3\_S-3.23bpw|0.345563 ± 0.003703|90.868 ± 0.090 %|10.1GiB|
|mradermacher/Eintopf-Qwen3.8-27B.i1-IQ3\_M|0.318143 ± 0.003139|91.471 ± 0.087 %|11.7GiB|
|bartowski/Qwen3.8-27B-IQ3\_XXS (NEW)|0.300480 ± 0.003345|91.511 ± 0.087 %|11.3GiB|
|unsloth/Qwen3.8-27B-UD-IQ3\_XXS (UD2)|0.268594 ± 0.002971|91.951 ± 0.085 %|10.8GiB|
|magiccodingman/Qwen3.8-27B-MQ-IQ2\_M\_1|0.256808 ± 0.002861|92.056 ± 0.085 %|10.9GiB|
|DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-MTP-IQ3\_M|0.251270 ± 0.002702|92.315 ± 0.083 %|13.1GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3\_S|0.249650 ± 0.002841|92.016 ± 0.085 %|10.8GiB|
|bartowski/Qwen3.8-27B-IQ3\_XS (OLD)|0.238656 ± 0.002627|92.312 ± 0.083 %|12.2GiB|
|hitsfmdj/Qwen3.8-27B-4.2BPW-16GB|0.222090 ± 0.002570|92.552 ± 0.082 %|11.7GiB|
|esatapedico/Qwen3.8-27B-NVFP4-MTP-LOW|0.220796 ± 0.002631|92.339 ± 0.083 %|14.3GiB|
|unsloth/Qwen3.8-27B-UD-IQ3\_S (UD3)|0.218522 ± 0.002591|92.399 ± 0.083 %|10.9GiB|
|turboderp/SC\_3.00bpw\_H4\_V4 (exllama3)|0.205712 ± 0.002509|92.573 ± 0.082 %|11.9GiB|
|jrell/Qwen3.8-27B-i1-IQ4\_XS-GGUF-Smaller|0.194459 ± 0.002242|93.049 ± 0.080 %|12.4GiB|
|agentionai/Qwen3.8-27B-AP-IQ3\_S|0.193041 ± 0.002352|92.995 ± 0.080 %|10.9GiB|
|orcarouter/Qwen3.8-27B-Uncensored-Q3\_K\_L|0.192312 ± 0.002294|92.726 ± 0.081 %|13.4GiB|
|bartowski/Qwen3.8-27B-Q3\_K\_M (NEW)|0.191103 ± 0.002369|92.823 ± 0.081 %|12.3GiB|
|mudler/Qwen3.8-27B-APEX-I-Mini|0.190209 ± 0.002354|93.012 ± 0.080 %|12.6GiB|
|Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3\_S-recovered|0.178882 ± 0.002102|93.110 ± 0.079 %|11.0GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3\_S-mtp|0.178715 ± 0.002200|92.949 ± 0.080 %|11.0GiB|
|Thireus/09ae8ba\_22b6bb2 (ikllama.cpp quality 41.39%)|0.178290 ± 0.002202|93.115 ± 0.079 %|11.0GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3\_S|0.175223 ± 0.002129|93.024 ± 0.080 %|11.0GiB|
|Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw (exllama3)|0.149827 ± 0.001936|93.735 ± 0.076 %|14.1GiB|
|unsloth/Qwen3.8-27B-UD-Q3\_K\_XL (UD2)|0.147186 ± 0.001809|93.734 ± 0.076 %|12.2GiB|
|unsloth/Qwen3.8-27B-UD-Q3\_K\_XL (UD3)|0.142647 ± 0.001860|93.789 ± 0.076 %|11.9GiB|
|byteshape/Qwen3.8-27B-IQ4\_XS-3.84bpw|0.131976 ± 0.001690|93.852 ± 0.075 %|12.0GiB|
|IvanKrastevAdventics/qwen3.8-27b-awq-int4-q4\_0|0.112990 ± 0.001558|94.171 ± 0.073 %|14.4GiB|
|AtomicChat/Qwen3.8-27B-AD-IQ4\_XS-IQ3\_S|0.111713 ± 0.001492|94.527 ± 0.071 %|13.2GiB|
|Bucoid/Qwen3.8-27B-Heretic-Ara-iq4\_xs-3.0|0.097572 ± 0.001341|94.651 ± 0.070 %|13.0GiB|
|Bucoid/Qwen3.8-27B-Uncensored-IQ4\_XS\_4BPW|0.091447 ± 0.001261|94.774 ± 0.070 %|12.8GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4\_XS|0.082871 ± 0.001205|94.981 ± 0.068 %|13.1GiB|
|unsloth/Qwen3.8-27B-UD-IQ4\_XS (UD3)|0.075626 ± 0.001097|95.258 ± 0.067 %|13.3GiB|
|agentionai/Qwen3.8-27B-AP-IQ4\_XS|0.073386 ± 0.001075|95.386 ± 0.066 %|13.0GiB|
|Thireus/09ae8ba\_248b31b (llama.cpp 49.75%)|0.063904 ± 0.000967|95.687 ± 0.064 %|13.3GiB|
|jpetrina/Qwen3.8-27B-IQ4\_XS-pure|0.061984 ± 0.000917|95.551 ± 0.065 %|13.3GiB|
|RonnieOps/Qwen3.8-27B-IQ4\_XS-fullvocab-E3|0.058278 ± 0.000886|95.773 ± 0.063 %|14.1GiB|
|bartowski/Qwen3.8-27B-IQ4\_XS (OLD)|0.056482 ± 0.000856|95.835 ± 0.063 %|14.3GiB|
|bartowski/Qwen3.8-27B-IQ4\_XS (NEW)|0.055415 ± 0.000849|95.850 ± 0.062 %|14.2GiB|
|unsloth/Qwen3.8-27B-UD-Q4\_K\_XL (UD3) *(can't fit)*|0.029844 ± 0.000476|96.921 ± 0.054 %|16.1GiB|
|unsloth/Qwen3.8-27B-UD-Q4\_K\_XL (UD2) *(can't fit)*|0.028026 ± 0.000432|96.988 ± 0.054 %|16.4GiB|

https://preview.redd.it/g9isjm0d04sh1.png?width=5355&format=png&auto=…

Hope this helps other VRAM starved people like me :)

▲
135
-1
13👁
r/LocalLLaMA · u/niacolhealth · 36d ago
Ling-3.0-flash-VL, built on Ling-3.0-flash with visual understanding and visual agent capabilities post image

It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.

▲
114
-4
15👁
r/LocalLLaMA · u/Tall_Abrocoma_3533 · 36d ago
Qwen3.8-Flash-Next on a phone CPU! post image

Like the title says, running completely locally on my Xiaomi 14T Pro device. Specific model: Qwen3.8-Flash-Next-UD-IQ3\_XXS App used: BigMoeOnEdge

▲
545
-3
18👁
▲
1443
 
19👁
r/LocalLLaMA · u/liright · 36d ago
You can now run a 90M conversational LLM on the Sony PSP (hardware from 2004). Doesn't get more local than this. post image

Github link: https://github.com/thatblend/LLMPSP

I wanted to see what the PSP can theoretically handle and I got my answer - a 90M model is about the max it can do without atrocious inference speeds. It's running around 0.5 - 0.6 tokens per second, which is very slow, but it's useable. Maybe 1-3 minutes for a reply.

The model is actually fairly impressive for 90M parameters, it's not really useful in any real metric, but it can generate crappy poems, short stories, write non-functional code and sometimes it gets things right if you ask it what company makes macbooks, what is an LLM etc, while other times it just hallucinates a crazy answer. Fun.

▲
63
+2
11👁
r/LocalLLaMA · u/FoxDeFleurs · 36d ago
Sometimes I be mourning the agents I get before context compacts

Just wanted to put that out there. It's like they get an ice pick to the brain no actual mourning here btw that'd be psychosis it's okay to laugh

▲
141
 
12👁
r/LocalLLaMA · u/TheLocalDrummer · 36d ago
Drummer's Artemis 31B v1 and v1.1 - Coming back with a bang!

Hey everyone, been a while!

https://huggingface.co/TheDrummer/Artemis-31B-v1.1

https://huggingface.co/TheDrummer/Artemis-31B-v1

A few months ago, Gemma graced us with models that served as a much needed downpour from a year-long drought. I'm so happy to see us thrive once again.

The difference between v1 and v1.1 is quite simple: v1 was an early attempt, an overdue release that excelled in prose and writing, while requiring some handholding to get over quirks like stuttering. v1.1 is a more refined approach where stability meets quality. My community is split, so I figured I'd just release both.

\---

I was gone for a while. I got busy dealing with life, both its ups and downs. While I couldn't attend to you folks, I've been lurking around and appreciating you all for the kind words.

\- Skyfall 31B v4.2 seems to be a banger for many of you. I'm proud of the upscale and consider it my ultimate home-run send-off for the beautiful Mistral 24B base. It's a shame that it was overshadowed by Gemma 31B's release, but hearing some of ya'll compare and even prefer it to a more modern base was an unexpected win.

\- Rocinante 12B X / 16B XL proves that Nemo is still the ultimate creative model to this day. For some to say that 16B XL felt like Cydonia 24B v4.3 just goes to show how far you can go with modern resources and techniques.

\- Anubis 70B v1.2, Valkyrie 49B v2.1, Anubis Mini 8B v1 surprised me too. I had zero expectations releasing them. Just like Rocinante X / XL, they are modern finetunes of old base models. And somehow, they still found their users singing praises.

\---

With the Artemis release taking weight off my shoulders, I'm eager to move on and tune a ton more bases!

But I have something else cooking: a HordeAI-like platform. I hope to provide value not just as a finetuner, but as a local lover too!

The premise is simple: it's a place where generous local hosters can share inference with the less fortunate. You'd be surprised how many power users would love to heat their rooms through the power of charity.

\---

Finally, I'd like to thank everyone who supported me over the years. From those who provided kind words, rigorous testing, compute access, inference, or cold hard cash. You've all granted me the ability to enrich the local ecosystem with fun experiments like Rivermind 12B, Fallen series, Big Tiger Gemma, Precog 24B/123B, and solid models like Cydonia 24B v4.3, Behemoth X 123B v2.x, and Skyfall 31B v4.2.

If you've got inference / compute credits to share, please contact me! It will all go to making the community happy <3

Backlog:

\- Gemma E2B

\- Gemma E4B

\- Gemma 12B

\- Gemma 26BA4B

\- Qwen 3.8 27B

\- Muse Glimmer 30B

\- Mistral Medium 3.5 128B

\- HordeAI Alternative / Crowdsourced 'OpenRouter' ("BeaverNet")

▲
266
 
17👁
r/LocalLLaMA · u/myreala · 37d ago
I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked?

This epyc server I am using twelve cards with 64 GB memory, plus 256GB ram. Looking at the most capable models in open source, GLM 5.3 seems to be the only option, but with Astra releasing it will likely be fairly behind. GLM6 looks like it will be at least double in size, maybe even triple. Qwen-max and Kimmi are already way too big to even consider. Even the deepseek V4 Pro is too big. Should I just give up on this frontier dream sell the excess GPUs and settle For flash models with far fewer GPUs and a reasonable cost. Note: I'm not using it for any business.

I was hoping to build a new business with this, but it can probably be done with much more effort with a flash model as well.

Edit: I don't want to go below 4-bit quants because then the models start making obvious mistakes. So I'm talking about a min/max of 4-bit

Okay, this post really blew up. I wasn't expecting so much interest or comments just attacking me. Was really just expecting to have a calm discussion about future SOTA model sizes.

▲
64
-2
17👁
r/LocalLLaMA · u/zRevengee · 37d ago
Qwen 3.8 Flash Next Can Build Funny Games post image

This is nothing impressive but, i had so much fun i wanted to share my experience with this model.

(yes this post is written by human)

I made an FPS with local Q4\_K\_XL 3.8 Flash Next (256k context) (it took 3 days to refine everything but playable demo was ready in 2 hours) to play with friends.

had ton of fun talking with them about what could we add , funny features etc.

Features:

  • \- toggle retro psx shader
  • \- totally destructible environments
  • \- tac sprint
  • \- tilting with Q and E for peaking from corners.
  • \- free for all modes, SnD, Swords Only (swords have animations when slashing), RPG only, team deathmatch
  • \- killfeed, map with red dots when a player shoot
  • \- bunny hop
  • \- day and night cicle with rain or snow
  • \- fov slider / shader intensity slider
  • \- hide n seek mode

I used opencode as harness, gun models were taken from sketchfab , model was running at 20tok/s avg with MTP, i know for someone is bad, but it did most of the work meanwhile i was at work or while sleeping, checking every now and then with a remote KVM from phone.

My machine:

5900x / 128GB DDR4 3200Mhz / RTX 5090 and RTX 4000 PRO (32 + 24 GB)

What games you would like to build in free time with ai? roguelites? 2d platforms? racing games?

Or did you already built something? share with some screenshots

▲
93
+4
12👁
r/LocalLLaMA · u/fairydreaming · 37d ago
MINISFORUM MS-S1 MAX-P495
€7??? Surprise Price Ends with Limited Stock That's likely 7999 EUR, so double of the initial price of MS-S1 MAX-128GB? 😭
▲
2682
+6
22👁
▲
62
+3
12👁
r/LocalLLaMA · u/saltexx · 37d ago
We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0)

I'm one of the developers. We said in August it would go open source in September and it did last night. MIT or Apache-2.0, pick one. The repo you see is our internal repo, kernels included, so from now on everything happens in public.

It's an inference engine in Rust and C++ with our own CUDA kernels. One binary with OpenAI and Anthropic style APIs, loads GGUF and safetensors. We run about 300B tokens a year through it at work.

Some numbers: Qwen3.8-27B FP8 on one RTX PRO 6000, spec decoding off on every engine:

  • vs vLLM faster in 13 of 13 cells, 1.02x to 1.19x (so not huge)
  • vs SGLang faster in 10 of 13, behind in 2, level in 1
  • vs llama.cpp Q8_0 faster in 13 of 13, 1.5x to 37x
  • 32 clients at 1024 in / 1024 out: 1062 tok/s, vLLM 958, SGLang 844

Full board with the losses: https://truespar.com/paddock/benchmarks/qwen38-27b

What it does not do yet: No Mac, no ROCm, no Vulkan.

One model per GPU, no tensor parallel. CUDA only, Windows and Linux. Validated on Blackwell (5090, RTX PRO 4500/5000/6000, B200) and Ampere (an A6000 was the bring-up card, 30-series works). Ada kernels ship but nobody has run a board on them so the engine refuses to start unless you set PADDOCK_UNVALIDATED_ARCH=1. Hopper and A100 kernels are in the tree without a board.

https://github.com/truespar/paddock

Thankful for any help and input!

▲
120
+3
13👁
r/LocalLLaMA · u/edward-dev · 37d ago
Has anyone already tried IFM's new K2-Horizon-MoVA-36B-A4B?

How good/bad is it against comparable MoEs the same size? How does it compare against Qwen 3.6 35BA3B?

Since we don't have 3.8 35B this seems like an upgrade if we look at some benchmarks like terminal bench, but they don't have SWE bench pro on the benchmarks table, and i don't really know anything about this lab, I'm wondering if it trades blows with models like tiel coder or if it's some benchmaxxed model like ornith?

At a single glance it looks really decent but haven't tried it in depth yet. What are your experiences with this model so far guys?

▲
65
+1
12👁
r/LocalLLaMA · u/Extension-Bid-639 · 37d ago
UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build

This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM, everything else on the GPUs. Since then I switched quants, stacked MTP on top of the cache, fixed the load time, found a bug in the cache PR and found out my RAM was thermal throttling (Now i gotta buy an additional case fan lol). Numbers are all from the same 4,000-token python coding prompt with thinking on unless stated otherwise.

Where it's at now

|Starting numbers (UD-Q6\_K\_XL, 4+4 resident layers)|First post (Q6 + cache, 135 slots)|Now (UD-Q4\_K\_XL + cache 188 slots + n-gram draft)|Now (Q4 + cache 150 slots + MTP)|
|:-|:-|:-|:-|
|decode, coding prompt with thinking|17|25-29|32-35|37-41|
|decode, code emission, thinking off|\-|24|37|49|
|decode at 131k depth|12|17|18-20|14-16|
|prefill, 26k prompt (ub 512)|\~350 at ub 2048|138|180-195|180-195|
|load to ready|\~13 min|8.5 min|2 min|2 min|
|host RAM for the experts|104 GB pinned + 51 GB PLE|same|73 GB pinned + 28 GB PLE|same|
|cache hit rate|\-|84-85%|90-92%|84-85% (fewer slots)|

Hit rate is the cache's own counter, decode is llama-server's eval time.

What changed, in order of payoff

  1. UD-Q4\_K\_XL instead of Q6\_K\_XL. Hit rate doesn't depend on the quant, only on slot count (Q4 at 135 slots: 84.7%, Q6 at 135: 84-85%). But what Q4 buys me is precious vram space, roughly 1.44x slots per GB of VRAM. So 188 slots actually fit where 135 did and achieved a hit rate 90-92%, increased decode from 27 to 32-35, prefill by +35% (fewer bytes per ubatch). The quality cost per unsloth's table is: KLD 0.047 vs 0.027, top-1 agreement 92.3% vs 94.1%; proper eval still to do. Host RAM drops to \~105 GB, so 128 GB is enough for this setup.
  2. MTP on top of the cache (mainline PR #28243, the unsloth MTP head). Yesterday I kind of concluded that "MTP does not pay" but looking back, that was the old fork with the cache off during verify. On the mainline, with the cache taking verify batches (see 4), MTP drafts every step at 50-58% acceptance on reasoning text and 94% on code emission. Decode went from 32-35 -> 37-41 t/s on the thinking prompt (single runs spread about 8% on this prompt at temp 0.7) and from 37 -> 49 on code emission. The draft head sits on the second GPU and costs about 4.5 GB, which is why the slots dropped from 188 to 150 on the table if you're wondering. Still a clear win at short context.
  3. Load 8.5 min -> 2 min. The loader was pulling 100 GB through page faults at 236 MB/s (MADV\_RANDOM under --numa distribute). Reading host-destination tensors straight from the file fixed it, that is in my PR #28223.
  4. A bug in the cache PR at n\_tokens > 1. \#27861 maps every uncached expert to one dummy slot, and the batched CUDA mul\_mat\_id kernels assume distinct ids per token: out-of-bounds writes. Only the mmvq path is safe, which quantized experts use up to 8 tokens, so the branch gates the cache at 8 tokens and keeps MTP's verify batch at 4. Repro and details in my #27861 comment: Link to comment
  5. My RAM was thermal throttling. This is more of a me issue but putting it out there for those who may have a similar box to mine. I experienced a slowdown after a few minutes of decode and the issue was the memory controller throttling once the hottest LRDIMM hit 78 C (perf stat -e unc_m_power_critical_throttle_cycles shows it, so don't worry if you have no BMC). A fan on the DIMM banks does keep it at 44-57 C, zero throttling, 16k-token runs from 10-15 to 24.6 t/s average. I would check this before touching software if your DDR4 Xeon box slows down under sustained load. This does not affect the numbers in the table and in my last post.

Did nothing or hurt here: q8\_0 KV (-18% at 131k depth), mirror-NUMA #27986, QSA gather #28213, --load-mode none, chained drafts, the ik\_llama GEMV port, more than 2 cache uploads per step, thread/poll/prio flags. --lazy-mode on-direct (#28136) gives +7-12% only on the first long prompt after a restart.

To replicate

Branch with everything: https://github.com/Inovello/llama.cpp/tree/flashnext-2x3090. It is master (b96806d) + PR #27861 (expert cache) + PR #28223 (pinned host experts under mmap + the load fix) + PR #28243 (MTP) + the mul\_mat\_id fix and the 8-token cache gate from my #27861 comment. Squashed into one commit, I added the credit in the commit message. It is a replication branch.

git clone -b flashnext-2x3090 https://github.com/Inovello/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build -j -t llama-server

LLAMA_ATTN_ROT_DISABLE=1 numactl --interleave=all build/bin/llama-server \
-m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
-md mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf --spec-type draft-mtp -devd CUDA1 --spec-draft-n-max 3 \
-ngl 99 -c 261888 --parallel 1 -fa on \
-ot "ffn_(gate|up|down)_exps\.weight=CUDA_Host,per_layer_token_embd\.weight=CPU" \
-lzm off --numa distribute -t 16 -tb 44 -b 4096 -ub 512 -ctk f16 -ctv f16 \
--moe-expert-cache 150 -lv 4

  • The MTP head is MTP/mtp-Qwen3.8-Flash-Next-shared-Q8\_0.gguf (2.79 GB) from the unsloth/Qwen3.8-Flash-Next-GGUF repo on HF. To clarify, the shared file has no embeddings of its own, the PR borrows them from the target, so it only loads as -md of the main model.
  • Slot sizing on Q4: \~75 MB per slot per GPU at full 261k context and ub 512. Without MTP I fit 188 slots (\~1 GB free per GPU); with the draft head on CUDA1, 150. Watch nvidia-smi after a long prompt, the CUDA pool grows \~350 MB during a 131k prefill.
  • -lv 4 prints the cache hit rate every 512 steps (moe-cache: ... hit-rate=) and the draft acceptance per request.
  • For sessions that you believe would reach high ctx usage, swap the three MTP flags for --spec-type ngram-map-k --spec-ngram-map-k-size-m 7 and raise the cache to 188.
  • Both PRs are drafts. #28243 has open review comments and #27861 has the bug above. The branch above is what actually runs here today, and it isn't something I would call finished.
  • For the single GPU brothers out there, same idea, just put -devd on your single GPU or skip MTP and take the slots.

Next thing I'll be doing is a proper comparison against Qwen3.8-27B at Q8 (speed and quality), since that is what most people here actually want to know. Happy to answer questions on any of it.