From: Polymarket on 𝕏: https://x.com/Polymarket/status/2095646821485842805 Julien Chaumond on 𝕏: https://x.com/julien\_c/status/2095822387895824836
From: Polymarket on 𝕏: https://x.com/Polymarket/status/2095646821485842805 Julien Chaumond on 𝕏: https://x.com/julien\_c/status/2095822387895824836
After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code (C code), the results were not completely unexpected but some quants were definitely underwhelming.
TLDR: Best overall: bartowski/Qwen3.8-27B-IQ4_XS. Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS. For a bit more context: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S or uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp
TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS, bartowski/Qwen3.8-27B-IQ3_XS, prism-ml/Ternary-Bonsai-27B-Q2_g64 and magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unslothunsloth/Qwen3.8-27B-UD-IQ3_S and ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_Smagiccodingman/Qwen3.8-27B-MQ-IQ2_M_1 and huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3_SAtomicChat/Qwen3.8-27B-AD-IQ4_XS-IQ3_SIvanKrastevAdventics/qwen3.8-27b-awq-int4-q4_0 (gguf of cyankiwi/qwen3.8-27b-awq-int4)bartowski/Qwen3.8-27B-IQ4_XS and bartowski/Qwen3.8-27B-IQ3_XXSbartowski/Qwen3.8-27B-Q3_K_M, Thireus/09ae8ba_22b6bb2 and Thireus/09ae8ba_248b31bbartowski/Qwen3.8-27B-IQ2_Shuihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtpmradermacher/Eintopf-Qwen3.8-27B.i1-IQ3_M and hitsfmdj/Qwen3.8-27B-4.2BPW-16GBJoakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3_S-recoveredMia-AiLab/Qwen3.8-27B-EXL3-3.5bpw and turboderp/SC_3.00bpw_H4_V4prism-ml/Ternary-Bonsai-2-27B-PQ2_0 and prism-ml/Ternary-Bonsai-2-27B-PTQ1_0, replaced prism-ml/Ternary-Bonsai-27B-Q2_g64 with prism-ml/Ternary-Bonsai-27B-PQ2_0Bucoid/Qwen3.8-27B-Heretic-Ara-iq4_xs-3.0agentionai/Qwen3.8-27B-AP-IQ3_S, agentionai/Qwen3.8-27B-AP-IQ4_XS and RonnieOps/Qwen3.8-27B-IQ4_XS-fullvocab-E3. This is probably the final edit as Qwen4 27B will likely come out soon.byteshape/Qwen3.8-27B-IQ4_XS-3.84bpw and byteshape/Qwen3.8-27B-IQ3_S-3.23bpw. Again, probably the last edit.(sorted by Mean KLD)
|Model|Mean KLD|Same top p|GGUF size (without MTP)|
|:-|:-|:-|:-|
|prism-ml/Ternary-Bonsai-27B-PQ2\_0|1.289582 ± 0.008684|82.849 ± 0.118 %|6.7GiB|
|prism-ml/Ternary-Bonsai-2-27B-PQ2\_0|1.096134 ± 0.007705|84.596 ± 0.113 %|6.7GiB|
|prism-ml/Ternary-Bonsai-2-27B-PTQ1\_0|1.095914 ± 0.007703|84.582 ± 0.113 %|5.5GiB|
|sdkyuan/qwen38-27b-qat-q2\_0|0.893177 ± 0.006948|85.727 ± 0.110 %|8.2GiB|
|bartowski/Qwen3.8-27B-IQ2\_Sbartowski/Qwen3.8-27B-IQ2\_S (NEW)|0.784060 ± 0.006457|87.016 ± 0.105 %|8.7GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2\_XS|0.767174 ± 0.006291|86.166 ± 0.108 %|7.8GiB|
|TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS|0.514311 ± 0.004864|89.023 ± 0.098 %|8.9GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2\_S|0.512614 ± 0.004909|88.802 ± 0.099 %|8.6GiB|
|empero-ai/Qwen3.8-27B-Ridge-3.7bpw|0.475767 ± 0.004483|89.612 ± 0.096 %|11.4GiB|
|magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4\_K\_S-Unsloth|0.419585 ± 0.004076|89.661 ± 0.095 %|13.1GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3\_XXS|0.379222 ± 0.003992|90.270 ± 0.093 %|9.4GiB|
|unsloth/Qwen3.8-27B-UD-Q2\_K\_XL (UD2)|0.350861 ± 0.003745|90.626 ± 0.091 %|9.6GiB|
|byteshape/Qwen3.8-27B-IQ3\_S-3.23bpw|0.345563 ± 0.003703|90.868 ± 0.090 %|10.1GiB|
|mradermacher/Eintopf-Qwen3.8-27B.i1-IQ3\_M|0.318143 ± 0.003139|91.471 ± 0.087 %|11.7GiB|
|bartowski/Qwen3.8-27B-IQ3\_XXS (NEW)|0.300480 ± 0.003345|91.511 ± 0.087 %|11.3GiB|
|unsloth/Qwen3.8-27B-UD-IQ3\_XXS (UD2)|0.268594 ± 0.002971|91.951 ± 0.085 %|10.8GiB|
|magiccodingman/Qwen3.8-27B-MQ-IQ2\_M\_1|0.256808 ± 0.002861|92.056 ± 0.085 %|10.9GiB|
|DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-MTP-IQ3\_M|0.251270 ± 0.002702|92.315 ± 0.083 %|13.1GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3\_S|0.249650 ± 0.002841|92.016 ± 0.085 %|10.8GiB|
|bartowski/Qwen3.8-27B-IQ3\_XS (OLD)|0.238656 ± 0.002627|92.312 ± 0.083 %|12.2GiB|
|hitsfmdj/Qwen3.8-27B-4.2BPW-16GB|0.222090 ± 0.002570|92.552 ± 0.082 %|11.7GiB|
|esatapedico/Qwen3.8-27B-NVFP4-MTP-LOW|0.220796 ± 0.002631|92.339 ± 0.083 %|14.3GiB|
|unsloth/Qwen3.8-27B-UD-IQ3\_S (UD3)|0.218522 ± 0.002591|92.399 ± 0.083 %|10.9GiB|
|turboderp/SC\_3.00bpw\_H4\_V4 (exllama3)|0.205712 ± 0.002509|92.573 ± 0.082 %|11.9GiB|
|jrell/Qwen3.8-27B-i1-IQ4\_XS-GGUF-Smaller|0.194459 ± 0.002242|93.049 ± 0.080 %|12.4GiB|
|agentionai/Qwen3.8-27B-AP-IQ3\_S|0.193041 ± 0.002352|92.995 ± 0.080 %|10.9GiB|
|orcarouter/Qwen3.8-27B-Uncensored-Q3\_K\_L|0.192312 ± 0.002294|92.726 ± 0.081 %|13.4GiB|
|bartowski/Qwen3.8-27B-Q3\_K\_M (NEW)|0.191103 ± 0.002369|92.823 ± 0.081 %|12.3GiB|
|mudler/Qwen3.8-27B-APEX-I-Mini|0.190209 ± 0.002354|93.012 ± 0.080 %|12.6GiB|
|Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3\_S-recovered|0.178882 ± 0.002102|93.110 ± 0.079 %|11.0GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3\_S-mtp|0.178715 ± 0.002200|92.949 ± 0.080 %|11.0GiB|
|Thireus/09ae8ba\_22b6bb2 (ikllama.cpp quality 41.39%)|0.178290 ± 0.002202|93.115 ± 0.079 %|11.0GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3\_S|0.175223 ± 0.002129|93.024 ± 0.080 %|11.0GiB|
|Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw (exllama3)|0.149827 ± 0.001936|93.735 ± 0.076 %|14.1GiB|
|unsloth/Qwen3.8-27B-UD-Q3\_K\_XL (UD2)|0.147186 ± 0.001809|93.734 ± 0.076 %|12.2GiB|
|unsloth/Qwen3.8-27B-UD-Q3\_K\_XL (UD3)|0.142647 ± 0.001860|93.789 ± 0.076 %|11.9GiB|
|byteshape/Qwen3.8-27B-IQ4\_XS-3.84bpw|0.131976 ± 0.001690|93.852 ± 0.075 %|12.0GiB|
|IvanKrastevAdventics/qwen3.8-27b-awq-int4-q4\_0|0.112990 ± 0.001558|94.171 ± 0.073 %|14.4GiB|
|AtomicChat/Qwen3.8-27B-AD-IQ4\_XS-IQ3\_S|0.111713 ± 0.001492|94.527 ± 0.071 %|13.2GiB|
|Bucoid/Qwen3.8-27B-Heretic-Ara-iq4\_xs-3.0|0.097572 ± 0.001341|94.651 ± 0.070 %|13.0GiB|
|Bucoid/Qwen3.8-27B-Uncensored-IQ4\_XS\_4BPW|0.091447 ± 0.001261|94.774 ± 0.070 %|12.8GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4\_XS|0.082871 ± 0.001205|94.981 ± 0.068 %|13.1GiB|
|unsloth/Qwen3.8-27B-UD-IQ4\_XS (UD3)|0.075626 ± 0.001097|95.258 ± 0.067 %|13.3GiB|
|agentionai/Qwen3.8-27B-AP-IQ4\_XS|0.073386 ± 0.001075|95.386 ± 0.066 %|13.0GiB|
|Thireus/09ae8ba\_248b31b (llama.cpp 49.75%)|0.063904 ± 0.000967|95.687 ± 0.064 %|13.3GiB|
|jpetrina/Qwen3.8-27B-IQ4\_XS-pure|0.061984 ± 0.000917|95.551 ± 0.065 %|13.3GiB|
|RonnieOps/Qwen3.8-27B-IQ4\_XS-fullvocab-E3|0.058278 ± 0.000886|95.773 ± 0.063 %|14.1GiB|
|bartowski/Qwen3.8-27B-IQ4\_XS (OLD)|0.056482 ± 0.000856|95.835 ± 0.063 %|14.3GiB|
|bartowski/Qwen3.8-27B-IQ4\_XS (NEW)|0.055415 ± 0.000849|95.850 ± 0.062 %|14.2GiB|
|unsloth/Qwen3.8-27B-UD-Q4\_K\_XL (UD3) *(can't fit)*|0.029844 ± 0.000476|96.921 ± 0.054 %|16.1GiB|
|unsloth/Qwen3.8-27B-UD-Q4\_K\_XL (UD2) *(can't fit)*|0.028026 ± 0.000432|96.988 ± 0.054 %|16.4GiB|
https://preview.redd.it/g9isjm0d04sh1.png?width=5355&format=png&auto=…
Hope this helps other VRAM starved people like me :)
Used qwen3.8-27b in Opencode to make this silly mini-game because I'm not sober:
```
We are going to play a game, it will be the Wikipedia game. The Wikipedia game has the following rules:
Your objective is to reach the the end point, which is an article completely separate from the starting point article.
Your only constraints are the following:
Use playwright to click the links.
```
Basically, Qwen needs to reach an ending article within 10 Wikipedia hyperlink clicks from the starting article, which is usually an unrelated article. It needs to use playwright (or some equivalent browser MCP) to click the Wikipedia hyperlinks without backtracking, using search or using external links.
I verified the links for accuracy and I can confirm it managed to complete this task within 6 turns. Thought it would get stuck in a loop. Its a dumb minigame but I think its a good, simple agent test to perform.
Ling 3.0 Tiny still seems to be leading the pack despite only having 1.3B active
€7??? Surprise Price Ends with Limited Stock That's likely 7999 EUR, so double of the initial price of MS-S1 MAX-128GB? 😭
How good/bad is it against comparable MoEs the same size? How does it compare against Qwen 3.6 35BA3B?
Since we don't have 3.8 35B this seems like an upgrade if we look at some benchmarks like terminal bench, but they don't have SWE bench pro on the benchmarks table, and i don't really know anything about this lab, I'm wondering if it trades blows with models like tiel coder or if it's some benchmaxxed model like ornith?
At a single glance it looks really decent but haven't tried it in depth yet. What are your experiences with this model so far guys?
Otaku is an LLM frontend, primarily designed for roleplay, an alternative to SillyTavern and the like. However, It also works for general-purpose chat with local backends (including Ollama) or cloud models, the way Open WebUI is used, once lore extraction is switched off in the settings.
Otaku offers two interfaces:
Both share the same functions; the difference is that in the terminal you execute them with slash commands (the reference is available with /help), while in the web UI the operations are available from the menu.
Install
Otaku is free and open source (MIT); it works on macOS, Linux and Windows. Install it with uv (uv tool install otaku) or see the GitHub README for other options: https://github.com/enclavum/otaku
Get started
Launch either otaku for the terminal or otaku web for the web UI; the web UI's default URL is http://localhost:9600. Two sample stories are imported on first start to give you an idea of the features and what play looks like, and you land right in the middle of one of them.
On first start, you choose a provider and a model: Otaku automatically detects local installations of Ollama, oMLX, LM Studio, llama.cpp and KoboldCpp, and lets you pick from their models. Cloud providers (OpenRouter, NanoGPT) are also there: enter an API key and their catalogs appear. After exploring the provided stories, you can start your own with the /new command.
Asking for feedback
Otaku is a personal side project, and I'd like to get feedback from the community on the product and on what to add.
I'm one of the developers. We said in August it would go open source in September and it did last night. MIT or Apache-2.0, pick one. The repo you see is our internal repo, kernels included, so from now on everything happens in public.
It's an inference engine in Rust and C++ with our own CUDA kernels. One binary with OpenAI and Anthropic style APIs, loads GGUF and safetensors. We run about 300B tokens a year through it at work.
Some numbers: Qwen3.8-27B FP8 on one RTX PRO 6000, spec decoding off on every engine:
Full board with the losses: https://truespar.com/paddock/benchmarks/qwen38-27b
What it does not do yet: No Mac, no ROCm, no Vulkan.
One model per GPU, no tensor parallel. CUDA only, Windows and Linux. Validated on Blackwell (5090, RTX PRO 4500/5000/6000, B200) and Ampere (an A6000 was the bring-up card, 30-series works). Ada kernels ship but nobody has run a board on them so the engine refuses to start unless you set PADDOCK_UNVALIDATED_ARCH=1. Hopper and A100 kernels are in the tree without a board.
https://github.com/truespar/paddock
Thankful for any help and input!
Anyone who has a 3D printer and get use of it finds it incredibly useful for those odd jobs around the house, a missing bracket, a cable router, steam deck holder and so on.
In the past if I was missing an app or useful software, a game I'd do the lazy thing, even though I can and have coded in the past, its "effort" I'll just go and buy or download the latest and greatest.
Earlier in the year I was lucky to snag a Minisforum MS-S1 395+ Max with 128GB Unified memory (currently setup 32gb system and 96gb Vram) before the price hike.
Was paired with a Qwen 3.6 27B or 3.6 35B moe but now a 3.8 27B uncensored. it can easily handle a Q8 with full 256k context.
Its now become my first instinct when I'm missing software to build it in a couple of hours local using the custom agent framework I setup.
Nothing I've created is for external use but every single day I find myself adding to it, while writing this post for example my framework finished an idea I had 2 hours ago when, I woke up this morning thinking I've got a lot of japanese visual novels and why don't I just design a combination hook into Exe or ocr the text app that translates via a local llm, and its done, ready for me to test.
I've written 12 adult games (don't code horny) a house AI, a coding framework, a game app to keep a track of all the games I play and download any faq or wiki to do with said game, 17 mods for my Skyrim install, 12 for my Fallout New vegas install, A temperature tracking system for the house that pulls rss local feeds and makes suggestions for my central heating system temps settings, A mapping software for my mobility scooter that checks my normal routes for issues and street work or maintenance that could make pavements impassable.
Plus hundreds of tweaks and test programs.
Anyone else out there using it like this ?
\---------Update-----
Awesome to see this kind of discourse one of the amazing strengths of these local llm's is it doesn't matter if they are slower, I can burn 50 million tokens over a 24 hours period on a new idea or problem and all it costs me is a little bit of electricity and time.
Just wanted to put that out there. It's like they get an ice pick to the brain no actual mourning here btw that'd be psychosis it's okay to laugh
I've been running Qwen3.8-27B as a local inference server for a production content intelligence pipeline (HVAC industry stuff, lots of long-context retrieval and structured extraction). I have been watching other redditors post their custom configurations, and I wanted to share what I tested to optimize for a single RTX 5090.
I was on llama.cpp (Q5\\\_K\\\_M GGUF, q5\\\_1 KV, 262K context, MTP), but it's limited to parallel=1 and I wanted concurrent serving. So I tested vLLM and NInfer as NVFP4 replacements and did a proper quality evaluation instead of just vibes.
\- RTX 5090 32GB (eGPU, OCuLink Gen4 x4) (yes, it's in an eGPU dock 😂 but that only affects model loading)
\- Ryzen 7 7840HS, 32 GB DDR5
\- Ubuntu 26.04, nvidia driver 610.43.02 (open)
## Engine configs
| | **\*\*llama.cpp\*\* | \*\*vLLM\*\* | \*\*NInfer\*\*** |
|---|---|---|---|
| Quant | Q5\\\_K\\\_M GGUF | NVFP4 | NVFP4 |
| KV cache | q8\\\_0 | FP8 | FP8 |
| Context | 196K | 262K | 240K |
| MTP | On (gate failed) | None | MTP3 (76% acceptance) |
| Concurrency | parallel=1 | Continuous batch | x2 lanes |
| VRAM | 31.6 GB | 29.6 GB | 30.5 GB |
I built a custom harness with 6 tiers, 50 items each, all from my actual production workload (not generic benchmarks):
Everything paired across engines: same prompts, same seeds (42), same gold labels. No cache\\\_prompt. Statistical comparison uses paired cluster-bootstrap CIs with pre-registered non-inferiority margins.
And when I do development with Claude Code, I often leverage multi-model consultations for design, planning, and code review. I did a 4-model review panel (Codex/GPT-5, DeepSeek V4 Pro, Grok 4.5, Kimi K3) audit the methodology mid-campaign. They found 10 issues, including 4 mislabeled gold items where the models were actually right and my labels were wrong. Fixed everything and reran.
| **\*\*Tier\*\* | \*\*llama.cpp\*\* | \*\*vLLM\*\* | \*\*NInfer\*\*** |
|---|---|---|---|
| Relevance | 86.0% | 84.0% | 86.0% |
| Needle (conditional) | 100% (29/29) | 100% (41/41) | 100% (41/41) |
| Transcript QA | 82.0% | 78.0% | 88.0% |
| Reasoning | 100% | 100% | 98.0% |
| Extraction | F1 0.300 | F1 0.350 | skipped |
| Tool replay | 0% | all errors | 0% |
Needle counts differ because llama.cpp's 196K context can't fit the 192K items (need room for max\\\_tokens + headroom). NInfer and vLLM both handle 192K fine. All engines score 100% on every needle they can fit.
Statistical comparison (NInfer vs llama.cpp, bootstrap):
\- Needle: delta = -0.29 (NInfer better, p=0.0006) - this is entirely from context capacity, not retrieval quality
\- Transcript QA: delta = -0.03, p=0.69 - no difference
\- Reasoning: delta = +0.02, p=0.72 - no difference
\- Relevance: McNemar p=1.0 - identical
\- Tool replay: delta = 0.0 - both fail equally
**\*\*Takeaway: quality is statistically indistinguishable across all engines.\*\***
| **\*\*Metric\*\* | \*\*llama.cpp\*\* | \*\*NInfer\*\* | \*\*Speedup\*\*** |
|---|---|---|---|
| **\*\*Decode 1K\*\*** | 114 tok/s | 158 tok/s | 1.4x |
| **\*\*Decode 32K\*\*** | 109 tok/s | 213 tok/s | 2.0x |
| **\*\*Decode 128K\*\* | 72 tok/s | 202 tok/s | \*\*2.8x\*\*** |
| Prefill 1K | 1,545 tok/s | 7,265 tok/s | **\*\*4.7x\*\*** |
| Prefill 32K | 2,155 tok/s | 6,892 tok/s | 3.2x |
| Prefill 128K | 1,528 tok/s | 3,904 tok/s | 2.6x |
| TTFT 1K | 670 ms | 138 ms | 4.9x |
| TTFT 32K | 15.2 s | 4.8 s | 3.2x |
| TTFT 128K | 85.9 s | 33.6 s | 2.6x |
vLLM speed excluded from the table because the perf probe used wall-clock timing (includes prefill + scheduling + decode) instead of server-side timings, so the numbers aren't comparable. From community reports and my own task-level measurements, vLLM does about 70 tok/s decode at short context, which actually matches NInfer's raw step rate (\~66 tok/s). The speed difference is entirely MTP3 speculative decoding.
**\*\*NInfer's speed advantage is all MTP.\*\*** The raw NVFP4 kernel speed is about the same between NInfer and vLLM (\~66-70 tok/s). NInfer's MTP3 speculation with 76% acceptance gets you to 158-213 tok/s. If vLLM or llama.cpp had working MTP on NVFP4, the gap would mostly close.
**\*\*The decode speedup grows with context.\*\*** At 1K context it's 1.4x. At 128K it's 2.8x. MTP acceptance stays high even at long context while llama.cpp's dense decode gets slower as context grows.
**\*\*NInfer's tokenizer endpoint is great.\*\*** It exposes \/v1/messages/count\_tokens\ (Anthropic Messages format) which gives exact token counts. No more \len(text)//3\ heuristics.
**\*\*NInfer does NOT support json\\\_mode (as far as I can tell).\*\*** \response\_format: json\_object\ returns 400. If you need structured JSON output, you'll need to route those calls elsewhere or use prompt-based enforcement.
**\*\*Don't trust vibes for quality.\*\*** I went in expecting NVFP4 might lose a few points vs Q5\\\_K\\\_M. It didn't. Not on any tier. The biggest delta across 250+ items was 2 percentage points, well within noise at n=50 (SE \~5.6pp).
NInfer NVFP4 replaces llama.cpp as my production engine. Same quality, 1.4-2.8x faster decode, 2.6-4.7x faster prefill, concurrent serving. The only gap is json\\\_mode.
I put together a detailed poster with all the charts and methodology details: \full results poster\
Setup if you want to try it:
\\\`
\# NInfer (from source)
git clone https://github.com/Neroued/ninfer && cd ninfer
mkdir build && cd build
cmake .. -DCMAKE\_BUILD\_TYPE=Release -GNinja && ninja
\# Model (HuggingFace)
\# https://huggingface.co/neroued/Qwen3.8-27B-nvfp4-NInfer (20 GB)
\# Run
./ninfer-serve /path/to/model.ninfer \\
\--model-id qwen3.8-27b \\
\--host 0.0.0.0 --port 8080 \\
\--max-context 240000 --kv-capacity 240000 \\
\--max-concurrency 2 --kv-dtype fp8 \\
\--spec mtp --draft-tokens 3 \\
\--vision --preserve-thinking
\\\`
This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM, everything else on the GPUs. Since then I switched quants, stacked MTP on top of the cache, fixed the load time, found a bug in the cache PR and found out my RAM was thermal throttling (Now i gotta buy an additional case fan lol). Numbers are all from the same 4,000-token python coding prompt with thinking on unless stated otherwise.
Where it's at now
|Starting numbers (UD-Q6\_K\_XL, 4+4 resident layers)|First post (Q6 + cache, 135 slots)|Now (UD-Q4\_K\_XL + cache 188 slots + n-gram draft)|Now (Q4 + cache 150 slots + MTP)|
|:-|:-|:-|:-|
|decode, coding prompt with thinking|17|25-29|32-35|37-41|
|decode, code emission, thinking off|\-|24|37|49|
|decode at 131k depth|12|17|18-20|14-16|
|prefill, 26k prompt (ub 512)|\~350 at ub 2048|138|180-195|180-195|
|load to ready|\~13 min|8.5 min|2 min|2 min|
|host RAM for the experts|104 GB pinned + 51 GB PLE|same|73 GB pinned + 28 GB PLE|same|
|cache hit rate|\-|84-85%|90-92%|84-85% (fewer slots)|
Hit rate is the cache's own counter, decode is llama-server's eval time.
What changed, in order of payoff
--numa distribute). Reading host-destination tensors straight from the file fixed it, that is in my PR #28223.perf stat -e unc_m_power_critical_throttle_cycles shows it, so don't worry if you have no BMC). A fan on the DIMM banks does keep it at 44-57 C, zero throttling, 16k-token runs from 10-15 to 24.6 t/s average. I would check this before touching software if your DDR4 Xeon box slows down under sustained load. This does not affect the numbers in the table and in my last post.Did nothing or hurt here: q8\_0 KV (-18% at 131k depth), mirror-NUMA #27986, QSA gather #28213, --load-mode none, chained drafts, the ik\_llama GEMV port, more than 2 cache uploads per step, thread/poll/prio flags. --lazy-mode on-direct (#28136) gives +7-12% only on the first long prompt after a restart.
To replicate
Branch with everything: https://github.com/Inovello/llama.cpp/tree/flashnext-2x3090. It is master (b96806d) + PR #27861 (expert cache) + PR #28223 (pinned host experts under mmap + the load fix) + PR #28243 (MTP) + the mul\_mat\_id fix and the 8-token cache gate from my #27861 comment. Squashed into one commit, I added the credit in the commit message. It is a replication branch.
git clone -b flashnext-2x3090 https://github.com/Inovello/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build -j -t llama-server
LLAMA_ATTN_ROT_DISABLE=1 numactl --interleave=all build/bin/llama-server \
-m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
-md mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf --spec-type draft-mtp -devd CUDA1 --spec-draft-n-max 3 \
-ngl 99 -c 261888 --parallel 1 -fa on \
-ot "ffn_(gate|up|down)_exps\.weight=CUDA_Host,per_layer_token_embd\.weight=CPU" \
-lzm off --numa distribute -t 16 -tb 44 -b 4096 -ub 512 -ctk f16 -ctv f16 \
--moe-expert-cache 150 -lv 4
-md of the main model.nvidia-smi after a long prompt, the CUDA pool grows \~350 MB during a 131k prefill.-lv 4 prints the cache hit rate every 512 steps (moe-cache: ... hit-rate=) and the draft acceptance per request.--spec-type ngram-map-k --spec-ngram-map-k-size-m 7 and raise the cache to 188.-devd on your single GPU or skip MTP and take the slots.Next thing I'll be doing is a proper comparison against Qwen3.8-27B at Q8 (speed and quality), since that is what most people here actually want to know. Happy to answer questions on any of it.
Github link: https://github.com/thatblend/LLMPSP
I wanted to see what the PSP can theoretically handle and I got my answer - a 90M model is about the max it can do without atrocious inference speeds. It's running around 0.5 - 0.6 tokens per second, which is very slow, but it's useable. Maybe 1-3 minutes for a reply.
The model is actually fairly impressive for 90M parameters, it's not really useful in any real metric, but it can generate crappy poems, short stories, write non-functional code and sometimes it gets things right if you ask it what company makes macbooks, what is an LLM etc, while other times it just hallucinates a crazy answer. Fun.
This epyc server I am using twelve cards with 64 GB memory, plus 256GB ram. Looking at the most capable models in open source, GLM 5.3 seems to be the only option, but with Astra releasing it will likely be fairly behind. GLM6 looks like it will be at least double in size, maybe even triple. Qwen-max and Kimmi are already way too big to even consider. Even the deepseek V4 Pro is too big. Should I just give up on this frontier dream sell the excess GPUs and settle For flash models with far fewer GPUs and a reasonable cost. Note: I'm not using it for any business.
I was hoping to build a new business with this, but it can probably be done with much more effort with a flash model as well.
Edit: I don't want to go below 4-bit quants because then the models start making obvious mistakes. So I'm talking about a min/max of 4-bit
Okay, this post really blew up. I wasn't expecting so much interest or comments just attacking me. Was really just expecting to have a calm discussion about future SOTA model sizes.
First, I'd like to thank the Qwen and Unsloth teams for the Qwen3.8 27b UD Q4\_K\_XL. Fits the poor 24GB of 3090 VRAM with 100k context at Q8 and works phenomenally well! Imho if theres anything that can threaten Anthropic/OpenAI profits is not another frontier model but actually these small ones you can run fast locally that can do 80..90% of mundane work for hours without paying a single dollar to any external company.
But next, if you want to jump up to a bigger smarter model I feel there is a gap now. Kimi-K3 is out of reach for many businesses let alone prosumers. So what frontier-like models do you use on what setups?
Is a DGX cluster (2..4 machines) or a GPU server with dual or quad GPU (\~96 ... 192 GB of VRAM + >256GB DDR4) a suitable setup to run something like MiniMax-M3 at reasonable speeds for agentic coding (>30tps)? And privacy aside, is hardware cost worth it?
I have a dual rtx3090 + 128GB ddr4 machine, running Qwen3.8-Flash-Next Q4 quite fast but despite being larger doesn't feel much smarter than the Qwen2.8 27b and while I \_can\_ run larger quantized models, Minimax-M2.7 being my workhorse, it way too slow for coding.
Hey everyone, been a while!
https://huggingface.co/TheDrummer/Artemis-31B-v1.1
https://huggingface.co/TheDrummer/Artemis-31B-v1
A few months ago, Gemma graced us with models that served as a much needed downpour from a year-long drought. I'm so happy to see us thrive once again.
The difference between v1 and v1.1 is quite simple: v1 was an early attempt, an overdue release that excelled in prose and writing, while requiring some handholding to get over quirks like stuttering. v1.1 is a more refined approach where stability meets quality. My community is split, so I figured I'd just release both.
\---
I was gone for a while. I got busy dealing with life, both its ups and downs. While I couldn't attend to you folks, I've been lurking around and appreciating you all for the kind words.
\- Skyfall 31B v4.2 seems to be a banger for many of you. I'm proud of the upscale and consider it my ultimate home-run send-off for the beautiful Mistral 24B base. It's a shame that it was overshadowed by Gemma 31B's release, but hearing some of ya'll compare and even prefer it to a more modern base was an unexpected win.
\- Rocinante 12B X / 16B XL proves that Nemo is still the ultimate creative model to this day. For some to say that 16B XL felt like Cydonia 24B v4.3 just goes to show how far you can go with modern resources and techniques.
\- Anubis 70B v1.2, Valkyrie 49B v2.1, Anubis Mini 8B v1 surprised me too. I had zero expectations releasing them. Just like Rocinante X / XL, they are modern finetunes of old base models. And somehow, they still found their users singing praises.
\---
With the Artemis release taking weight off my shoulders, I'm eager to move on and tune a ton more bases!
But I have something else cooking: a HordeAI-like platform. I hope to provide value not just as a finetuner, but as a local lover too!
The premise is simple: it's a place where generous local hosters can share inference with the less fortunate. You'd be surprised how many power users would love to heat their rooms through the power of charity.
\---
Finally, I'd like to thank everyone who supported me over the years. From those who provided kind words, rigorous testing, compute access, inference, or cold hard cash. You've all granted me the ability to enrich the local ecosystem with fun experiments like Rivermind 12B, Fallen series, Big Tiger Gemma, Precog 24B/123B, and solid models like Cydonia 24B v4.3, Behemoth X 123B v2.x, and Skyfall 31B v4.2.
If you've got inference / compute credits to share, please contact me! It will all go to making the community happy <3
Backlog:
\- Gemma E2B
\- Gemma E4B
\- Gemma 12B
\- Gemma 26BA4B
\- Qwen 3.8 27B
\- Muse Glimmer 30B
\- Mistral Medium 3.5 128B
\- HordeAI Alternative / Crowdsourced 'OpenRouter' ("BeaverNet")
I was running into a lot of posts that praised both the Fixed template and the Sharp template in comparison to the stock one, so I put them to the test.
It's not as extensive as it should be for a paper-grade analysis, but it gives out the point of each template.
I used SWE-bench Verified with mini-SWE-agent 2.4.6, slice 0:100 (the identical 100 tasks for all runs)
I containerized jpezzulli/sglang-rtxpro6000 and ran Flash Next with RadixArk/Qwen3.8-Flash-Next-NVFP4 on CUDA 13.3.
I ran all templates at both medium and xhigh reasoning efforts.
|Metric|Stock (medium)|Stock (xhigh)|Stock Δ|Fixed (medium)|Fixed (xhigh)|Fixed Δ|Sharp (medium)|Sharp (xhigh)|Sharp Δ|
|:-|:-|:-|:-|:-|:-|:-|:-|:-|:-|
|Resolved|91|99|\+8|87|98|\+11|94|94|\+0|
|Resolution rate|91%|99%|\+8 pts|87%|98%|\+11 pts|94%|94%|\+0 pts|
|Median output tokens|5,691|13,855|\+143.5%|6,956|14,819|\+113.0%|8,596|12,008|\+39.7%|
|Median reasoning tokens|3,050|8,759|\+187.2%|3,809|9,063|\+137.9%|5,437|7,967|\+46.5%|
|Median wall time|38s|1m 46s|\+180.4%|43s|1m 47s|\+152.3%|1m|1m 32s|\+53.4%|
|Total wall time|1h 47m 1s|4h 31m 22s|\+153.6%|1h 59m 53s|4h 4m 52s|\+104.3%|2h 29m 18s|3h 11m 36s|\+28.3%|
https://preview.redd.it/02geu81o8qnh1.png?width=1152&format=png&auto=…
https://preview.redd.it/v2mt6mgo8qnh1.png?width=1152&format=png&auto=…
https://preview.redd.it/ph1z36zo8qnh1.png?width=1152&format=png&auto=…
https://preview.redd.it/6ydu12mp8qnh1.png?width=1152&format=png&auto=…
Disclaimer: I wrote the post myself then used AI to format it properly for readability
I've been trying the latest models from the frontier labs and honestly, after extensive testing I can not tell the difference between the best open source options.
I think the differences are now marginal but the labs are doing heavy marketing to convince the public into paying more for tokens as they prepare to go public.
Can't help but see the similarities between the dot com bubble and AI in terms of a very insular environment where the technology will survive but the business models may not.
I've been building a cybersecurity network and we definitely know that even local AI models like Deepseek V4 flash do an excellent job and are really neck and neck with the best the frontier labs can provide.
Will be interesting to see how this all turns out! Exciting time nonetheless.
I see a new one being launched every few days... How do these new harnesses compare to claude code, pi etc. has anyone switched from these?
which harness to prefer and why
edit: Ive tried several different ones claude code, deepagents(langgraph), opencode, pi, and trueforge
my thoughts-
claude code - strongest on maturity and the managed experience but cost and token burn is high
deepagents - interesting middle ground if you want a more structured agent framework and the flexibility of an open-source stack. im interested in testing it more extensively on longer-running workloads fs
trueforge - this is a recent one, this was interesting to me because of its runtime-efficiency, also it allows separate the model from the runtime, which makes experimenting with different models much easier
https://github.com/truefoundry/trueforge
why?? - i also ran a benchmark on a real agent workload same model, same prompt, same tasks to compare these
adding the results of benchmarking i ran to compare this
so I tried to do this by running 14 cross-system tasks, three mcp servers behind them - a crm, an issue tracker, and a doc store through claude's managed agents, langchain's deepagents and trueforge, both open-source agent harnesses
the result that was most surprising:
Claude Managed Agents + Opus 4.8:
11/14 tasks solved | $11.8/run | 10.0M tokens/run
TrueForge + Opus 4.8:
11/14 tasks solved | $8.6/run | 3.7M tokens/run
Same model. Same benchmark. Same average solve rate, to my surprise trueforge used about 63% fewer tokens and cost about 30% less per run.
similar difference in tool usage: trueforge averaged 19 tool calls per task vs 32 for Claude Managed Agents.
Then I tried changing the model.
trueforge + GLM-5.2:
11.7/14 solved | $3.0/run | 3.8M tokens/run
On this benchmark, that was a slightly higher average solve rate than Claude Managed Agents + Opus at roughly 75% lower cost.
The token savings alone make this sooo interesting especially because the solve rate stays comparable
so this one was worth checking out ig
but this is still v early and the OSS runtime does not yet have first-class tracing/eval tooling. They don't ship their own code-execution sandbox, so you need to plug one in and context compaction is intentionally lossy.
So it is definitely not a replacement for a mature managed agent platform or other harnesses in the comparison, feature-for-feature today btu what I do find interesting is that the core runtime can already be competitive on these tasks while staying open, model-neutral, and deployable on my own infrastructure
this was their benchmark kit i used https://github.com/truefoundry/trueforge/tree/main/benchmark
It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.
Modified 4090 48GB has been out for a while. I remember a lot of people were buying them at the time. A lot of people were also complaining that they are meant to fail, that they scam etc.
I have a few questions to people people who bought these.
Full credits to @artificialisabel from X!
This is nothing impressive but, i had so much fun i wanted to share my experience with this model.
(yes this post is written by human)
I made an FPS with local Q4\_K\_XL 3.8 Flash Next (256k context) (it took 3 days to refine everything but playable demo was ready in 2 hours) to play with friends.
had ton of fun talking with them about what could we add , funny features etc.
Features:
I used opencode as harness, gun models were taken from sketchfab , model was running at 20tok/s avg with MTP, i know for someone is bad, but it did most of the work meanwhile i was at work or while sleeping, checking every now and then with a remote KVM from phone.
My machine:
5900x / 128GB DDR4 3200Mhz / RTX 5090 and RTX 4000 PRO (32 + 24 GB)
What games you would like to build in free time with ai? roguelites? 2d platforms? racing games?
Or did you already built something? share with some screenshots
Title help me decide and avoid making impulse purchases 😩 I already have dual 3090 which I can sell to help EDIT: ty all I’ll just wait it out, doesn’t seem worth it right now
Like the title says, running completely locally on my Xiaomi 14T Pro device. Specific model: Qwen3.8-Flash-Next-UD-IQ3\_XXS App used: BigMoeOnEdge
Along with everyone's favorite here, qwen3.8-27B