Title help me decide and avoid making impulse purchases 😩 I already have dual 3090 which I can sell to help EDIT: ty all I’ll just wait it out, doesn’t seem worth it right now
Title help me decide and avoid making impulse purchases 😩 I already have dual 3090 which I can sell to help EDIT: ty all I’ll just wait it out, doesn’t seem worth it right now
After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code (C code), the results were not completely unexpected but some quants were definitely underwhelming.
TLDR: Best overall: bartowski/Qwen3.8-27B-IQ4_XS. Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS. For a bit more context: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S or uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp
TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS, bartowski/Qwen3.8-27B-IQ3_XS, prism-ml/Ternary-Bonsai-27B-Q2_g64 and magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unslothunsloth/Qwen3.8-27B-UD-IQ3_S and ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_Smagiccodingman/Qwen3.8-27B-MQ-IQ2_M_1 and huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3_SAtomicChat/Qwen3.8-27B-AD-IQ4_XS-IQ3_SIvanKrastevAdventics/qwen3.8-27b-awq-int4-q4_0 (gguf of cyankiwi/qwen3.8-27b-awq-int4)bartowski/Qwen3.8-27B-IQ4_XS and bartowski/Qwen3.8-27B-IQ3_XXSbartowski/Qwen3.8-27B-Q3_K_M, Thireus/09ae8ba_22b6bb2 and Thireus/09ae8ba_248b31bbartowski/Qwen3.8-27B-IQ2_Shuihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtpmradermacher/Eintopf-Qwen3.8-27B.i1-IQ3_M and hitsfmdj/Qwen3.8-27B-4.2BPW-16GBJoakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3_S-recoveredMia-AiLab/Qwen3.8-27B-EXL3-3.5bpw and turboderp/SC_3.00bpw_H4_V4prism-ml/Ternary-Bonsai-2-27B-PQ2_0 and prism-ml/Ternary-Bonsai-2-27B-PTQ1_0, replaced prism-ml/Ternary-Bonsai-27B-Q2_g64 with prism-ml/Ternary-Bonsai-27B-PQ2_0Bucoid/Qwen3.8-27B-Heretic-Ara-iq4_xs-3.0agentionai/Qwen3.8-27B-AP-IQ3_S, agentionai/Qwen3.8-27B-AP-IQ4_XS and RonnieOps/Qwen3.8-27B-IQ4_XS-fullvocab-E3. This is probably the final edit as Qwen4 27B will likely come out soon.byteshape/Qwen3.8-27B-IQ4_XS-3.84bpw and byteshape/Qwen3.8-27B-IQ3_S-3.23bpw. Again, probably the last edit.(sorted by Mean KLD)
|Model|Mean KLD|Same top p|GGUF size (without MTP)|
|:-|:-|:-|:-|
|prism-ml/Ternary-Bonsai-27B-PQ2\_0|1.289582 ± 0.008684|82.849 ± 0.118 %|6.7GiB|
|prism-ml/Ternary-Bonsai-2-27B-PQ2\_0|1.096134 ± 0.007705|84.596 ± 0.113 %|6.7GiB|
|prism-ml/Ternary-Bonsai-2-27B-PTQ1\_0|1.095914 ± 0.007703|84.582 ± 0.113 %|5.5GiB|
|sdkyuan/qwen38-27b-qat-q2\_0|0.893177 ± 0.006948|85.727 ± 0.110 %|8.2GiB|
|bartowski/Qwen3.8-27B-IQ2\_Sbartowski/Qwen3.8-27B-IQ2\_S (NEW)|0.784060 ± 0.006457|87.016 ± 0.105 %|8.7GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2\_XS|0.767174 ± 0.006291|86.166 ± 0.108 %|7.8GiB|
|TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS|0.514311 ± 0.004864|89.023 ± 0.098 %|8.9GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2\_S|0.512614 ± 0.004909|88.802 ± 0.099 %|8.6GiB|
|empero-ai/Qwen3.8-27B-Ridge-3.7bpw|0.475767 ± 0.004483|89.612 ± 0.096 %|11.4GiB|
|magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4\_K\_S-Unsloth|0.419585 ± 0.004076|89.661 ± 0.095 %|13.1GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3\_XXS|0.379222 ± 0.003992|90.270 ± 0.093 %|9.4GiB|
|unsloth/Qwen3.8-27B-UD-Q2\_K\_XL (UD2)|0.350861 ± 0.003745|90.626 ± 0.091 %|9.6GiB|
|byteshape/Qwen3.8-27B-IQ3\_S-3.23bpw|0.345563 ± 0.003703|90.868 ± 0.090 %|10.1GiB|
|mradermacher/Eintopf-Qwen3.8-27B.i1-IQ3\_M|0.318143 ± 0.003139|91.471 ± 0.087 %|11.7GiB|
|bartowski/Qwen3.8-27B-IQ3\_XXS (NEW)|0.300480 ± 0.003345|91.511 ± 0.087 %|11.3GiB|
|unsloth/Qwen3.8-27B-UD-IQ3\_XXS (UD2)|0.268594 ± 0.002971|91.951 ± 0.085 %|10.8GiB|
|magiccodingman/Qwen3.8-27B-MQ-IQ2\_M\_1|0.256808 ± 0.002861|92.056 ± 0.085 %|10.9GiB|
|DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-MTP-IQ3\_M|0.251270 ± 0.002702|92.315 ± 0.083 %|13.1GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3\_S|0.249650 ± 0.002841|92.016 ± 0.085 %|10.8GiB|
|bartowski/Qwen3.8-27B-IQ3\_XS (OLD)|0.238656 ± 0.002627|92.312 ± 0.083 %|12.2GiB|
|hitsfmdj/Qwen3.8-27B-4.2BPW-16GB|0.222090 ± 0.002570|92.552 ± 0.082 %|11.7GiB|
|esatapedico/Qwen3.8-27B-NVFP4-MTP-LOW|0.220796 ± 0.002631|92.339 ± 0.083 %|14.3GiB|
|unsloth/Qwen3.8-27B-UD-IQ3\_S (UD3)|0.218522 ± 0.002591|92.399 ± 0.083 %|10.9GiB|
|turboderp/SC\_3.00bpw\_H4\_V4 (exllama3)|0.205712 ± 0.002509|92.573 ± 0.082 %|11.9GiB|
|jrell/Qwen3.8-27B-i1-IQ4\_XS-GGUF-Smaller|0.194459 ± 0.002242|93.049 ± 0.080 %|12.4GiB|
|agentionai/Qwen3.8-27B-AP-IQ3\_S|0.193041 ± 0.002352|92.995 ± 0.080 %|10.9GiB|
|orcarouter/Qwen3.8-27B-Uncensored-Q3\_K\_L|0.192312 ± 0.002294|92.726 ± 0.081 %|13.4GiB|
|bartowski/Qwen3.8-27B-Q3\_K\_M (NEW)|0.191103 ± 0.002369|92.823 ± 0.081 %|12.3GiB|
|mudler/Qwen3.8-27B-APEX-I-Mini|0.190209 ± 0.002354|93.012 ± 0.080 %|12.6GiB|
|Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3\_S-recovered|0.178882 ± 0.002102|93.110 ± 0.079 %|11.0GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3\_S-mtp|0.178715 ± 0.002200|92.949 ± 0.080 %|11.0GiB|
|Thireus/09ae8ba\_22b6bb2 (ikllama.cpp quality 41.39%)|0.178290 ± 0.002202|93.115 ± 0.079 %|11.0GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3\_S|0.175223 ± 0.002129|93.024 ± 0.080 %|11.0GiB|
|Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw (exllama3)|0.149827 ± 0.001936|93.735 ± 0.076 %|14.1GiB|
|unsloth/Qwen3.8-27B-UD-Q3\_K\_XL (UD2)|0.147186 ± 0.001809|93.734 ± 0.076 %|12.2GiB|
|unsloth/Qwen3.8-27B-UD-Q3\_K\_XL (UD3)|0.142647 ± 0.001860|93.789 ± 0.076 %|11.9GiB|
|byteshape/Qwen3.8-27B-IQ4\_XS-3.84bpw|0.131976 ± 0.001690|93.852 ± 0.075 %|12.0GiB|
|IvanKrastevAdventics/qwen3.8-27b-awq-int4-q4\_0|0.112990 ± 0.001558|94.171 ± 0.073 %|14.4GiB|
|AtomicChat/Qwen3.8-27B-AD-IQ4\_XS-IQ3\_S|0.111713 ± 0.001492|94.527 ± 0.071 %|13.2GiB|
|Bucoid/Qwen3.8-27B-Heretic-Ara-iq4\_xs-3.0|0.097572 ± 0.001341|94.651 ± 0.070 %|13.0GiB|
|Bucoid/Qwen3.8-27B-Uncensored-IQ4\_XS\_4BPW|0.091447 ± 0.001261|94.774 ± 0.070 %|12.8GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4\_XS|0.082871 ± 0.001205|94.981 ± 0.068 %|13.1GiB|
|unsloth/Qwen3.8-27B-UD-IQ4\_XS (UD3)|0.075626 ± 0.001097|95.258 ± 0.067 %|13.3GiB|
|agentionai/Qwen3.8-27B-AP-IQ4\_XS|0.073386 ± 0.001075|95.386 ± 0.066 %|13.0GiB|
|Thireus/09ae8ba\_248b31b (llama.cpp 49.75%)|0.063904 ± 0.000967|95.687 ± 0.064 %|13.3GiB|
|jpetrina/Qwen3.8-27B-IQ4\_XS-pure|0.061984 ± 0.000917|95.551 ± 0.065 %|13.3GiB|
|RonnieOps/Qwen3.8-27B-IQ4\_XS-fullvocab-E3|0.058278 ± 0.000886|95.773 ± 0.063 %|14.1GiB|
|bartowski/Qwen3.8-27B-IQ4\_XS (OLD)|0.056482 ± 0.000856|95.835 ± 0.063 %|14.3GiB|
|bartowski/Qwen3.8-27B-IQ4\_XS (NEW)|0.055415 ± 0.000849|95.850 ± 0.062 %|14.2GiB|
|unsloth/Qwen3.8-27B-UD-Q4\_K\_XL (UD3) *(can't fit)*|0.029844 ± 0.000476|96.921 ± 0.054 %|16.1GiB|
|unsloth/Qwen3.8-27B-UD-Q4\_K\_XL (UD2) *(can't fit)*|0.028026 ± 0.000432|96.988 ± 0.054 %|16.4GiB|
https://preview.redd.it/g9isjm0d04sh1.png?width=5355&format=png&auto=…
Hope this helps other VRAM starved people like me :)
It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.
Like the title says, running completely locally on my Xiaomi 14T Pro device. Specific model: Qwen3.8-Flash-Next-UD-IQ3\_XXS App used: BigMoeOnEdge
Github link: https://github.com/thatblend/LLMPSP
I wanted to see what the PSP can theoretically handle and I got my answer - a 90M model is about the max it can do without atrocious inference speeds. It's running around 0.5 - 0.6 tokens per second, which is very slow, but it's useable. Maybe 1-3 minutes for a reply.
The model is actually fairly impressive for 90M parameters, it's not really useful in any real metric, but it can generate crappy poems, short stories, write non-functional code and sometimes it gets things right if you ask it what company makes macbooks, what is an LLM etc, while other times it just hallucinates a crazy answer. Fun.
Just wanted to put that out there. It's like they get an ice pick to the brain no actual mourning here btw that'd be psychosis it's okay to laugh
Hey everyone, been a while!
https://huggingface.co/TheDrummer/Artemis-31B-v1.1
https://huggingface.co/TheDrummer/Artemis-31B-v1
A few months ago, Gemma graced us with models that served as a much needed downpour from a year-long drought. I'm so happy to see us thrive once again.
The difference between v1 and v1.1 is quite simple: v1 was an early attempt, an overdue release that excelled in prose and writing, while requiring some handholding to get over quirks like stuttering. v1.1 is a more refined approach where stability meets quality. My community is split, so I figured I'd just release both.
\---
I was gone for a while. I got busy dealing with life, both its ups and downs. While I couldn't attend to you folks, I've been lurking around and appreciating you all for the kind words.
\- Skyfall 31B v4.2 seems to be a banger for many of you. I'm proud of the upscale and consider it my ultimate home-run send-off for the beautiful Mistral 24B base. It's a shame that it was overshadowed by Gemma 31B's release, but hearing some of ya'll compare and even prefer it to a more modern base was an unexpected win.
\- Rocinante 12B X / 16B XL proves that Nemo is still the ultimate creative model to this day. For some to say that 16B XL felt like Cydonia 24B v4.3 just goes to show how far you can go with modern resources and techniques.
\- Anubis 70B v1.2, Valkyrie 49B v2.1, Anubis Mini 8B v1 surprised me too. I had zero expectations releasing them. Just like Rocinante X / XL, they are modern finetunes of old base models. And somehow, they still found their users singing praises.
\---
With the Artemis release taking weight off my shoulders, I'm eager to move on and tune a ton more bases!
But I have something else cooking: a HordeAI-like platform. I hope to provide value not just as a finetuner, but as a local lover too!
The premise is simple: it's a place where generous local hosters can share inference with the less fortunate. You'd be surprised how many power users would love to heat their rooms through the power of charity.
\---
Finally, I'd like to thank everyone who supported me over the years. From those who provided kind words, rigorous testing, compute access, inference, or cold hard cash. You've all granted me the ability to enrich the local ecosystem with fun experiments like Rivermind 12B, Fallen series, Big Tiger Gemma, Precog 24B/123B, and solid models like Cydonia 24B v4.3, Behemoth X 123B v2.x, and Skyfall 31B v4.2.
If you've got inference / compute credits to share, please contact me! It will all go to making the community happy <3
Backlog:
\- Gemma E2B
\- Gemma E4B
\- Gemma 12B
\- Gemma 26BA4B
\- Qwen 3.8 27B
\- Muse Glimmer 30B
\- Mistral Medium 3.5 128B
\- HordeAI Alternative / Crowdsourced 'OpenRouter' ("BeaverNet")
This epyc server I am using twelve cards with 64 GB memory, plus 256GB ram. Looking at the most capable models in open source, GLM 5.3 seems to be the only option, but with Astra releasing it will likely be fairly behind. GLM6 looks like it will be at least double in size, maybe even triple. Qwen-max and Kimmi are already way too big to even consider. Even the deepseek V4 Pro is too big. Should I just give up on this frontier dream sell the excess GPUs and settle For flash models with far fewer GPUs and a reasonable cost. Note: I'm not using it for any business.
I was hoping to build a new business with this, but it can probably be done with much more effort with a flash model as well.
Edit: I don't want to go below 4-bit quants because then the models start making obvious mistakes. So I'm talking about a min/max of 4-bit
Okay, this post really blew up. I wasn't expecting so much interest or comments just attacking me. Was really just expecting to have a calm discussion about future SOTA model sizes.
This is nothing impressive but, i had so much fun i wanted to share my experience with this model.
(yes this post is written by human)
I made an FPS with local Q4\_K\_XL 3.8 Flash Next (256k context) (it took 3 days to refine everything but playable demo was ready in 2 hours) to play with friends.
had ton of fun talking with them about what could we add , funny features etc.
Features:
I used opencode as harness, gun models were taken from sketchfab , model was running at 20tok/s avg with MTP, i know for someone is bad, but it did most of the work meanwhile i was at work or while sleeping, checking every now and then with a remote KVM from phone.
My machine:
5900x / 128GB DDR4 3200Mhz / RTX 5090 and RTX 4000 PRO (32 + 24 GB)
What games you would like to build in free time with ai? roguelites? 2d platforms? racing games?
Or did you already built something? share with some screenshots
€7??? Surprise Price Ends with Limited Stock That's likely 7999 EUR, so double of the initial price of MS-S1 MAX-128GB? 😭
From: Polymarket on 𝕏: https://x.com/Polymarket/status/2095646821485842805 Julien Chaumond on 𝕏: https://x.com/julien\_c/status/2095822387895824836
I'm one of the developers. We said in August it would go open source in September and it did last night. MIT or Apache-2.0, pick one. The repo you see is our internal repo, kernels included, so from now on everything happens in public.
It's an inference engine in Rust and C++ with our own CUDA kernels. One binary with OpenAI and Anthropic style APIs, loads GGUF and safetensors. We run about 300B tokens a year through it at work.
Some numbers: Qwen3.8-27B FP8 on one RTX PRO 6000, spec decoding off on every engine:
Full board with the losses: https://truespar.com/paddock/benchmarks/qwen38-27b
What it does not do yet: No Mac, no ROCm, no Vulkan.
One model per GPU, no tensor parallel. CUDA only, Windows and Linux. Validated on Blackwell (5090, RTX PRO 4500/5000/6000, B200) and Ampere (an A6000 was the bring-up card, 30-series works). Ada kernels ship but nobody has run a board on them so the engine refuses to start unless you set PADDOCK_UNVALIDATED_ARCH=1. Hopper and A100 kernels are in the tree without a board.
https://github.com/truespar/paddock
Thankful for any help and input!
How good/bad is it against comparable MoEs the same size? How does it compare against Qwen 3.6 35BA3B?
Since we don't have 3.8 35B this seems like an upgrade if we look at some benchmarks like terminal bench, but they don't have SWE bench pro on the benchmarks table, and i don't really know anything about this lab, I'm wondering if it trades blows with models like tiel coder or if it's some benchmaxxed model like ornith?
At a single glance it looks really decent but haven't tried it in depth yet. What are your experiences with this model so far guys?
This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM, everything else on the GPUs. Since then I switched quants, stacked MTP on top of the cache, fixed the load time, found a bug in the cache PR and found out my RAM was thermal throttling (Now i gotta buy an additional case fan lol). Numbers are all from the same 4,000-token python coding prompt with thinking on unless stated otherwise.
Where it's at now
|Starting numbers (UD-Q6\_K\_XL, 4+4 resident layers)|First post (Q6 + cache, 135 slots)|Now (UD-Q4\_K\_XL + cache 188 slots + n-gram draft)|Now (Q4 + cache 150 slots + MTP)|
|:-|:-|:-|:-|
|decode, coding prompt with thinking|17|25-29|32-35|37-41|
|decode, code emission, thinking off|\-|24|37|49|
|decode at 131k depth|12|17|18-20|14-16|
|prefill, 26k prompt (ub 512)|\~350 at ub 2048|138|180-195|180-195|
|load to ready|\~13 min|8.5 min|2 min|2 min|
|host RAM for the experts|104 GB pinned + 51 GB PLE|same|73 GB pinned + 28 GB PLE|same|
|cache hit rate|\-|84-85%|90-92%|84-85% (fewer slots)|
Hit rate is the cache's own counter, decode is llama-server's eval time.
What changed, in order of payoff
--numa distribute). Reading host-destination tensors straight from the file fixed it, that is in my PR #28223.perf stat -e unc_m_power_critical_throttle_cycles shows it, so don't worry if you have no BMC). A fan on the DIMM banks does keep it at 44-57 C, zero throttling, 16k-token runs from 10-15 to 24.6 t/s average. I would check this before touching software if your DDR4 Xeon box slows down under sustained load. This does not affect the numbers in the table and in my last post.Did nothing or hurt here: q8\_0 KV (-18% at 131k depth), mirror-NUMA #27986, QSA gather #28213, --load-mode none, chained drafts, the ik\_llama GEMV port, more than 2 cache uploads per step, thread/poll/prio flags. --lazy-mode on-direct (#28136) gives +7-12% only on the first long prompt after a restart.
To replicate
Branch with everything: https://github.com/Inovello/llama.cpp/tree/flashnext-2x3090. It is master (b96806d) + PR #27861 (expert cache) + PR #28223 (pinned host experts under mmap + the load fix) + PR #28243 (MTP) + the mul\_mat\_id fix and the 8-token cache gate from my #27861 comment. Squashed into one commit, I added the credit in the commit message. It is a replication branch.
git clone -b flashnext-2x3090 https://github.com/Inovello/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build -j -t llama-server
LLAMA_ATTN_ROT_DISABLE=1 numactl --interleave=all build/bin/llama-server \
-m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
-md mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf --spec-type draft-mtp -devd CUDA1 --spec-draft-n-max 3 \
-ngl 99 -c 261888 --parallel 1 -fa on \
-ot "ffn_(gate|up|down)_exps\.weight=CUDA_Host,per_layer_token_embd\.weight=CPU" \
-lzm off --numa distribute -t 16 -tb 44 -b 4096 -ub 512 -ctk f16 -ctv f16 \
--moe-expert-cache 150 -lv 4
-md of the main model.nvidia-smi after a long prompt, the CUDA pool grows \~350 MB during a 131k prefill.-lv 4 prints the cache hit rate every 512 steps (moe-cache: ... hit-rate=) and the draft acceptance per request.--spec-type ngram-map-k --spec-ngram-map-k-size-m 7 and raise the cache to 188.-devd on your single GPU or skip MTP and take the slots.Next thing I'll be doing is a proper comparison against Qwen3.8-27B at Q8 (speed and quality), since that is what most people here actually want to know. Happy to answer questions on any of it.