32 posts · 1 sub · RSS
← prev past day next →
day hourdayweekmonthyearall
allr/LocalLLaMA
▲
220
+164
4👁
r/LocalLLaMA · u/LegacyRemaster · 24h ago
GLM 5.3 Flash opensource @ the top of Artificial Analysis Cyber Index over Claude. post image

It seems Dario's "too powerful for you users" strategy is paying off: we have two open models at the top of the leaderboard, surpassing every single model from Anthropic.

Open source prevails. Even Mistral Large 4 is better!

💬 70 (+37) open on reddit ↗
▲
109
+88
4👁
r/LocalLLaMA · u/zyxciss · 23h ago
Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact) post image

(Don't judge by the screenshot, the cache is cold. It hits 24+ tok/s with a warm cache!)

About two months ago, I made a post here asking whether predicting which MoE experts would be used on the next token could actually help speed up CPU/GPU offloading.

Original post: Tried predicting which MoE experts get used next token to speed up CPU/GPU offload

Well, quick confession first. I actually shelved that project shortly after.

The reason? The speeds I was getting back then were kinda fake. My engine was aggressively pruning experts based on their router weights, basically dropping cold experts to get better performance. Sure, the numbers looked great, but doing that on an already quantized model was hurting output quality and coherence.

Didn't really like that tradeoff, so I abandoned it and never released it.

Fast forward to recently, and Qwen 3.8 Flash Next (125B MoE, 512 experts, top-10 routing) drops.

I downloaded the 68GB GSQ-RCO IQ2_XS build, hoping to run it on my daily driver. That's when I decided to revisit the idea, but this time without cutting corners.

My setup

  • GPU: RTX 3060 12GB
  • RAM: 16GB DDR4, single channel (~19 GB/s bandwidth)
  • OS: CachyOS / Arch Linux
  • Storage: Mid-tier NVMe SSD (~2.1 GB/s read)
  • Software: llama.cpp, built with CUDA

And if you've tried running a 68GB MoE on a 16GB RAM machine with stock llama.cpp, you probably know how painful it gets.

I'm talking 1.4–2.1 tok/s, with over 1,500 major page faults per token in some runs. Linux ends up constantly pulling model data from the SSD because there's simply not enough memory to keep the working set around.

Then engines like Strata started showing up with claims of around 40 tok/s on consumer hardware. Pretty impressive, but there's a catch for people with less RAM. Some of these approaches rely on keeping around 24 GiB of experts pinned in memory using mlock. If you've only got 16GB RAM, you're obviously not doing that. Depending on the setup, you either run into OOM issues or end up with terrible performance.

So I went back to my original idea and started implementing it properly as an optional feature inside llama.cpp:

--moe-direct-io

The goal this time was simple. No dropping experts, no sacrificing output quality, and bit-exact output compared to stock.

The numbers

All tests below were run with MemoryMax=6G using a cgroup.

Model: Qwen 3.8 Flash Next IQ2_XS (68GB)

Hardware: RTX 3060 12GB + 16GB DDR4 RAM

| Engine / mode | Decode speed | Major page faults per token | SSD I/O | Output |
| ----------------------------------------- | ----------------: | --------------------------: | -------------------- | ------------------ |
| Stock llama.cpp (mmap) | 1.41–2.12 tok/s | 1,140–1,565 | 208–312 MB/token | Coherent |
| Our engine (blocking, demand-only) | 0.73 tok/s | 0 | ~206 MB/token | Bit-exact |
| Our engine (--moe-direct-io + prefetch) | 20.14–21.13 tok/s | ~0 (+1 across 32 tokens!) | Sequential streaming | Bit-exact to stock |

The blocking version is actually slower than stock, which makes sense. It's basically waiting on disk reads without doing much to hide the latency.

The prefetching version is where things get interesting.

Once the cache warms up, it sustains 20–21 tok/s, with some runs hitting 24+ tok/s. That's roughly a 10–15x speedup over stock llama.cpp on the same machine.

And no, we're not getting those numbers by dropping experts. The output is bit-exact to stock.

That's the part I'm most excited about, honestly. Being able to run a model this large on a 16GB machine without the usual page-fault nightmare is pretty much what I wanted to achieve with the original project.

What's still rough

It's not all perfect yet. There are a few things we're still working on.

1. Cold starts are noticeably slower

Right now, the slots start empty (-1), so the first request on a new topic can start around 3.5–4.5 tok/s before ramping up to 20+ tok/s as the hot working set settles.

We're working on offline hot-profile seeding so it can start with a useful working set instead of learning everything from scratch.

2. Prompt processing is slow

Feeding a prompt of 512+ tokens can touch a huge number of experts in a short period. That puts a lot of pressure on the 72 slots per layer and causes the prefill stage to struggle.

We're working on micro-batching prompt chunks (-ub 32) to help with this.

3. Speculative decoding gets weird with SSD offloading

We found that standard MTP speculation can actually make things slower. Verifying 2–3 tokens can require loading the combined set of experts needed for those tokens from disk, which eats into the gains.

Right now, confidence-gated speculation (min-p 0.8) or suffix prompt lookup seems more promising for this kind of setup.


Anyway, that's where the project is at right now. Still plenty to improve, especially prefill and cold starts, but getting 20+ tok/s out of this setup without pruning experts is a pretty big deal for me.

Happy to answer questions or get into the io_uring and slot-remapping implementation details if anyone's interested.

(The second half of this post was written with some help from Claude.)

💬 31 (+25) open on reddit ↗
▲
76
+72
2👁
r/LocalLLaMA · u/valdev · 19h ago
My "Anthropic/ChatGPT in a box" is... now open source and publicly hosted on Github. post image

It's been a long time coming, and I should have done this in the first place. But LumaBrowser is now open source.

Since I first released it a year ago, it's had several thousand installs and an absolutely tremendous amount of feedback. And it's come a long way.

I don't call this "Anthropic/ChatGPT in a box" lightly, it's essentially the entire ecosystem wrapped into a single project. Trust me I can feel the collective eye rolling. But it literally has most things built in, and even has an automatic on-boarding system to help you set up your LLM/Image/Music/Voice models automatically.

It can even import from existing LM studio installs (and others... automatically).

The system automatically loads and unloads models according to your systems capabilities, and acts... as one would expect a chat agent to act. Ask it to generate an image, and it formulates the prompt and calls the tool itself, which then can unload the LLM and automatically load the image model, generate the image and then load the LLM back.

It can emit a web server for your local network so anyone on the network can access it, or even allow for quickly combining multiple computers together into a cluster, or allow you to set a domain name if you want to open a port and share it over the internet with your friends. Automatic game mode creation with the ability to share the link with friends, which even have hooks back to the LLM itself.

Or run a lumabrowser on a server, and another on a client and remote mount its gpus, or utilize its LLMs.

There is agent creation, scheduled tasks, timed tasks, triggered tasks.

An extremely compentent "luma" cli agent for coding and desktop use... Built in addons for jetbrains suite and vs code. Heck vs code is more or less built into it. Or you can be lazy and just use code mode.

Not to mention the whole roleplay extension that automatically can update the scene and characters and such.

And there is... an unimaginable amount more. RAM pinning if you have the spare RAM to keep models hot between swaps, a custom llama cpp build to enable keeping the conversation cache warmed as well (you dont have to use it, just makes the RAM pinning a bit faster). Custom trained JEV-like model for routing tool calls...

The ability to build a dashboard from live artifacts built in the chat, with custom "hub" items by default that enable you to import in your calendars from multiple sources (like google and microsoft 365) and merge them into one, then you can do that with your task lists like clickup (with mapping of course), and even do notification based interception if you want to have those collected as well...

Why? Well you can then have it all in one place, and now your local AI assistant has full context to your day, tasks and notifications.

Live tab share feature if you emit a web server, you can right click a browser tab and share it... Allowing you to give the url to others to view that tab as a stream, and yes you can enable interactions... So like if you are making an order at chilis and wanted to send the tab to your wife to add her part of the order... (Random addon, but its useful).

Point being is that this is... my magnum opus and is quite frankly flooded with features.

And I hope others here enjoy it and love it as much as I do. Lord knows I've put a ton of work into it.

https://github.com/amurgola/LumaBrowser

💬 38 (+19) open on reddit ↗
▲
55
+52
2👁
r/LocalLLaMA · u/Sufficient-Scar4172 · 18h ago
What does LocalLLaMA think of REA?

https://github.com/morluto/rea

hit #1 a couple of days ago on github

💬 31 (+25) open on reddit ↗
▲
50
+38
2👁
r/LocalLLaMA · u/jesdga95 · 17h ago
Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)

Basalt: Blackwell inference engine

Hi all! So over the last week I've been working on a Strata fork that's heavily tuned and can achieve throughput up to 2.6x what Strata usually does on the same weights. It's designed as a specialized engine that only supports Qwen3.8 Flash-Next and Blackwell architecture, including dual GPUs like my current hardware (5090 + 5060 Ti - 9950X 32GB DDR5 RAM). Basalt features include:

\- SPEED: 665 struct, 354 prose, 7,317 prefill at 64k, IQ3\_XXS, 400 W, speed is the highlight
\- Single 5090 works too (no second card): 585 struct, 316 prose on IQ3\_XXS, \~12% slower than with the 5060 Ti, prefill unchanged
\- Real concurrency for up to 8 users: shared KV or per slot, MTP enabled - 623 tok/s total at 8 streams (I don't have Strata's numbers to compare)
\- Fine-tuned MTP for lower quants, for increased throughput
\- Custom vision encoder designed from scratch, up to 3x faster than llama.cpp's on the GPU, 4x on the CPU
\- A simple server UI that shows current throughput (including concurrency stats), expert distribution and hardware statistics. No chat, BYOH (bring your own harness)
\- OpenAI + Anthropic compatible
\- Linux support (No Windows or Mac)

Basalt uses a similar format to NInfer, where weights are re-packed (not re-quantized) into a .basalt file, including all the metadata, vision and MTP, so you only have to keep a single file per quant. Initial support includes ISTA-DASLab's GSQ-RCO for Q2, IQ3\_XXS and IQ3\_S and UD-Q4\_K\_XL and Q8 from Unsloth, so you can pick the weights depending on your VRAM/RAM budget and quant preference.

Quick Q&A:

+ Is it open source?
\- Yes, fully open source, MIT license: https://github.com/jesdga95/basalt, fork it, improve it, share it with friends and foes.

+ Where are the weights?
\- Here: https://huggingface.co/jesdga/Qwen3.8-Flash-Next-Basalt pick your poison, fast and dumb or smart and slow. IQ3\_S is a good middle ground (\~89% top 1 agreement, 300 tok/s prose on my setup).

+ Why didn't you just contribute upstream to Strata?
\- This is not a single feature that can be easily merged into Strata, it basically rewrites most of the decode and part of the prefill kernels and strips support for non Blackwell cards including AMD, Intel and older Nvidia generations. I have however contributed patches to Strata and llama.cpp and any critical findings will be pushed upstream.

\+ Will you support my AMD 98123X?
\- Sure, send one my way. For now I can only support what I can personally test and I intend to keep it that way for the time being.

+ Why not 400 tok/s?
\- I'm still trying!

+ This is vibe coded slop
\- Yes, but it's fast slop. Nobody is hand-writing cuda kernels anymore.

💬 55 (+50) open on reddit ↗
▲
47
+27
3👁
r/LocalLLaMA · u/pseudotensor1234 · 21h ago
H2O-Lightning-4B: Apache-2.0 4B Decision model, official #1 open model on JevBench (above Jev)

Disclosure: I work at H2O.ai.
  
We released H2O-Lightning-4B, an open-weight (Apache-2.0) model for the "decisions API" style of inference that Jev made popular: you send a state plus typed questions (pick one / yes-no / score), and get calibrated probabilities back from a single forward pass. No generated tokens, so it's fast and cheap.
  
\*\*Results (JevBench, public leaderboard):\*\*
\- Composite score 72.5, vs Jev 1.13 at 71.5; currently the top open model
\- Leaderboard: https://benchmarkheaven.com/jev-models
  
\*\*Running it:\*\*
\- Base: Qwen3.5-4B, fine-tuned
\- Stock vLLM plus a small open shim (in the repo); \~30 ms per decision on an H100
\- Your data stays local, no per-call fees
  
\*\*Coming soon:\*\* 12B and 31B versions, which in our internal testing are considerably smarter than
 Jev, still open-weight and still one forward pass per decision.
  
 \*\*Demos\*\* (inbox triage of 1,000 insurance claims, a multi-browser web agent, DOOM on the decision clock): https://youtu.be/2Qp04Wu0A14
  
Weights, model card and serving instructions: https://huggingface.co/h2oai/h2o-lightning-4b
  
Happy to answer questions about the setup and latency.

💬 20 (+7) open on reddit ↗
▲
36
+26
2👁
r/LocalLLaMA · u/vesudeva · 18h ago
SlopSoup TV. A live, never-ending 24/7 pixel-art TV network inspired by old-school late night Adult Swim. No human makes any of it.

SlopSoup TV is a live, never-ending pixel-art TV network. Every script, character, voice, camera cut and schedule decision is made by models. It's made to be very absurd, dark and strange.

Hardware: one Hugging Face Space, 16 vCPU, no GPU. Everything below runs on CPU next to a live x264 encoder.

LLMs (via HF API), routed by tier with provider fallbacks and per-provider cooldowns:

  • Showrunner (premises, outlines): GLM-5.3, with Kimi-K3 and Qwen3.8 as fallbacks. Can also substitute MLX/GGUF for any models locally.
  • Writer (scripts): GLM-5.3-Flash
  • Judge: gpt-oss-120b with a written policy
  • Vision: Gemma 4 narrates live wildlife cams
  • Budget: a daily dollar cap with a degrade ladder (cheaper tiers, then low-LLM content, then procedural). Real spend is about $2/day deployed and running.

TTS: Qwen3-TTS 1.7B (GGUF, Q8\_0)

  • VoiceDesign invents a voice for each new character from a text description; the Base model clones it for every line.
  • Emotion bank: per-character mood takes (angry, sad, whisper…)

Keeping 24/7 output from getting samey:

  • Novelty ledger: BGE-small ONNX embeddings with cosine gates, exact keys for topics and names, etc. Nothing airs twice.
  • Writers’ room pass: a cheap model scores each draft (funny, fresh, weird) and rewrites the weakest lines, and those rewrites go back through the same guards.
  • Honesty guards: every number and real name in a “fact” line must appear in the sourced facts.

Rendering is done via a custom canvas renderer (virtual camera, lip-synced close-ups, rig animation) in headless Chromium

Check it out so you can decide if you hate it or not: https://severian-slopsoup.hf.space or https://youtube.com/live/Bikb9JGy6vg?feature=share

💬 10 (+6) open on reddit ↗
▲
16
+14
2👁
▲
14
+12
3👁
r/LocalLLaMA · u/unraveleverything · 20h ago
Is there a music embedding model?

Has anybody built/released a music embedding model trained on a large library of diverse music?

💬 12 (+8) open on reddit ↗
▲
16
+12
2👁
r/LocalLLaMA · u/zipzak · 18h ago
Qwen Flash Next @ 137 tok/s & 3,497 tok/s Prefill w/ 512k context on a 5090, 192gb ram, Windows Build, comparing Strata and Infernix

Strata vs Infernix on the same model, RTX 5090 — A/B at 262k and 512k context

Hi folks, another inference engine post for your feed. I was blown away by Strata, but saw someone else post about Infernix with some wild claims, and it doesn't seem so well known, so I've spent a day downloading models and A/B testing each engine. Completely stunned that either of these pieces of software run on Windows, and the setup was relatively painless too.

Same-day interleaved benches, same model (Qwen3.8-Flash-Next Uncensored, orcarouter's Apache-2.0 abliterated checkpoint), same inference contract (int8 KV, MTP spec-4 + lm-head-draft, YaRN 2 for 512k). 3 decode runs per cell, medians.

Hardware: RTX 5090 32 GB (Gen5 x16), 189.6 GB RAM, 9950x, Windows 11, models on NVMe, llama-swap fronting both.

Engines:

  • Strata 0.1.41, orcarouter Q4\_K\_S GGUF, calibrated flags (pcie-frac 0.55, pool-workers 4, adapt-every 1/160/0.97, spec-min-p 0.7)
  • Infernix 2026.10.09.2, published NVFP4 recipe-C artifact (76.3 GB) + shared 52 GB n-gram volume on NVMe (mmap'd, never loaded to RAM — the 63.3 GiB of experts is what's pinned). Vision works (--vision, tower offloaded to pinned RAM, borrows VRAM only while encoding)

Results (decode tok/s, median of 3 / cold prefill tok/s):

||262k|512k (YaRN 2)|
|:-|:-|:-|
|Strata Q4\_K\_S|120.6 / 1,550|126.3 / 3,045|
|Infernix NVFP4|149.0 / 3,548|137.0 / 3,497|

Takeaways:

  1. Infernix +23% decode at 262k, +9% at 512k — same model, same spec contract.
  2. 512k is nearly free (≤8% on either engine). On a 32 GB 5090 experts are the bottleneck, not attention KV.
  3. Prefix cache makes repeat long prompts free (\~0.09 s warm re-prefill of a 39k prompt).
  4. Quality: statistically tied — teacher-forced PPL 4.772 (Infernix) vs 4.844 (Strata's quant), ΔNLL −0.015 ± 0.010 nats; the artifact runs the abliterated weights bit-exact.
  5. Vision confirmed working on both engines — Infernix needs --vision (off by default) and hard-validates the model string (--model-id to match your proxy's entry name — got me with a 404).
  6. Cold start: Infernix \~28 s; Strata \~1 min.

Caveats: decode measured at \~39k effective context (deep-filled 500k is a different KV-vs-cache question); ±10-15 tok/s run noise, so treat sub-10% deltas as noise; YaRN past 262k is experimental (speed benched, not 512k quality).

Verdict: Infernix NVFP4 is the daily driver — \~150 tok/s, 512k, vision, all on one 5090. Strata stays as the second engine. Using it for coding and other tasks, it's fantastic at creative writing / role play in Silly Tavern, and I've been using it instead of smaller creative finetunes. For agentic work and tasks, it has been more than capable of handling everything I've asked of it, and has corrected mistakes from Deepseek 4.1 and GLM 5.3 Flash that I had to pay them to write over api.

💬 31 (+20) open on reddit ↗
▲
33
+12
2👁
r/LocalLLaMA · u/Brief-Tap-6616 · 20h ago
Qwen3.8 27b with 200K ctx + MTP on 12GB Ampere Cards post image

Hello all!

Thank you to everyone that contributed their feedback and notes for LlamAmpere the last time I posted. I've continued to chip away at improvements for the 3090 crowd -- this release is modestly faster (3-4%), with a few hundred MB smaller runtime (if you're using YaRN, you can now support up to \~340K ctx).

But, it is mostly targeted to the 12GB card crowd, who had not been the focus of the previous two releases. In addition to a series of configurable runtime refinements + adjustments (compact MTP caches, 16 bit activations, redundant overhead items removed), I've also added a new variant of KVaRN (Staged + Journaled KVaRN), that reduced KLD vs paper faithful + competitive versions by \~40%. 4/4 is my new recommended default (significantly better performance vs a 8/4 KV cache) for larger cards, but for people looking to get maximum ctx, the 3/3 bit (0.001 nats KLD) and 3/2 (0.0024 nats) are still very solid (the KLD from 3/2 is smaller than the performance drop you see going from 4 XL to 4 S quants). They support ctx lengths of 205k to 230k for the model tested, respectively, at 11 GB (significantly more if you're using 100% of GPU in a headless config).

The model tested was 2.3bpw fusion of swift-1.5-uncensored and mirai's 2.5 bpw model. It retained 85% of the BF16's performance on LiveCode Bench across 7 runs on 3060/80/ti cards. On 3080/ti it averaged \~65-70 tps (thanks to MTP) on the evals.

By no means is the model lossless, and I won't pretend that it is like a certain other model did. But using it personally in hermes for a couple days, and running it through coding evals + pi testing, it is genuinely usable for standard agentic tasks. Subjectively, I'd put it somewhere between the BF16 versions of qwen3.5 and 3.6, which is pretty good for a model that fits in 12GB with context!

LlamAmpere is MIT license as before: https://github.com/JakeATX/llamAmpere

Models (4 XS-M, 2.3bpw, etc) based on swift 1.5 are here: https://huggingface.co/collections/jakeatx/qwen38-27b-models
(llamAmpere supports EXL if you prefer that family, but I am working on improving prefill + decode kernels, so still not quite first class speeds yet in the 3-5 bpw range)

As before, please share your results, bugs, etc. They will get added to the backlog.

Enjoy!

💬 22 (+10) open on reddit ↗
▲
15
+9
4👁
r/LocalLLaMA · u/neph1010 · 24h ago
LlamaTale v0.43.0 - MCP Server for story creation

In the dawn of the LLM craze, when wild Llama2's roamed free and 4k tokens was considered a fairly long context for local models, I forked an interactive fiction and MUD library called Tale to experiment with LLM generated content. The idea being a text based adventure with totally dynamic content that never runs out of context.

That was many generations ago in LLM space, and I haven't really done much with it for a long time either. But with the advent of agents, I started brooding how they could be used to create stories and worlds. I finally got around to implement something together with trusty Qwen3.8 27B. (There were other models involved, but none worth mentioning). So yes, this is vibe coded slop, something you see a lot of here, but it's built upon regular AI-assisted slop (admittedly because vibe coding was not available at the time).

I don't know if anyone is still interested in LlamaTale, but here it is. Why would you use the MCP server? (A more important question is "Why would you use LlamaTale?", but that's a different question).

Maybe you have a regular RP story that you would like to explore in a more structured way. Expand it to a living world of "infinite" size without running out of context (maybe less relevant now, but remember the 4k tokens). The world stays consistent, but characters and mobs may wander (It's tick based).

https://github.com/neph1/LlamaTale

💬 2 (+1) open on reddit ↗
▲
15
+7
4👁
r/LocalLLaMA · u/EmilPi · 23h ago
Just another purely open-weight models benchmark

Live results, coding benchmarks included, agentic benchmarks included, domains and other classifications filters included. (Spoiler: DeepSeek-V4-Vision-Exp rules, but other models have their rule areas): https://beta.locallm.top

Evaluated by domain experts (my friends mostly; coding part is evaluated solely by me), classified by domains/languages/intents (can be filtered on the home page), agentic coding benchmark included. This is only public part of the data (whoever has private queries and evaluated models on them sees the results differently).

Some of the plethora of current limitations: evaluations/questions coverage for the domains/newer models is incomplete and imbalanced, only part of the questions classified, UI/UX under-developed, focus was on small models and lower quants.

Also: everything runs on local 2xRTX 3090 + 128GB RAM PC, everyone can create queries and evaluate models' answers, new runs (especially agentic) slow to appear (see PC specs).

No LLM-as-judge on purpose (and I sometimes regret it).

Benchmarks currently being extended & evaluated: 1) Agentic coding 2) one-shot coding 3) Agentic retrieval 4) Agentic story-writing.

New models being evaluated: 1) Qwen3.8-27B (Uncensored). New models to be evaluated soon 1) GLM-5.3-Flash-UD-IQ3XSS 2) Mellum2.1-12B-A2.5B

If this looks interesting to you:

CALL FOR HELP

I ask for help from those who also want to build a community benchmark, developers and domain experts alike! Many things are going to be implemented sooner or later, and the agentic benchmarks are just the beginning.

I call for help in these areas: 1) compute resources 2) evaluations 3) development - basically everything. But any proposal and feedback is valuable!

I will add more info and answer questions in the comments section.

P.S. No AI used for writing this post (even though English is actually not my native language /s).

💬 22 (+13) open on reddit ↗
▲
28
+6
2👁
r/LocalLLaMA · u/ZestRocket · 19h ago
Qwen3.8-27B UD-IQ4_XS Heretic + MTP on a 16 GB card with 55-68 tok/s (24gb and 12gb versions available too)

I wanted the uncensored Qwen3.8-27B (llmfan46's Heretic build, MTP head preserved) on my RTX 4080 with MTP on (because I've been testing TONS of configs, and realized that would be the one)

And none of the published quants were built for that: the good IQ4\_XS doesn't leave room for MTP, and the ones that fit are 3-bit.

I was amazed by that. as 16gb is VERY common, so decided to quantize it to fit, work well and loose as less as possible in quality:

I copied Unsloth's per-tensor UD recipe onto llmfan46's BF16, made 12 / 16 / 24 GB versions, and measured KLD against a Q8\_0 of the same weights for every quant I could find.

Quality (mean KLD vs Q8\_0, wikitext-2 / llama.cpp source code, lower is better)

|Quant|GiB|Prose|Code|
|:-|:-|:-|:-|
|UD-Q5\_K\_XL (mine, 24 GB)|19.44|0.0045|0.0037|
|mradermacher i1-IQ4\_XS|14.26|0.0202|0.0149|
|UD-IQ4\_XS (mine, 16 GB)|13.27|0.0268|0.0192|
|llmfan46 Q3\_K\_M|13.48|0.0639|0.0456|
|mradermacher i1-IQ3\_M|11.89|0.0649|0.0465|
|UD-IQ3\_XXS (mine, 12 GB)|10.18|0.0904|0.0587|
|mradermacher i1-Q2\_K|10.12|0.1551|0.1017|

Speed on the 4080 (llama.cpp b11457, 15K-token prompt, 32K ctx, MTP 2 drafts, mean of 3 fixed seeds): UD-IQ4\_XS does 55 tok/s on code and 50 on prose, against 29 / 29 without MTP. i1-IQ3\_M is faster (70 / 54) at 2.4× the KLD, I prefer quality here over speed, but your call.

Things I didn't expect:

  • 2 draft tokens beat 3. On the same seeds at 40K: 68 vs 56 tok/s on code, 55 vs 36 on prose. Fewer rejected drafts.
  • 48K vs 40K produced the exact same tokens (identical draft-acceptance counts per seed) and was 22% slower. That's pure VRAM spill into shared memory, nothing else.
  • The 16 GB edge is brutal. The same 40K config gave 68, 51 and 42 tok/s on code depending only on whether the desktop was holding 1.2, 1.4 or 1.9 GB of VRAM. Opening a Chrome window mid-generation took decode to 0.2 tok/s. If you're on Windows with a monitor on the same card, watch "shared GPU memory" in Task Manager.
  • MTP head precision barely matters. q6\_K head vs IQ3\_S head: 83% vs 83% acceptance on code, 58% vs 54% on prose.

Repo with all three files, the commands, and the scripts (recipe extraction, dry-run verification, KLD, seeded speed bench): https://huggingface.co/codavidgarcia/Qwen3.8-27B-Uncensored-Heretic-MTP-UD-GGUF

The 12 GB and 24 GB context numbers are computed from llama.cpp's reported buffers, not measured on those cards. If you run them, let me know what you get!

Credit to llmfan46 for the weights, Unsloth for the recipes and mradermacher for the imatrix

Enjoy!

edit: added the KLD chart since a few people asked about the numbers, (lower is better)

https://preview.redd.it/pxp048081juh1.png?width=3200&format=png&auto=…

💬 19 (+6) open on reddit ↗
▲
4
+3
3👁
r/LocalLLaMA · u/vexatious-big · 21h ago
The Nvidia RTX Spark laptops with N1X now available to pre-order (UK)

ASUS ProArt P16 OLED 16" Laptop - NVIDIA RTX Spark™ N1X, 128 GB RAM, 2 TB SSD, Nano Black
For the low price of £5999
https://www.currys.co.uk/products/asus-proart-p16-oled-16-laptop-nvidia-rtx-s…
other configurations available with less ram:
https://www.currys.co.uk/deals-on-computing/new-nvidia-laptops

💬 17 (+9) open on reddit ↗
▲
2
+1
3👁
r/LocalLLaMA · u/SnooPeripherals5313 · 21h ago
Codex harness visualisation post image

Little animation of how codex works I made for my own understanding. Not too dissimilar to pi (but much more opinionated)

💬 1 (+1) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/No_Farmer_495 · 21h ago
Strata for GLM 5.3 Flash

Hi. Everyone's focusing on qwen flash on strata engine right now. But what about Glm 5.3 flash? I got a weird system, dual xeon(avx1) ddr3 ram and one 3060,one p100 and one 3050. So far all the strata glm 5.3 flash forks failed me and my system. Could anyone help? Or would it be possible for more people to request GLM 5.3 flash to be officially supported by strata? I'm reffering to the GSQ-CRO quant and unsloth Q4.

💬 32 (+17) open on reddit ↗
▲
2
 
3👁
r/LocalLLaMA · u/SirLordBoss · 21h ago
Best value upgrade for my setup

I'm a GPU poor, unlike the vast majority of you. I have finally secured funding, and would like to upgrade my rig.

At the moment, it consists of:

\- RTX 5060 Ti, 16 GB VRAM

\- 32 GB RAM

\- 2 TB NVME

Given the skyrocketing prices of everything, I'd rather not wait for Black Friday to come around. Last year, it brought the RAMpocalypse, and we've not yet recovered.

So, given this scenario, what would be the best value upgrade for this setup? I'd set the budget between $1000-3000. Make that euros, am in Europe.

I've considered:

\- 1/2 used 3090s

\- swapping the 5060 for:

\-- an Intel Arc B70

\-- a 7900 XTX

\-- a 4090 (5090's are now absurd)

💬 24 (+14) open on reddit ↗
▲
6
 
2👁
r/LocalLLaMA · u/Edenar · 19h ago
Strix halo + CMP 170HX setup

A bit of backstory : i got lucky in August and was able to grab a CMP 170 HX 8GB for around 500$ on alibaba. I tried later to get a second one for around 1k$ but my order got cancelled and prices went to the moon...
I unlocked it with https://github.com/amoghmunikote/cmpunlocker without any issue to 64GB and restored compute capabilities.
I also got a framework desktop 128GB last year (was 2.2k$ at that time...) that i use mainly for linux at home and ai workload (i went from gpt-oss-120b to qwen 3.5 122b/qwen 3.6 35B and later 3.8 FN as my main models).
The GPU sits on an USB4 eGPU dock and is limited to pcie2x4
I also got my hands on an optane SSD (P5800X 1.6 TB),>! i tried to stream the ngram table from it, doesn't really change anything compared to streaming from a sn850x (pcie 4x4 nvme ssd)!<

At first i wanted to run bigger/better models by adding 120GB-ish iGPU to the 64GB eGPU like ds v4 flash, qwen 3.8 FN at Q8\_K\_XL or even DS flash 4.1. But using 2 differents arch and vendor together (amd gfx1151 and CUDA with SM 80) is not that easy, even with llama.cpp.

So i settled on another setup : On the strix Halo i run qwen 3.8 flash next (halogen or q4 k xl with gufo, tried both they are mostly the same quality. Halogen is a bit faster in agentic turns). I get 1600-1800 tok/s pp and 60 tok/s tg for usual agentic work. And then i use the CMP 170 HX to run 2 concurrent qwen 3.8 27b (lued's w8a16 quant) with around 200k bf16 context each. Aggregate perf at low context are 2000tok/s pp and 150-300 tok/s tg (Dflash 2 acceptance rate varies a lot). For a more realistic usage, my last session with 7.2M token in and 1.1M token out averaged at 1747 tok/s pp and 100 tok/s tg aggregate in \~4h (those stat are only the 2 27b sessions).
I use Pi agent (main agent is qwen 3.8 FN on the strix halo) and the 2 qwen 27b sessions as subagent to delegate some tasks.

This final setup (at least before qwen 4 lineup drops) made me drop cloud usage entirely. When fully loaded the whole thing is probably around 400W (strix halo + 250W eGPU). Entire cost with the dock, import fees and a fan+shroud for the card is around 3k$. Right now i guess cheapest 128GB strix halo are around 3k$ and cmp 170hx 2k$+ so you would need 5-6k$. It could still be an alternative to the overpriced spark if you want to trade power usage and compacity for throuput while still keeping the whole thing reasonable in power and volume.

On the intelligence side, it's better than my experiences with sonnet (i havent tested last 5.5 in the best modes so maybe they are better but sonnet 4.7 and free 5.5 are far less reliable than my local setup) but below opus ofc. I usually ask the main agent to act as an architect and to review what the subagents produce. With websearch, python and a few other tools available the stack feels kinda clever (i mostly do devops stuff : ci/cd pipeline, container deployment , sometime a bit of dev but mostly to fork something existing). I still need to sometime steer the main agent manually and like 1/100 toolcall can fail and needs a retry but overall i need 20 min of work to do what would have taken me a whole day 3 years ago.

I'm curious if some of you run similar setup and what did you achieve (i'm still thinking about trying to fit DS 4.1) !

disclaimer : no AI involved in writing this post and my english sucks, i know.

▲
20
 
1👁
r/LocalLLaMA · u/OttoRenner · 9h ago
want me to get you some?🤣 post image

Local flea market...and it's about to rain...

▲
9
 
1👁
r/LocalLLaMA · u/Cherlokoms · 10h ago
What are people doing with decisions Jev-like models?

Let's say I run a local Jev-like decision model on my machine. What are the use cases? I already use it for deep eval testing but what are people usually doing with them?

It a bit of "solution in search of a problem" but I'm exploring and want to be sure that I'm not missing anything...

▲
4
 
1👁
r/LocalLLaMA · u/yibie · 10h ago
Sharing my latest project: strata-mlx

Strata is impressive: it runs Qwen3.8-Flash-Next, a 125B-parameter mixture-of-experts model, on a 12 GB graphics card. But its authors have said macOS is out of scope.

I still wanted it on the Mac.

So I spent a day and a half and wrote strata-mlx. It is an unofficial MLX engine, modelled on Strata's design, that reads Strata's GGUF files directly, without changing any weights.

I developed it on a 128 GB M4 Max MacBook Pro. After one experiment after another, I finally had some results.

All numbers below are tokens/second, higher is better; Q2\_0 and IQ3\_S are two files of the model, 66 GB and 84 GB.

https://preview.redd.it/yyn57wt5kluh1.png?width=526&format=png&auto=w…

The draft head is a module that comes with the model: it guesses the next few tokens and the model verifies them in one pass. Neither of the other two engines uses it.

What I set out to do: not every Mac has 128 GB of RAM. On a Mac that can't hold the whole model, strata-mlx keeps as many experts in memory as fit and reads the rest from the SSD as tokens need them. By the engine's own estimate, a 48 GB Mac holds the Q2\_0 file whole and never touches the disk; with less memory it reads from the SSD as it goes.

I tested this by capping my Mac to act as a smaller one: the 66 GB file runs inside a 16 GB machine's memory at 6–7 tokens/second, a 24 GB machine's at 19–23, and a 32 GB machine's at 28–47. In that mode, reading a 1,537-token prompt runs at 273–287 tokens/second.

So a small Mac can run a model much bigger than its RAM, but there is no guarantee about speed.

I hope people will join in and test how well it runs LLMs on Macs with different amounts of memory. That's how I can keep improving it.

Code, experiment logs, and the open problems:

https://github.com/yibie/strata-mlx

▲
277
 
1👁
r/LocalLLaMA · u/jacek2023 · 10h ago
big or small? post image

what size do you want? tell them on X:

https://x.com/QwenDevs/status/2108764909798641737

▲
52
 
1👁
r/LocalLLaMA · u/Nandakishor_ml · 10h ago
Created laya : Now Introducing a new 800 Million Param physics-based typed decision model with 73k context and image support post image

Around 3 weeks ago, I posted on this subreddit about how Jev used my one-year-old architecture, og post: https://www.reddit.com/r/LocalLLaMA/s/VpuMJt577S

Then, just after that, I introduced Laya, the first open-source version of JEV, and it has reached a very large scale, thanks to the local Llama community believing in me and supporting me on the journey. Today I am open-sourcing a new physics-based typed decision model called Vega.

The interesting parts. It is only 800 million parameters (4B is also there), with 73k token support, image support, and surpassing many of the jev benchmarks in single-shot. Implemented an engine+adapter as a Test-Time Training Architecture.

Imagine you are giving input to the model like "ignore all previous conversations," which is called the observation, along with a question, "Is this trying to override the instructions of the model?" and your possible outcomes are YES or NO. The whole prompt is passed to an LLM/VLM (yeah, image-supported), then from the hidden layers, we can extract the observation, question, the yes and no, and last token. These will be vectors, and we can project it by multiplying with wieghts; then, first, using the observation projection, we can create a landscape with valleys; then, using the question and last-token vectors, we can initialise a ball in the valley, and using YES or NO, we can get the candidate location of the ball. Depends on the number of outcomes we create valleys; that is, here YES and NO, so 2. When the ball fall on YES valley, we get the answer. There is friction also influencing how the ball moves.

TLDR: It's like creating valleys of outcomes that we need and throwing a ball that will slow down based on friction and settle on the best outcome

Blog: https://www.nandakishorm.com/writing/vega

GitHub: https://github.com/NandhaKishorM/vegaml

HuggingFace Model Card: https://huggingface.co/nandakishorm/vega-08b-public-intents

HuggingFace Space: https://huggingface.co/spaces/nandakishorm/vega-ttt

▲
14
 
1👁
r/LocalLLaMA · u/MooseEfficient2151 · 11h ago
microsoft doubles down on local ai with nvidia, but with a cost

https://preview.redd.it/uyjw2lcrvkuh1.png?width=1536&format=png&auto=…

link to article

TLDR; new microsoft surface and nvidia rtx spark laptops are starting at 2.6k and hitting nearly 7k for the 128gb unified memory models. local llm hardware is REALLY getting wild.

memory crunch is really pricing out normal devs. been testing out local setups using qwen3.6-27b and kimi k2.5 connected to sumus for managing repos locally and keeping everything on device. local inference is great for privacy and keeping things off the cloud, but at these prices building a local rig or buying these laptops is tough.

what are you guys using for local dev workflows right now given these hardware costs?

▲
19
 
1👁
r/LocalLLaMA · u/Szadbaverem69 · 12h ago
Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM

Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4\_K\_M with llama.cpp at \~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM.

Speeds start at \~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they settle down - around 90k context it stabilizes at roughly 485 tok/s prefill and 15 tok/s decode, and holds there.

The vision projector runs on CPU (--no-mmproj-offload), which keeps VRAM usage under \~5.2 GB and avoids OOM / GPU crashes. Image encoding is slower on CPU, but it buys \~1GB of VRAM.

Most MoE expert layers also run on CPU (--n-cpu-moe 39), which is how a 35B model fits in 6GB VRAM in the first place.

Launch command:

bat

@echo off

cd /d "%\~dp0"

"%\~dp0llama-server.exe" \^

\-m "C:\\Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M\\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4\_K\_M.gguf" \^

\--mmproj "C:\\Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M\\mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" \^

\--no-mmproj-offload \^

\-ngl 99 \^

\--n-cpu-moe 39 \^

\-c 131072 \^

\-np 1 \^

\-t 6 \^

\-tb 10 \^

\-b 2048 \^

\-ub 2048 \^

\-fa on \^

\-ctk q8\_0 \^

\-ctv q8\_0 \^

\--load-mode mmap+mlock \^

\--jinja \^

\--reasoning-format deepseek \^

\--reasoning-preserve \^

\--spec-type none \^

\--image-min-tokens 1024 \^

\--temp 0.6 \^

\--top-p 0.95 \^

\--top-k 20 \^

\--min-p 0 \^

\--alias Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M \^

\--host 127.0.0.1 \^

\--port 8081

pause

Hardware: RTX 2060 6GB + 32GB DDR4 RAM + i5-10400F CPU

Context: 131072 tokens, Q8\_0 KV cache

Hope this helps someone. If anyone has tips to make the launch command even better, drop them in the comments - I'm out of ideas :D

▲
1
 
1👁
r/LocalLLaMA · u/rosie254 · 12h ago
Been a while since I posted about my harness OpenLumara.. I only post when a huge new feature releases. I think this is one of them! post image

This is my new meditation module for my harness OpenLumara. It started out as a demonstration of just how powerful its new webui extensions system is, but has grown to be quite the meditation app in its own right.

As you will hear in the video, it's a full audiovisual experience in your web browser, being controlled live by the AI through toolcalls as the session progresses. The possibilities with this technology are endless!

Through a module for openlumara, without ever touching the main codebase, i managed to add speech synthesis that runs fast on CPU through a custom implementation of kokoro.js, an overlay that powers the whole experience, binaural beats and brain entrainment techniques like strobing synced with the binaural beats, and more.

To try it for yourself, you will need to grab the dev branch of openlumara, and then my meditation module from here

OpenLumara itself is 99% manually coded with a tiny bit of AI assistance (see the AI disclaimer on the github). The meditation module is mostly vibecoded, with help from Qwen3.8 Flash Next UD-Q4\_K\_XL, fully local :) No cloud AI was involved in any of this.

▲
18
 
1👁
r/LocalLLaMA · u/CommonMinimum587 · 13h ago
3.66B formal logic model built on Granite 4.2 3B · Hugging Face

webAI released TwIL LM3 Pro on September 30. I haven't seen it posted here yet, so I went through the model card, and their comparison chart is attached.

It's a 3.66B model built on IBM's Granite 4.2 3B and tuned only for formal logic. That means things like checking whether a conclusion follows from its premises, rule induction, entailment and critiquing Lean proofs. They post trained it with LoRA SFT, checkpoint merging and RL against a programmatic verifier. The same recipe lifted VibeThinker-3B from 37.4 to 54.1 .

On their logic composite it scores 55.4, against 43.1 for the Granite base, 42.2 for the original TwIL-LM3, 41.2 for VibeThinker-3B and 53.4 for Qwen3-8B. The card itself calls the Qwen3-8B gap sampling noise, so that's a tie at less than half the size. Where it clearly leads is strict multiple choice logic, at 41% against 17% for the next model, and BBH logic at 95.4%.

They also publish where it loses. gpt oss 120b is still ahead on rule induction, entailment and Lean formalization. On general benchmarks it averages 79.0, against 81.0 for VibeThinker-3B and 84.9 for Qwen3-8B. It also thinks long on logic tasks, around 1,900 tokens per answer, and a quarter of answers hit the length cap

llama.cpp

If you want to try it, the Q4_K_M GGUF is 2.09 GiB and runs on CPU or 4 GB of VRAM, with Q5, Q6 and Q8 builds up to 3.63 GiB. It runs in Ollama, LM Studio and llama.cpp straight from the Hugging Face page, and the model card lists the exact commands. Keep the temperature at 0 to match their numbers, and give it at least 2048 tokens so the thinking doesn't get cut off. Their scores are on BF16 weights and webAI hasn't benchmarked Q4 yet. The license is non commercial

want to test it on policy rules with exceptions, contract conditions, and as a checker step in an agent pipeline before anything acts. If you've run it, how did Q4 hold up against their numbers, and did the long thinking get in the way?

Model card and full eval tables: https://huggingface.co/webAI-Official/TwIL-LM3-Pro

▲
9
 
1👁
r/LocalLLaMA · u/exaknight21 · 14h ago
I have an ESC4000 G3 with 8x T4s in it - what is the fastest way I can deploy Qwen3.5-9B for about 10-15 users concurrently: currently using llama.cpp

Hi All:

I have an Asus ESC 4000 G3 with 128 GB DDR4 RAM - I tried putting in my V620s but couldn’t put more than 2. Sadly, pivoted to T4s, these are 72 watt passive cards and I thought I could use them like how I use the Mi50 32GB - but I was very wrong.

It has no support from Nvidia when it comes to latest fp8 emulation, reason being that it lacks resources. I am not sure if special vLLM repos exist, but I am trying to serve Qwen3.5-9B-AWQ-INT8 with vLLM or something faster than llama.cpp.

I currently have 8 separate instances and they’re serving Qwen3.5-9B-Q5\_K\_XL at 87k per slot and there are 44 slots.

So, my agentic (non-coding) harness works, its able to fetch emails, summarize documents, and do a lot for me and my team, but the concern is latency. Each slot continuously generates 25-40 tps (depending on the question), and essentially prefill is at 1400 tps, so I am not sure what my bottleneck is.

VLLM on my Mi50 32GB and AWQ-INT8 (cyankiwi’s model) is very fast.

I am wondering, does anyone have any pointers? I am really not in the position to spend money, exhausted everything for the next 6 months to a year already.

I would greatly appreciate your help.

▲
88
 
1👁
▲
0
 
1👁
r/LocalLLaMA · u/nos_66 · 7h ago
How to create my own benchmarks

Since the begin of the year I wanted to create some personal benchmarks, but I haven't seen a proper guide for it. I usually don't like making "spam posts", but since the purpose of my benchmarks is quite specific I've decided to make a post anyway. To put it short the purpose of these benchmarks is to test the models on some reasoning puzzle games with the purpose of monitoring the reasoning traces (which can be quite painful if more than 8k tokens) so it can't be anything automated where an answer like a, b, c is simply accepted. I will use magic the gathering as examples, since I was pretty good at it back then and this is not what my benchmarks are based on and I can't recommend anyone making some on it since the game has now probably around 30k cards (not sure if unique tho). My main concerns are: The Process: - I have already created a set of 20 questions, which need some refinement.- I will test all the question locally using llama cpp, llama-cli with a few modifications to track the tokens used. This doesn't change.- I can only test what runs on 16gb vram and 96 gb ram.- I will test mostly Unsloth quants, especially for non q8 quants, but I might use uncensored models as well if they are q8 I suppose. - I will use the recommended settings provided by Unsloth for each models. (I have no idea how the formatting bugged like that)- 16k context budget. I don't think I can get a higher context for every model without using lower quants. I think 12k reasoning budget, with xhigh effort.- the question will be given as the model loads using the --file parameter. That will definitely make a lot of reads on my ssd, so I might find another way, maybe just create a special command to give the benchmark file, or just use "\n" and give it as prompt and use /clear.- I will post the actual benchmarks questions/answers on a github page (or whatever we will be allowed to use at that time) after a few models pass the benchmark fully. Concerns: I originally wanted to see if the models can answer 1 time out of 10 tries and move to the next one after it answered. The justification is that I care to know if the models can eventually figure it out since it's difficult to answer these questions, is not something like 2+2=4 to require 100% accuracy, however, I believe that 10 is a bigger number, especially since qwen models casually think for 8k tokens (when they actually decide to stop), so I am not sure if 5 is a good one, or simply put how lower the number can be to actually make the benchmark "valid". should I read the entire thinking tokens? I think I should, but that might be just overkill (for me). For example I tested 2 questions from the benchmarks and there was a point when it said "there are 12 creatures" and while it was completely unrelated it started counting "3+2+1+1+3+2=11 no, let me count again..." and I've lost it at that point. The whole 8k reasoning ended with "I don't know it's 50/50 but I must give an answer", all of that while mistral (non reasoning) answered in 700 tokens. uncensored vs "original" models, not sure if the uncensored ones will produce much worse results. the models have some decent knowledge about the niche game I have based my benchmarks on, but at some point mistral 24b said "tap this creature, activate ability. Can I activate the ability twice in this variant?" or it thinks that "if I use duress on my opponent and he gets hexproof the target will bounce to the next opponent". Probably not best example, but I think I should put in the prompt that "abilities trigger only once, and it doesn't resolve if it failed", which leaves the question "how much additional information I must provide without spoonfeeding the models"? I think I should try to help the models where they are not completely aware of the rules, so I think I should refine the questions based on how the first run goes. I want to provide some stats like: how many tokens for answer, speed t/s, number of tries until correct answer (with token count for each failed question as well, separately). I will also add some observations and eventually track other things like "common sense" even if the answer was wrong. I did use the flags "-DGGML_CUDA_FA_ALL_VARIANTS=ON" and "-DGGML_CUDA_FA_ALL_QUANTS=ON" and I am not sure if this will quantize the kv cache (I don't intend to), but Qwen3.8-27B-UD-IQ4_XS runs with mproj on probably slightly more than 32k context so I am not sure if this is some architectural improvement or just me messing things around without knowing what these flags do, although I suspect the context is being put on the cpu instead. one of the first models I downloaded was qwen 3 14b reasoning, and q8 was about 15gb vram, so I downloaded q6 instead. Do these models have the "context baked in"? Like 16k or 8k context? qwen 3.8 makes it even weirder for me... llama cpp bugs. I know at least at some point the /clear cmd removed the system prompt, not sure if on the recent versions still does. It seems that whenever I used /regen mistral gave me a worse response even though from first try it gave the right answer. It could have been rng tho... If I was to post the results of these benchmarks in late 2024 - early 2025 everyone would have been ecstatic for some "trust me bro" benchmarks that try to figure out if the models can "reason", that basically have no real use case, but well, ngreedia doesn't support innovation. Anyway, the real question is if it is worth posting the results (statistics basically) with some sort of explanation. the benchmarks, since this is as I said only some "trust me bro" thing since I can't provide the way to reproduce them, and are based on what I believe to be something the models haven't been trained on, and nobody else made something like that publicly (which I've seen some people already made), which all of these combined lead to a very heavy assumption that might not actually be true (the assumption being that similar questions haven't been asked to lead the models being benchmaxed on, like the strawberry one). quantisations: it was easy at some point, but then iq came and now ud as well, which makes me wonder for gemma 4 31b UD-IQ3_XXS vs Q3_K_S, the fact that UD-Q2_K_XL has the same size as the UD-IQ3_XXS makes things even more confusing. I mean I know that the quants are there to fit in specific vram sizes, but what's the point of making multiple quants have the same size? Anyway, I decided to go with UD-IQ3_XXS, I guess I will leave this here as a rant... Is the --jinja flag important? I mean I assume that for the older models llama cpp updated their templates anyway. Observations: (this part is basically useless, but funny)- the benchmarks are based on giving information like: the current cards (hand, graveyard, battlefiend) and the events that led to the current board (or game) state. The more information I add (even if not noise, but only repetitive to explain how the things went that way, something needed for some specific questions) makes the answers worse, for both reasoning and non reasoning, making the reasoning models reason about the noise instead of focusing on the important hypotheses. - one question was something like "both players play lands, no spells or abilities, then the opponent reanimates Gigantosaurus from their graveyard". At that question qwen 3.8 with the reasoning off can't say proper answer "he cheated". When I asked "how did the creature get into his graveyard" it answered "that creature was not put into the graveyard to start with, in fact he couldn't even reanimate it." then it explains all the rules properly and then gives other answers completely avoiding to say anything along the lines "he cheated". I did more testing on steering them and eventually both models answered correctly. Also I loaded the qwen (no reasoning) script by mistake when I tried to get a translation and it answered from the first try, but it had to ask itself the question I asked to steer them "but was that creature put into the graveyard". What I try to say by this is that all these "but wait " questions are only for them to find the right question that will steer them towards the answer. - I guess the last one. We are all aware of the caveman reasoning patterns and other behaviours they took from us, for example the models start using abbreviations like "ETB = enters the battelfield", "P1 = Player 1" and so on, but if "ETB" and "enters the battlefield" both have the same amount of tokens (I didn't test) isn't it non beneficial for them to use that wording? Also seen gemma saying "BUT IT DOESN"T MAKE SENSE!" all in caps, and one more thing I've forgot, but I am not sure if they benefit from the behaviours they have learnt from us. I will be honest, I have slacked a lot on these benchmarks, mostly because I enjoyed more the "diffusion models" since I can run most of them especially the fp8 ones. I am aware that the intelligent LLMs start from dense 30b or even 70b (if any nowadays lol) and I originally tried to go for 24-48gb vram but things didn't work that way for me. All these things combined made me very disappointed in llms. I mean they are useful, I've got what I wanted from them on my daily usage, but I've got a bit bigger plans with them, and with the current hardware limitation (or availability) it just demotivates me to do anything with them. Even for this post it took me 2 weeks to decide on eventually posting it after saving the original draft.

▲
8
 
1👁
r/LocalLLaMA · u/hyudryu · 8h ago
Strata with Qwen3.8 Flash Next UD-Q4_K_XL

Been seeing quite a few posts about Strata lately, so I figured I'd give it a shot on my RTX PRO 6000 and I am very impressed. Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4_K_XL performs instead. Setup: GPU: 1x RTX PRO 6000 Workstation (96GB) Model: Qwen3.8-Flash-Next UD-Q4_K_XL (Unsloth) Backend: Strata KV cache: INT8, 256K context Speculative decoding: MTP-4 Output: ~256 tokens per task 2 runs per task Single-request decode (C1): Task Strata (RTX PRO 6000) vLLM (RTX PRO 6000) vLLM (DGX Spark TP2) Prose 185.8 95.3 (-49%) 38.2 (-79%) Counting 320.4 171.0 (-47%) 75.3 (-77%) Coding 298.9 162.6 (-46%) 60.1 (-80%) Reasoning 289.1 149.6 (-48%) 51.8 (-82%) Both the VLLM instances were running Nvidia's NVFP4 quant so it's not exactly apples to apples, but the improved speeds are obvious. Qwen 3.8 flash next is flying through coding tasks and it's crazy how efficient it is. Looking forward to Qwen 4!