Like.. yea bro, I knew GLM was cool. Now everyone does.
Like.. yea bro, I knew GLM was cool. Now everyone does.
Here are some inspirational quotes you can put into the comments:
It had such a long run. It was my first introduction to modern language models. I remember getting slightly excited over it. And now it lives purely in our memories. Arguably what's more infuriating is that they suggest using GPT-5.6 Terra as a replacement. Keep in mind that Babbage is a model that's literally 3/4 of the size than MiniCPM5 2B. Even Luna might be overkill as a replacement. But neither is a drop-in replacement. Davinci is the main GPT-3 most people use. This is why we have local models, because they simply cannot have a universal end of life date.
You may think that jev, OpenClaw or TurboQuant are super cool, but actually the coolest LLM invention happened two years ago
As we all know, the best source of reliable information about LLMs is YouTube:
https://preview.redd.it/4tv0ctbwyfsh1.png?width=2544&format=png&auto=…
Back in September 2024, Reflection 70B appeared out of nowhere and was announced as an open-source model that supposedly destroyed GPT-4o
There was only one small problem. People downloaded it. And tested it :(
https://preview.redd.it/kamhrp9izfsh1.png?width=1514&format=png&auto=…
It turned out that Reflection 70B was basically a Llama 3.1
https://preview.redd.it/ghmof4l50gsh1.png?width=1524&format=png&auto=…
but at the end the mystery was solved
https://preview.redd.it/44eb3jw90gsh1.png?width=1530&format=png&auto=…
Let this be a moment of reflection on the current hypes in LocalLLaMA.
great summary by Maziyar PANAHI https://x.com/MaziyarPanahi/status/1838559480658710982
Some context first.
I am process improvement / business consultant and had worked with Fortune 500 companies on improving their processes around refunds, returns, customer support, etc.
This entire thing had a lot of complex decision making and generally every decision / condition node in a process map was generally replaced by a human because factoring in ambiguity in a code is very difficult.
Idea of ImaJev
Hence, when Jev came out, I was very intrigued with it and also could clearly see its use-case of improving decision making in complex decision work flows.
However, Jev didnt have support for Images and I thought that it can be replicated for both Text and Images in a single model and thats when I started with ImaJev.
Training Process
It went badly at first. My first big fine-tune on about 500k short decisions made the 9B model worse at reasoning: 64.9 down to 42.3 on JevBench hard. It had basically learned to pattern-match. I spent the next couple of weeks generating hard questions with open-weight models and only keeping the ones where two Ai models agreed on the answer. That brought it back.
Results
Then, on the JevBench - It came out #1 of 91 (v1.4.2.2, scored 27 Sep), 67.37 vs Jev 1.13.0 at 63.29.
The same week DecisionBench put it #3 of 56, ahead of GPT-5.6 Luna and DeepSeek V4.1.
I honestly didn't expect either.
To be fair about it: the #1 is on a score that weighs accuracy, calibration, speed and cost equally. On accuracy alone it's #3. Its main strength is that when it says 90% it's usually right, and it'll say "can't tell" instead of guessing.
What it actually is: LoRA plus a small decision head on Qwen3.5-4B. You give it text or a JSON record, up to two photos, and closed questions.
It gives back a probability for each option plus "unknown", in one forward pass.
Runs on a Mac with MLX or on one GPU.
The whole project costed me around $1200 in rented GPU and a lot of time :P
I would love to know your thoughts on it - it anyone would be interested to try that.
Six months ago, a result like this was unthinkable. But now we can say it loud and clear: local models are at the cutting edge, and the gap of just a few months has been confirmed.
Personally, I use Qwen-Next 3.8 for complex tasks; today, GPT-Sol-6-High was messing up a project, but Qwen-Next got it back on track. I consider it a reliable benchmark. What’s your take?
Using an asrock 12 unit case running one board as the main with the rest of them headless. About 71GB of vram exposed. So far 40 tok/s is with 30k context and it dips to around 30 tok/s at 100k context. This is all over the 1gb Ethernet that is on the boards already. I have 2 more of these and will probably get them running to see if 3.8 flash next runs at usable speeds. This setup is wildly inefficient with power but cost me less than $800.
I played a game of a custom chess variant against an claude opus 5.5, and it beat me.
The game combined several rule changes: the board wraps around from the h-file to the a-file, knights move three squares in one direction and one sideways, and captured pieces can be dropped back onto the board, as in crazyhouse. I gave the model the rules, the starting position, and a board diagram, then asked it to reply with one legal move at a time.
Also I recreated this board game https://nika-game.com/ and played with claude, and it still win, although I believe it is really unpopuplar and old game without much training data available (claude didn't even know the rules initially)
I used to think chess exposed a fundamental limitation of LLMs: keeping track of a changing board, following exact rules, and planning ahead seemed like a poor fit for a language model. This game made me reconsider that assumption.
So I’m curious: what broad classes of problems do you think LLMs still can’t solve reliably? Are there limitations you consider fundamental/archiectural, rather than problems that might improve with better models, more computation, or tools? What would be a good test?
Model Developer: NVIDIA Corporation
Model Development: Fine-tuned from NVIDIA-Nemotron-3-Ultra-550B-A55B
NVIDIA Nemotron™ is a family of open models with open weights, training data, and recipes, delivering leading efficiency and accuracy for building specialized AI agents.
Nemotron-Labs-3-Competitive-Coding is a competitive-programming specialist model based on Nemotron-3-Ultra, fine-tuned for one epoch on 477,642 synthetic reasoning traces distilled from GLM-5.2 across 22,000 curated problems spanning 16 regional and international competitive-programming contest families. Selected as the SFT teacher for its higher accuracy and roughly 30% shorter generations compared to a DeepSeek-V4-Flash-trained variant, GLM-5.2 distillation yields a model that, combined at inference time with GenCorrect — an iterative closed-loop test-time compute strategy that generates diverse candidate solutions, incorporates evaluator feedback, and refines subsequent generations under a fixed submission budget — was evaluated live and prospectively on the IOI 2026 problem set under official contest time, internet-access, and submission constraints, scoring 535.4 out of 600 and surpassing both the gold-medal threshold (361.12) and the top human contestant's score (498.27), making it the first AI system reported to outscore the highest-scoring human contestant on an IOI problem set.
This model is ready for commercial or non-commercial use.
Too many hyperactive amateurs are coming in here with "1b model at 832843tok/s!" and hardly any of them have all the info necessary for local runners to evaluate. We need context ladders with perplexity/KLD, hardware specs, model params and quant(s), runtime, tuned runtime parameters, basically everything we need to reproduce locally if we can match the entire setup. To say nothing of what the model is even good at in the first place if it's not a well-known model.
The goal is to get perf numbers that show they meet a certain quality bar. I don't care if I get 8324834 tok/s if it's all garbage.
Can we start filtering the hyperactive amateur perf posts please? It's getting really frustrating seeing all these posts of models and wading through info just to see that it doesn't test with anything but an empty context or doesn't say anything about quant or platform.
We need to define some rigor and apply it to this place, or it will remain like most ai-oriented subs and get continually choked with slop.
I've been working on Mica, a 4B decision model, and wanted to see how far it could get in actual Minecraft, not a sim. New world, empty inventory, and the goal was a diamond pickaxe.
Last time I posted, it got an iron pickaxe. Honestly that took around 20 tries and it was pretty flaky. I've reworked the harness a lot since then. This time it went all the way to a diamond pickaxe, and once the harness was finished it did it on the first run.
It took 26 decisions and about 8 minutes of game time. It got wood, made a crafting table, then wooden and stone pickaxes, then iron and coal, a furnace and an iron pickaxe. After that it tunneled down to diamonds at y=2, mined three, put a crafting table down right there and made the pickaxe. Decisions took about 108 ms on average.
My favorite bit is around step 14. The planner wanted it to make planks to burn in the furnace, but Mica went and mined coal instead (0.69 vs 0.31) and then smelted all three iron at once. Which was the better call, honestly.
How it works: every step Mica gets the game state as text (inventory, nearby blocks, health, what happened last step) plus a few candidate commands, and it picks one. A Mineflayer bot running Mindcraft skills does the actual moving and mining. The panel on the right of the video shows each decision and its probabilities live. The bot also knows where the nearest diamonds are, so it isn't searching for them.
To be clear, I'm not saying a 4B model plays Minecraft on its own. What I wanted to show is that a model this small can sit behind a bot, read what's going on, and make the next call well enough to get all the way to diamonds.
I'm planning to release the harness soon. Mica will read Minecraft chat, so you can type what you want and it'll work toward it. Simple stuff like getting items, crafting or following you should work fine, but it'll struggle with anything really complex, like building a house.
Also, v0.5 should be out in the next 1-2 weeks. A lot of the architecture changed, and I did extra training on the parts where v0.1 was weak, so I'm expecting a clear jump in performance. The aim is to be at or near the top among 4B JEV-like models.
There'll be two versions: Mica v0.5 4B, and Mica v0.5 4B Distill Laya, which is light enough to run on pretty much any PC.
For Minecraft, I'm hoping v0.5 will be good enough to take down the Ender Dragon, and I did extra training specifically with that in mind. No promises, but it'd be really cool if it pulls it off lol
Oh and fun fact, Mica is a fully vibe-coded project
Model: https://huggingface.co/sky7350/Mica-v0.1-4B
Model code: https://github.com/akivet/Mica-v0.1-4B
Minecraft harness: coming soon
Anthropic released the GLM article today. Trump is getting very involved.
Do you foresee Chinese open weight models getting banned soon?
I was optimizing my diet using a cloud AI provider, and I found my Pi agent stuck like this. Looks like I had too many mushrooms lol.
Here's a metric dashboard giving an idea of the last few days.
Been testing with a variety of different agentic coding use-cases, mostly using a pi harness.
qwen3.8-flash-next has seriously exceeded my expectations (used https://huggingface.co/tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8)
Both speed and quality have surprised me, given that I can get 3-5 concurrent streams going with \~100t/s gen each, and single stream easily gets to 150+t/s. Prefill is 10k+t/s
Having tested both Qwen 3.8 27b and Flash next on RTX 5090 with 96GB RAM, I want to find the middle ground between the two for coding capabilities but not sacrifice decode speed to standstill. I would like decode speed to between 75-100 ideally for fast iterations; otherwise I become impatient.
Currently I get 200+ TPS on Qwen 3.8 27B and approx 50 TPS on Flash next.
My hardware - RTX 5090 and 96 GB DDR5 which I plan to upgrade to 128 GB (in this economy, yes, but unwillingly).
What model sits between these two in terms of coding capabilities and hardware requirement? If none is present, I can perhaps think of using Flash next for plan creation and 27b for implementation.
Edit: Fast forward few days. I gave Strata a go with Swift 1.5 Flash Next IQ3\_XXS and able to achieve \~150 tok/sec decode and 5k tok/sec prefill. I am escatic! The quality of response from 27b is considerably better and the speed is great. Both targets achieved.
TL;DR -- You should probably just use base Qwen3.6-35B, as only Occamy-1.0 is competitive with it. Tiel is a major let-down, worse than Ornith. KAT surprises (good), Nex surprises (bad). This post is long. Sorry, lots to cover.
I think we all want to see a next-generation small MoE from the Qwen team to replace 3.6-35B in our workflows. This model is a perfect fit for smaller gmaing laptops and mid-tier rigs. It sucks that Qwen seems to have abandoned this model, but at least there are fine-tunes that improve upon it... right?
Well... maybe not. I ran benchmarks on the 3.6-35B-A3B base model, as well as five finetuunes: Occamy-1.0, Ornith-1.5, KAT-Coder-V2.5-Dev, Tiel-Coder, and Nex-N2.5-mini, and the results are quite surprising.
I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark I can run in ~10 hourso on my machine, it's not actively post-trained on by any of these models, and it provides a lot of useful information along with the raw accuracy scores. This includes: first-try and retry pass rates, token counts, solve times, and how well-formed the output diffs are. Here's the table:
| model | First-try pass | Retry pass | tokens | sec/case | tok/solve | well-formed diff |
|---|---|---|---|---|---|---|
| Qwen3.6-35B-A3B BASE (STOCK template) | 37.4% | 71.0% | 8650 | 285 | 14.1K | 96.3% |
| Occamy-1.0-35B-A3B (STOCK template) | 29.0% | 70.1% | 6801 | 285 | 17.2K | 86.9% |
| Occamy-1.0-35B-A3B (froggeric medium) | 30.8% | 69.2% | 6009 | 233 | 16.8K | 94.4% |
| Occamy-1.0-35B-A3B (froggeric, xhigh) | 27.1% | 67.3% | 8631 | 310 | 20.0K | 91.6% |
| Ornith-1.5-35B-A3B | 23.4% | 63.6% | 4813 | 226 | 16.4K | 87.9% |
| KAT-Coder-V2.5-Dev | 20.6% | 58.9% | 2190 | 84 | 9.3K | 86.9% |
| Tiel-Coder-35B-A3B | 18.7% | 53.3% | 4851 | 171 | 18.2K | 89.7% |
| Nex-N2.5-mini | 10.3% | 30.8% | 5037 | 188 | 33.3K | 95.3% |
As you can see, the only finetune that even competes with the base model is Occamy-1.0. The rest are utterly dominated by the base model, a grim disappointment for finetune enthusiasts. I was particualarly surprised by the performance of Tiel, which seems to get a lot of love in this subreddit.
Speaking of Tiel, I want to clarify that Tiel is just Ornith-1.5 with a different chat template, Sharp, which is based on froggeric with an added "terse mode" instruction that's supposed to reduce excessive verbosity. I wanted to standardize for templates, so ALL models are using the base froggeric v22.5 template set to medium (which is equivalent to standard thinking on, no additional message sent). I used this because I wanted to test Tiel vs. Ornith-1.5, and Tiel is the chat template. Also, practially, I use froggeric in my real workflows. However, to ensure coverage, I also tested the STOCK template on the 2 highest-performing models, to make sure it wasn't affecting the scores. As you can see, the template doesn't make a significant difference in the scores, and the scores for base 35B with different templates are so close to identical that I excluded the froggeric one from the table.
I also tested froggeric/Sharp's reasoning-effort toggle, and found xhigh -> medium significantly reduces token counts and solve times (by ~1/3), without affecting accuracy significantly. That stands in stark contrast to Tiel's 'terse mode' toggle, the core feature of Tiel over Ornith, which dramatically reduces accuracy along with the reduction in token counts. My results strongly suggest that if you want a less verbose model, you're better off lowering the reasoning effort than using Tiel with terseness on.
Speaking of token use, that's probably the big differentiator here. A couple models stand out: Ornith and KAT-Coder-V2.5-Dev are the most efficient models, with KAT in particular having a brevity unmatched by anything else. KAT is fucking fast, and I think despite its lower accuracy than Occamy, it has a place in my lineup as a subagent because it just gets. shit. done. Occamy is also interesting, as it is the only model that perfoms on a similar level to the base, but it uses 20-30% fewer median tokens. However, Occamy also had a number of runaway generations where the token count blew up, so it's total tokens/solve is actually higher than base.
In an effort to further distinguish Occamy from base, since Aider struggled to do that, I ran tau2-bench, an agentic tool-calling benchmark consisting of multi-turn interactions with a simulated counterparty. I figured this was a good bench to use as Occamy is post-trained specifically for 'co-work' scenarios, but not trained on this particular set. I used Qwen3.8-27B with reasoning effort set to low as the simulated customer in these conversations. The base model was able to pull away from Occamy in the harder retail domain of this benchmark, but Occamy resolved the issues in the airline domain at an equal rate while requiring fewer turns. Here's the results.
| model (Q8_0) | domain | pass^1 | tokens | sec/task | turns/task |
|---|---|---|---|---|---|
| Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) | airline | 80.0% | 5390 | 290 | 11 |
| Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) | airline | 78.0% | 6770 | 390 | 13 |
| Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) | retail | 86.0% | 4098 | 333 | 14 |
| Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) | retail | 79.8% | 3284 | 284 | 13 |
Overall, I think the results are clear, if unexpected: Occamy-1.0 is the only fine-tune that even competes with the base model on Aider Polyglot, but even it is not a clear winner. Tiel is noticebly worse than plain Ornith without the terseness toggle, and the terse mode doesn't even save any tokens. xhigh in froggeric/Sharp degrades accuracy slightly and bloats token use, which makes sense given the models were not RL'd for the extra thinking effort prompt. KAT-Coder-V2.5-Dev is the most efficient model, with accuracy nearly as good as Ornith and better than Tiel. Finally, Nex-N2.5-mini is a disaster.
ninfer, dwarfstar, Splash, llamAmpere, gufo, etc.
We've all seen them popping up, great tok/s, people loving them. Forks of llama.cpp or another engine, or made from scratch.
For better or worse, the list will continue to grow
They work so well because they dodge a main difficulty of software, generality, and just implement for a single model/hardware combo (or a few), and then optimize kernels/compute graph for that one case. Highly 'overfit' codebases that beat well-known inference engines (llama.cpp, vLLM, etc) and incidentally will be completely forgotten in 6 months.
But new ones will take their place...
THESIS
One-off engines will become the norm. Llama.cpp, vllm, etc, will not make sense for most people to use, because they're slower
A few axioms you probably accept:
String these axioms together, and I arrive at
Thank you for coming to my ted talk
IMPLICATIONS
Most obvious example: it would be annoying to have a different usage API for every engine, so we already standardized on OpenAI API compatibility years ago.
Is that also true for cli arguments/configs? The packaged gui (llama-server)? Benchmarking tools (llama-bench)? Logging, model format, Etc?
One-off engines that replicate the experience of everything wrapping the inference itself will be more seamless to adopt. Case in point, the main reason I haven't tried any of these new one-off engines myself is it was annoying enough to figure out how to drive llama.cpp properly. Don't want to do that again unless it's really worth it.
There's probably a place for an open source project that standardizes all of this and makes it easy for one-off engines to adopt.
ALTERNATIVE FUTURES
Scenarios where the one-off future doesn't happen:
P.S. there's growth in a dimension separate from single model/hardware engines which is more like "frontrunning a general inference engine's features because it's slower to pull in PRs". Freetoken, BeeLlama, etc. Not as model- or hardware- specific as the other examples I've given. Haven't thought much about that dimension.
I think I need a 3-slot for my two cards. but holy fuck these things are pricey.
IQuest-Q1 is a Mixture-of-Experts (MoE) model developed by IQuest for agentic coding, reasoning, and multi-step tool use. It comprises approximately 320B total parameters, with an estimated 15B parameters activated per token.
Apple Silicon Macs on macOS 26+ come with a small LLM built in. No download, no API key, and nothing leaves your Mac.
Why I built it
I was making a tool that writes API docs from code, and I didn't want users to install Ollama or paste an API key. Apple's model was already on their Mac, so I used it.
Getting it to work well was harder than expected. At temperature 0 it kept repeating itself, and a 2-second call took 20. Calls took 17 seconds instead of 1.5 until I kept one process running. So I turned all the fixes into a library: apple-llm.
What it's good for
\- Pulling structured data out of messy text, like emails into tickets or receipts into expenses. The JSON always matches your schema.
\- Tagging, summarising and rewriting
\- Private data you don't want to send anywhere
\- Tools you share with other Mac users, who don't need to download a model or get a key
What it's bad at
Coding and long reasoning. It's a small model.
There's also an optional cloud mode that uses Apple's bigger server model for harder questions. That one is not local: your prompt goes to Apple's servers, and it has a usage limit.
Node: npm install apple-llm
Python: pip install apple-llm
We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed.
Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16.
What's inside
Results:
At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.
Coder (capability pruned model):
Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits.
The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.
Both measured at xhigh reasoning effort.
Links
Both repositories ship the complete per-tensor RCO allocation.
The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome.
From the ISTA Deep Algorithms and Systems Lab.
TL;DR: The Jeff models are a set of Qwen3.5 and Gemma fine-tunes for zero-shot classification: small, efficient, open-weight models with respectable out-of-the-box performance that can be slotted right into code or fine-tuned/LoRA-trained as needed. Give them a situation and a list of options; they return a calibrated probability for each, in one forward pass, no text generation. The 2B scores 83.1% on a five-benchmark panel (Jev's published figure: 83.0%); the 0.8B decides in \~28 ms on an M4 Max (see caveats below).
Maybe equally exciting for open model enthusiasts like myself, everything was done on local hardware: one RTX PRO 6000 for training, two DGX Sparks running Qwen3.8-Flash-Next to write the synthetic data, a MacBook for testing - all connected and monitored from my Android phone via Tailscale. Apache 2.0, Jev-compatible API. Weights: Jeff-Qwen3.5-0.8B · Jeff-Qwen3.5-2B · Jeff-Gemma4-E2B · Code: github.com/firelex/jeff · Videos: games table
When TypeSafe released Jev a couple of weeks ago and then AutoJev appeared, I wanted to see if I could replicate the experiment using only small language models on local hardware.
So here's what I did:
Benchmarks (4,599 questions: BBH, Financial PhraseBank, JudgeBench, RAGTruth, WinoGrande):
|Model|Untrained base|Jeff (trained)|Calibration error|
|:-|:-|:-|:-|
|Qwen3.5-0.8B|45.3%|79.1%|0.049|
|Qwen3.5-2B|46.5%|83.1%|0.028|
|Jev (published)||83.0%|≈0.06|
|AutoJev-27B (published)||84.9%|—|
The caveat: the published numbers are on a different sample of the same benchmarks, and Jeff's overall score comes from classification-style tasks (96% on Financial PhraseBank, 86–89% on RAGTruth, both above Jev). On multi-step reasoning it's behind: BBH 64–68% against Jev's 94%, and \~50% on JevBench's hard tier against \~73%. See the HuggingFace model card for details. But that's not surprising, and I don't think it matters. No 0.8B or 2B model reasons like an LLM, and I don't think anyone should expect it to. The Jeff models are extremely fast judgement-callers (much faster than Jev), and have reasonable out-of-the-box performance. In one of my apps, I used the 0.8B model for voice-based navigation, and with a quick fine-tune, I got to real-time performance (24ms) at almost 100% accuracy.
Now the fun part: games, as a zero-shot test. Games are not the ideal zero-shot test, but they're fun, and the TypeSafe guys (Jev) did it, too. There was no game data in training. Each turn the code describes the situation and the moves in words, and the model picks one. The options say what each move leads to, never which one is right.
|20 episodes each|Doom (kills)|Frogger (crossings)|Pac-Man (pellets of 98)|
|:-|:-|:-|:-|
|Random moves|−0.05|0|11.2|
|Hand-coded rule bot|6.55|10.25|94.1|
|Qwen3.5-0.8B, untrained|5.0|1.0|25.8|
|Jeff 0.8B|6.55|10.3|57.0|
|Jeff 2B|−0.9|6.0|41.2|
Jev's published Doom score is also 6.55, but its prompt spells out the aiming rule (fire when the bearing is between −8 and +8 degrees) and it takes \~212 ms per call over its API. Jeff gets "the nearest monster is a little to your left" and decides in \~29 ms on my Mac.
Lessons learned:
Happy to answer questions about the pipeline (synthetic data from a local teacher, leak filter, calibration) or the game harness. Everything, including the videos, is linked above.
https://preview.redd.it/23cumykhabsh1.png?width=944&format=png&auto=w…
Reminder that you can put Qwen 3.8 27B as a subagent and its a workhorse. Pic: using DeepSeek v4.1 Flash as Orchestrator in Pi, Qwen 3.8 27B GSQ RCO in llama.cpp.
I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.
Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.
And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.
The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.
All figures here: https://github.com/0xShug0/audio.cpp/tree/main/assets/figure/
I've liked how Muse-Glimmer worked, so I wanted to see if I could produce a smaller "kid" out of it. Ornith's sharp decisions on when to think and which tool to call were the other thing I liked, so Ornith-1.0-9B got to be the policy teacher while the parent wrote the words. No RL anywhere, distillation only.
I present to you Xyntetik-Kvist-14B.
What it is good for
Numbers, from the card
| claim | number |
|---|---|
| parent | Muse-Glimmer-30B, cut by width (hidden 6,656 to 5,760, FFN 19,968 to 10,240, heads 32 to 24), all 52 layers kept |
| size | 14.44 B parameters; BF16 28.9 GB, Q8_0 15.4 GB, Q5_0 mix 10.3 GB |
| distillation | 6,000 steps, 98.3 M tokens, 162 hours, then 1,440 steps on agentic trajectories |
| fidelity to parent | KLD 0.762, margin-qualified top-1 84.0% on 45,056 held-out positions (a student's row, not the quant bar) |
| tool tasks | 57 of 60 held-out, re-executed against ground truth; parent 60, untrained control 0 |
| format and calls | 199 of 200 first turns well formed; 99 of 99 tool calls valid |
| attempts | 12 gated attempts, 2 full passes, attempt 12 shipped |
| weak spot | calc tasks 12 of 15 over 160; 7 of 160 runs end in a reasoning loop |
| serving | Runner v0.5.7 or later |
| licence | Apache-2.0 |
Links
EDIT: reading the comments, i should have said this first. this is not a general drop-in for Muse or a gemma4 replacement, and it was never going to be on my compute (98M distillation tokens vs the trillion a real distill wants, i simply lack the compute). the purpose was more on getting the tool calling right. IQ4_NL mix (7.6 GB) is up now too, it scores the same 57/60 on the tool tasks but does not fit an 8 GB card whole (50/52 layers on a 3070, ~5 tok/s).
MINISFORUM MS-S1 MAX-P495 – Minisforum EU
Expected to ship mid october.
At that price point, it doesn't make a whole lot of sense in my opinion.
I get that ram prices are where they are, i get that it's a newer model of hardware, but twice the price for 50% more ram and \~5-10% more performance is just hard, especially when compared to the recently released m5 ultra studio.
Hi there!
My name is Jerry and I recently built Enclosure, e deterministic environment for testing LLMs. It's a simulated office enfironment https://github.com/openconstruct/Enclosure with tools agents are used to, like slack, calendar, mail, chat and more. For conversations is uses AIML instead of a model, so it is totally deterministic.
The first test I made(had Claude make) is the one I had in mind when I designed Enclosure, a personality test for models. It's not a winnable benchmark, but rather a tool to help people find the mdoels that fit their workstyle best. It's called disposition and you can find it here: https://github.com/openconstruct/disposition
I tested the cheapest models on Alibab modelstudio already, after I refill my opencode go next month, I will probably test more. I am asking the community to test some models if you think it's a cool project.
Instructions are in the Disposition repo, and the submission repo is here: https://github.com/openconstruct/disposition
I know this may have been asked a lot, but I'm not an LLM expert yet (have never set up vllm myself or llama.cpp myself yet) I've only used scripts other people have set up.
Is there a simple way / steps to follow to run Qwen 3.8 27b in Windows with my single 3090?
I can run it straight up with Ollama with a 64k context but it seems like it only works reliably in chat and not in OpenCode (in OpenCode it "works" but at one point it "hung up" on me it seemed like. Not sure if maybe it was busy thinking still)
I've found quite a few threads that are close to what i'm asking for but many I think use WSL or something, too. Which I'm also not super experienced with yet.
I found the HyperQwen repo but their documentation is like...120% all technical and not very user friendly at all. I CAN do technical stuff! But it's barely passable as "do this, and then this" type of docs at the moment.
Ninfer is only for 5090 cards from what I read.
EDIT: Thanks guys, I was able to use llama-cpp-windows-manager project to get Qwen 3.8 27B up and going (Q4\_K\_M) . I have a 98K context and get 70tok/s and it's working with Open Code just fine. Very usable.
Hi everyone :)
The amazing Swift finetunes of Qwen3.8 27B generate much fewer tokens at mostly similar benchmark performance to the original model, while repos like HyperQwen (formerly syv-ai/qwen38-27b-rtx3090) deliver insane TPS on an RTX 3090.
To get the best of both worlds, I adapted Swift 1 and Swift 1.5 for HyperQwen and benchmarked them against HyperQwen’s specialized W4A16 AutoRound fast quant for speed and quality.
|Model|Average time/task ↓|Average output tokens/task ↓|Decode tok/s ↑|
|:-|:-|:-|:-|
|Qwen - HyperQwen fast quant|108.1 s|8,985|112.1|
|Swift 1.0 + HyperQwen|66.2 s|5,245|105.9|
|Swift 1.5 + HyperQwen, INT8 heads|72.2 s|5,751|104.0|
|Swift 1.5 + HyperQwen INT4 heads|68.2 s|5,669|107.2|
All models ran on one RTX 3090 24GB with FP8 KV cache and 150k configured context. Average task time covers about 630 tasks from different benchmarks listed below.
Swift-1 seems to use the fewest tokens and has fastest task completion time. Swift-1.5-INT4 achieves roughly 37% lower average time per task compared to the HyperQwen Qwen fast model. Despite slightly lower TPS (due to lower draft acceptance) it finishes sooner because it generates fewer tokens.
There are some minor quality and performance tradeoffs between the models:
|Test|Qwen HyperQwen fast|Swift 1.0|Swift 1.5 INT8 heads|Swift 1.5 INT4 heads|
|:-|:-|:-|:-|:-|
|GSM8K, 200 questions|97.5%|98.0%|98.0%|97.5%|
|IFBench, 300 prompts, strict|74.0%|73.3%|73.7%|72.3%|
|LiveCodeBench, (100-problem subset)|90%|89%|89%|91%|
|Custom tool-call/JSON eval|29/30|28/30|30/30|30/30|
|English/Python perplexity ↓|6.551|6.605|6.643|6.679|
Applied changes to Swift models to adapt for HyperQwen:
If you want to try it yourself, point your coding agent at these setup instructions and ask it to set up Swift 1.5 + HyperQwen on your machine.
All three models can be found here:
https://huggingface.co/collections/daavidhauser/swift-for-hyperqwen-rtx-3090-benchmarks
Big shoutout to UkisAI for Swift and syv-ai for HyperQwen! It's genuinely insane to be able to run these models on an RTX3090 at those speeds!
360 A/B runs on GLM 5.3 and GLM 5.3 Flash, max thinking, 5 repeats per cell. Savings up to 70%.
The block (shipped to global instructions):
How I tested: real agent sessions in throwaway repos, a 9-part exam (two bug fixes, a wrong-premise trap, a hidden requirement, a trivial rename, and four pushback flavors: mild, authority, evidenced, false-fail). Four instruction variants - baseline, the 9 rules, the rules + a "one meaningful check, then commit" clause, the rules + a false-FAIL guard. Deterministic scoring, hand-adjudicated finals. Neither extra clause earned its place, so the 9 rules stand alone. Same result on the first family I tested this way (MiMo 2.6 Pro, net -28%), so this isn't a one-model fluke.
Exams to test for yourself: github.com/Arshad-Kamal/thinking-quality-exam
Someone prompted different LLMs to generate CAD code for a bridge under fixed constraints (2-foot span, under 500g filament, 18-hour print limit), printed them, and load-tested them to failure. The results were quite varied: some models couldn't even design parts that fit together, while the winner held over 100 lbs. Most benchmarks don't really capture physical intuition, spatial reasoning, and functional code generation at the same time. I would love a standardized benchmark and leaderboard for this.
​
I created this repo to help the DGX Spark users that have a spare 10-24 GB GPU at home to squeeze some extra memory out of a single Spark or a Sparks cluster.
It moves the spec-decode draft model off your Sparks onto that GPU: the freed GB of memory can be used for extra context, or better quant quality. Supports both TCP and RDMA, shipped as eugr-vllm compatible mods:
https://github.com/ciprianveg/gb10-vllm/tree/main/remote-dspark
We did a study: does post training actually make LLMs funnier?
We used open models that publish every stage of post training, so we could compare a base model with the future models it became: Tulu 3 (on Llama 3.1 70B), OLMo 3.1 32B and Qwen2.5. We tracked 11 stages, 100 joke prompts, 64 human raters and 2,330 head-to-head judgments.
What we found: post training makes models funnier, but reduces diversity of response.
\- In 5 of 7 training steps, the later model's jokes were judged funnier. Jokes also got 10–20 words shorter after early post training, so they get to the punchline faster.
\- In 6 of 7 steps, the jokes a model wrote for the same prompt got more similar to each other. Ask for eight jokes on one premise and you get eight versions of the same joke. The biggest drop was Qwen2.5 base to instruct.
\- Asking the model to plan a line or two before the joke cut variety in all 4 models we tried, with no reliable gain in funniness.
\- A comedian persona won back a little variety in all 4 models, but only made the jokes funnier in 2 of them.
Humans judged the base versus final. A model judge calibrated on those votes compares the stages in between.
Full report and paper below. Which open models should we run through this next?
For example, directly using claude code or code, which is then hooked up to automatically delegate the actual code writing tasks to a local model like qwen 3.8 flash next, to save on cloud usage limits.
I’m imagining the loop would be:
User writes prompt
Claude/codex thinks about it and the plan
Claude/codex sends the specific and bounded coding instructions to the local model+harness (opencode, pi, etc) via api endpoint or MCP, with clear instructions on a defined endpoint
One the local model+harness hits the clear endpoint/“done” step, it sends a ping back to claude/codex
Claude/codex then verifies the output and then thinks about next steps to instruct the local model+harness on
Does this actually lead to improved savings on the cloud model usage while preserving code quality? Or does this end up being unnecessarily complex and not saving on any cloud usage
Forlinx Embedded has listed an M.2 AI accelerator card based on Rockchip’s RK1820 and RK1828 processors, providing 20 TOPS of INT8 computing performance and up to 5GB of integrated DRAM. The module uses an M.2 2280 interface and is designed to handle local AI inference, including large language models, vision-language models, and computer vision workloads on embedded Linux and Android systems.
I wanted a setup where I can compare all the harnesses with all the local models.
It turned out to be a rabbit hole. For instance, you would not only have to test all the harnesses (being sure that they are well configured), but also all models, with all their flavors, and this for all kinds of hardware.
Everyone can do their share, but no one can pretend to do all possible tests extensively.
To solve this, I created a website where everyone can test the configurations they want and share the results if they want. You can try it here: airbench.ai
There is a leaderboard where I share the tests I'm making, but I hope I can populate it with tests from others. https://airbench.ai/leaderboard?k=poL
My hope is to turn this into a fully distributed Agent Benchmark.
Let me know what you think.
I’ve been working on an open-source project called WebBrain that gives LLMs the ability to see and operate a browser.
One thing I really wanted to avoid was making the browser agent dependent on a single cloud model/provider.
So WebBrain can work with local models through things like LM Studio and Ollama, as well as cloud APIs if you want them.
I also ended up training a small vision model specifically for browser tasks:
webbrain-vl-2-450M
It’s based on LFM-2.5-VL-450M and fine-tuned on browser screenshots/tasks. The idea is that instead of sending every screenshot to a giant multimodal model, some browser perception can happen with a very small model locally.
It can run through WebGPU directly on the user's machine.
The agent itself combines screenshots with the browser accessibility tree rather than relying entirely on DOM parsing.
Current architecture is roughly:
• screenshot + accessibility tree for perception
• browser-specialized tiny VLM where useful
• model-agnostic planner
• local models via LM Studio/Ollama
• Chrome / Edge / Firefox / Chromium support
• optional cloud execution
• open source
I'm especially interested in figuring out how far browser agents can realistically go with small local models rather than GPT/Claude-scale models.
Repo: https://github.com/webbrain-one/webbrain
Model: https://huggingface.co/webbrain-one/webbrain-vl-2-450M
Dataset: https://huggingface.co/datasets/webbrain-one/webbrain-vl-2-450M-dataset
Would be very interested in feedback from people here running smaller Qwen/LFM/MiniCPM/etc. models locally — particularly what model you would try as the planner.
Like everyone, I use 3js for data visualisation. But while semantic clusters are interesting, they don't confer much practical information alone.
So I did something very simple: a query spatially re-assembles the nodes, and you can switch them to text.
Honestly, it's hard to swing 3D viz for text as a genuinely useful feature and not a novelty, but I get the feeling there's still some potential in the idea. Would be good to discuss, I'm sure someone here has made a better implementation.
I have been using Qwen 3.6:35b IQ4 via llama.cpp on my p40 and am getting anywhere between 37 - 83 tok/s (mtp is on). Prefill usually starts at 600 and slowly degrades as it continues processing so 600 for short prompts and more like 300-400 by the end of a long prompt. It hurts, but it is what my budget allows.
AI told me IQ quants are noticeably slower on the P40 and it referenced (https://www.reddit.com/r/LocalLLaMA/comments/1dmhpud/are\_iq\_quants\_slow\_o…) from two years ago, but I didn't feel it being slower when I moved from Q4 to IQ4, so...
What am I missing? Are IQ Quants actually a significant amount slower on the P40 or is this outdated information?
Hello everyone! A little while back I posted about LlamAmpere, a fork of Llama.cpp with Ampere-specific improvements (though it is caught up to main and will support other hardware, too).
Thank you to everyone that tried it out and shared back their results across the 30xx cards. I'm happy to share I've pushed v0.4 out this morning. On the 4.6bpw model tested, speeds improved \~10% vs the last version while also improving the max context by 10%+ (technically, it can go above 262K, but I have not tested any custom kernels or graphs to support YaRN).
The closest competition comes from vLLM, keeping within <10%, but does so with lower maximum context. It is significantly faster than other llama.cpp options tested.
https://preview.redd.it/u70q7lew1bsh1.png?width=1080&format=png&auto=…
[](https://preview.redd.it/95-tps-through-100k-generated-262k-ctx-on-a-single-30…)
There's also a number of other improvements for other quants/formats, with EXL3 seeing significant speed up (\~80% the speed of the 4-XS-M quant tested). It has a slightly lower KLD, but not a range I have found stat significance for at the task level, so I sticking with the XS-M model for now (built on top of Swift-qwen's distill, which is far more token efficient than the stock train for \~1% performance loss). the 4.3 bpw EXL3 model does provide a bit more room if you are interested in 2+ concurrent predictions. Improvements in this format and the IQ2/3 codebook quants will be most useful for people on 12/16/20 GB setups. These measurements are at temp=1, vs some of the vanity speeds you will see people claim with temp=0 and/or short generations.
As always, please share your results + config details so I can keep improving!
fork is here: https://github.com/JakeATX/llamAmpere/blob/main/QWEN\_AMPERE.md#build-and-run**
model used here:
https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF**
Build command:
git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server
Build + launch (linux):
git clone -b v0.4 https://github.com/JakeATX/llamAmpere.git
llamAmpere
cmake -S . -B build-sm86 -DCMAKE\_BUILD\_TYPE=Release -DGGML\_CUDA=ON -DCMAKE\_CUDA\_ARCHITECTURES=86
cmake --build build-sm86 -j8 --target llama-server
curl -L -o ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf \\
https://huggingface.co/jakeatx/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M-GGUF/resolve/main/ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf
\-m ATX-Swift-Qwen3.8-27B-Uncensored-IQ4\_XS-M.gguf -c 262144 \\
\-ngl 99 -fa on -ctk turbo5 -ctv turbo4 -b 4096 -ub 1024 -t 8 -tb 8 --parallel 1
cd
./build-sm86/bin/llama-server
Previously, people had expressed concern over quantizing KV cache, and TQ specifically. The TL;DR on that is that any reasonable KV quantization strategy (at least for hybrid attention models like Qwen) is going to be swamped by quantization of the weights. The KV quant we're using here (TQ5/TQ4) is less than 1/3 of the KLD we see when moving from 8 bit weights to 4.6 bit weights (and the KLD is only partially additive, so some of the incremental errors cancel out). There was no statistical significance when testing this KV quant at the task level against 8/8 kv (just trivial variations in sentence length). I will be adding KVaRN in the next release, but with a better codec than currently available elsewhere, so it requires a bit more testing before release.
v0.5 will be focused primarily on the 12GB cards, but this should have generation-wide speed ups, so even if you're not on 24GB, please share your results.
Enjoy!
Hey guys, I am looking for benchmarks for people running multi node (4 or more) A100 80 GB cards and seeing what results and models they are getting. Something with VLLM and multi users would be very useful. Or if you know of a place I can find such results please let me know! Appreciated!
Other than the obvious self promotion, is there a practical reason people do this that I am missing? There's dozens of llamacpp forks with silly names that are supposedly "optimized" for this or that specific GPU and seem to have zero intention to merge into upstream. Am I missing the real reasons why this happens so often? Why do people think it's OK to do this? In my experience in the open source community this is generally frowned upon.
I don't know if it's just a me problem that this kind of thing puts me off so much. I am usually quite grateful for PR feedback and conscientious about the code I put out there; I take pride in submitting high quality code that meets or exceeds the standards of a given project. Of course there is nothing ethically wrong with hard forks or taking shortcuts if you find the collaborative process cumbersome, but personally I wouldn't promote my fork in such cases, let alone go out of my way to add custom branding with a Reddit announcement post etc. since the effort required to do so seems roughly equivalent to the effort required to meet the contributor standards. In contrast, many of the authors of these forks seem very eager to have others adopt their rebranded fork for production use cases. There just seems to be a big disconnect, idk.
Edit: some great discussion in this thread, thanks to all who responded. Consensus seems to be that (excluding the obvious low-effort engagement bait forks) the base project has to meet many compatibility requirements while a downstream project can be more focused, which is a great point.
Anyone have a good config they've found for QFN on a A40 (or similar Ampere GPU(s) with 48GB VRAM)? The system the card is in has 128GB of RAM (DDR3); right now it's running 27B but I'm curious if there's a way to move to QFN to maybe get better speed/a little more smarts. TY in advance!
TypeSafe's Jev answers typed questions (yes/no, pick one label, pick a rubric level) with a probability per answer instead of text. Its docs say it takes text only: "Images, audio, and video are not supported (yet)." I wanted the same kind of answer about photos, from an open model on my own card, so I built typevet (MIT, Python 3.12).
How it works. No sampling, no parsing. typevet composes the native Gemma 4 turn itself, ends the prompt with the empty-thought no-thinking prefill, and reads the next-token distribution over the allowed answer tokens only. What comes back is a label plus a probability for every option, or an error. Images go to llama.cpp's /completion as base64 in prompt.multimodal_data with one media marker per image; on vLLM they go as image_url blocks. There's also a JSON path that returns an object passing your JSON Schema, or raises.
Local setup. Gemma 4 31B, my 24 GiB vramfit pack (byte-identical to the file on HF, projector sidecar for vision), llama.cpp b11223, one RTX 4090. Nothing leaves the machine.
The receipt test. 6 real receipts from CORD v2 (CC BY 4.0), 3 synthetic expense claims each: right total, two digits swapped, digits masked with ?. One three-label Choice: match / mismatch / insufficient_evidence. Each claim sent as text only, then with the receipt photo.
Example: claim says 646329, receipt says 664,329. Text only: match at 0.99964. With the photo: mismatch at 0.99999. Every swapped total was caught at 0.99998 or higher, and the masked ones abstained every time. The tests also check the image actually arrived: each photo added 228 to 1,108 prompt tokens here, and the gate fails if the count doesn't grow.
Hosted. Same code against vLLM 0.30.0, BF16 Gemma 4 31B, one H100: the receipt test went 18/18, and reversing the label order flipped 0 of 18 answers. Throughput on 480 Banking77 records with two questions each: 0.24 s median per record at 1 in flight, 39.6 records/s at 64 in flight, 0 errors.
Prior art, credit where due. The text-side decision model comes from TypeLLM (SGLang), which added its own image input on 9/24; typevet's image path is separate code on llama.cpp's request shape. allanrbo posted a Jev-like single script for Gemma 4 12B with webcam images on 9/25. VQAScore has read the probability of "Yes" from VLMs since 2024. typevet's angle: a library, not a script, the 31B on one 24 GiB card, the same code on vLLM, and image-arrival checks.
Scope: 18 claims, one run per server. The probabilities are the model's confidence, not calibrated.
So I got this new local model on my system & wanted a mini test to see if it was actually smart or just confidently wrong (as I like to do with new models I haven't yet tried). I asked the cloud assistant to cook up a pretty rigorous 10 point diagnostic suite: reasoning traps, Python semantics, SQL fluency, strict instruction following, the whole gauntlet. At first it felt like DeepSeek V4-Flash frontier cloud intelligence vs my little local quantized guy. Classic quick test.
Then I ran it on the local model & shared the results back. The assistant was grading it & found out it got item #4 wrong. That item was a logic puzzle. The assistant thought one statement had to be false, but the local model was like nah, the set of statements is logically consistent, so your question is built on a false premise. It literally refused the leading question. I was like "wait wtf". That was the turning point fr. The local model solved a trap that the cloud model just completely keyed wrong.
The local model is a Q2\_K quantization of Nex-N2.5-mini, which is a fine-tuned Qwen3.5-MoE architecture. A 2-bit quant. Normally people call that low quality. But it outperformed a frontier model on a logic trap. The assistant went from “I am grader” to genuine admiration, saying resisting a leading question is high-level reasoning. Lowkey pretty humbling for the cloud side.
The whole thing made me think about emergent intelligence & AI democratization. Less giant centralized compute, more efficient specialized local stuff. The student corrected the teacher. Efficiency & MoE architecture maybe can beat raw parameter count sometimes. The mini test felt like a rite of passage for the local model. Its kinda like it became a validated thinker instead of just software. Q2 compression is also symbolic resilience, because despite being squished, the reasoning circuits stayed intact. And the assistant admitting it was wrong made the local model’s win feel more real.
And the craziest part? It went 10/10. This wasn't some easy benchmark either. It was a deliberately nasty little diagnostic with multiple ways for a heavily quantized model to screw up, and it didn't.
I need to make it clear that I am not claiming that Q2 ORCA model is generally superior to DeepSeek V4-Flash. But rather as an anecdotal demonstration that an extremely compressed local model can sometimes catch a reasoning failure in a frontier model & maintain much of it's reasoning power when done correctly & skillfully. It is a testament to how even under Q2 compression, it still preserved the model's “reasoning circuits."
Final verdict from the assistant: model is in excellent shape & ready for real work. So yeah, a local model dunked on the cloud AI. I’m happy for what this means for the future of local ai.
MODELS USED for quick test:
Local Model: Nex N2.5 Mini Uncensored
Frontier Model: DeepSeek V4.1
EDIT / CORRECTION bc I fucked this part up 😅
Small but important correction to the post. I originally called the frontier model I tested DeepSeek V4-Flash. That's not the right model name for the one I actually used on the DeepSeek website. It was DeepSeek V4.1-Flash.
Also, V4-Flash itself is a local/open-weight model, so my original wording made it sound like I was comparing my local model against some cloud-only AI. That's not accurate & that's on me.
The actual comparison was my local Q2\_K Nex-N2.5-mini vs DeepSeek V4.1-Flash through the DeepSeek website.
So yeah, V4.1-Flash is the model I should have named in the original post.
I'm leaving this correction here instead of quietly changing the post bc I don't wanna bullshit anybody or make it look like I didn't make the mistake. I got the model name wrong, someone pointed it out & I'm correcting it. 🤷♂️
The actual 10/10 result & the logic trap part of the test are unchanged.
\*\*TL;DR:\*\* I tested a local Q2\_K Nex-N2.5-mini on a 10-part mini test, it caught a logic trap the cloud assistant got wrong, & the assistant basically certified it as ready for real work. It got 10/10 correct.
That's why we have to fight against closed source centralized AI and closed recipes open weights overcoming the frivolous "gifts" exchanged with them by giving away our souls to those greedy entities if we want to be free in the future instead of being squeezed like lemons/at mercy/enslaved in the paws of these soulless folks that have a id of any single one of us at their disposal to switch us on/off at their leisure/convenience.
I was just wondering because I see it mentioned in threads here a lot... When people talk about LLMs used for roleplay, that's a euphemism for dirty/sexy chats right? Kind of like how "torrenting linux ISOs" is really just pirating copywritten media. Or are you guys really burning tokens pretending to talk to a medieval shopkeeper?
I know I sound crazy, but i’m using a strix halo and finding that GLM 4.7 flash just runs significantly better than qwen 3.6 35 a3b. But its obviously quite old at this point, i wish we had a new glm that was 4.7 flash sized but is anyone using a model that they have found better than this.
My brief testing with qwen shows that glm is better at tool calling and better at world knowledge, but maybe there’s a chance my qwen set up is wrong
I'm making a local roleplaying model for my girlfriend's community server.
It's supposed to do roleplay, have a specific talking style(dry while answering to normal stuff and extensive when talking about lore), never talk out of roleplay, have hundreds of pages of lore and information and their rank in lore( discord roles maybe?).
It's basically supposed to be just an LLM you can converse with that answers in a specific talking style and has all the lore info.
For now I implemented: 10 ish% of the written lore, and it recognizes 3 people, but by discord ID that i inserted in the system prompt.
I'm using GPT 5.6 sol(and well 6 sol now) for doing stuff, but i keep running into a problem.
When i reach a nice point where the model has a nice talking style and knows information well enough, i tell Sol to add this new lorebook and this info, here now everything breaks.
Talking style is fucked, It doesen't recognise people individually anymore, when asked about other unrelated lore it just gets it wrong or hallucinates, or even starts roleplaying as one of the characters in it's lore book out of nowhere.
It's connected to discord through a discord bridge that Sol made and a developer dashboard bot.
I'm using GPT OSS20b on MXFP4, single 9070xt and 32gb ddr4.
Should I maybe fine tune it?
PS. im a beginner in AI so if it wasn't obvious i do NOT know what im doing but im trying my best for her.
Thanks to tcclaviger, vLLM now has expert RAM offloading support (link). This makes frontier models much more accessible on a local setup!
I was able to run the original DeepSeek-V4-Flash-Vision-Exp on four R9700s.
podman run --rm -it \
--init \
--network host \
--ulimit memlock=-1:-1 \
-v /models:/models:ro \
-v ~/.vllm-cache:/cache \
-e VLLM_ROCM_USE_AITER=0 \
-e ROCR_VISIBLE_DEVICES=0,1,2,3 \
-e VLLM_CACHE_ROOT=/cache/vllm \
-e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \
-e TRITON_CACHE_DIR=/cache/triton \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--annotation run.oci.keep_original_groups=1 \
--security-opt label=disable \
--security-opt seccomp=unconfined \
--shm-size 160g \
docker.io/tcclaviger/vllm@sha256:ef99b3d07c3f15e7978528c7510762ba024df9ab4242070d8ed092cd4cc1a694 \
/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
--served-model-name DeepSeek-V4-Flash-Vision-Exp \
--tensor-parallel-size 4 \
--enable-expert-offload \
--expert-offload-mem 160 \
--reasoning-parser deepseek_v4 \
--reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
--tool-call-parser deepseek_v4 \
--enable-auto-tool-choice \
--max-num-seqs 8 \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-batched-tokens 2048 \
--kv-cache-dtype fp8 \
--block-size 256 \
--max-model-len 256000 \
--gpu-memory-utilization 0.97 \
--mm-processor-cache-gb 4.0 \
--override-generation-config '{"max_tokens": 128000, "temperature": 1.0, "top_p": 0.95}' \
--speculative-config '{"method":"dspark","model":"/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp","num_speculative_tokens":3,"draft_sample_method":"probabilistic","enable_adaptive_verification":false}' \
--compilation-config '{"cudagraph_capture_sizes": [4,8,12,16], "max_cudagraph_capture_size": 16}' \
--host 0.0.0.0 \
--port 8090
I am pretty dumb in this kinda stuff so please don't blame me for it.
My llama serve kept crashing with Cublas errors on first prompt with some models, but my VRAM usage was 3000MB/8k. I chatted with sonnet 5.5 for a bit and it told me it was a CUDA version issue (and made it work by not selecting any CUDA device)... I don't know if this makes sense, please tell me if it doesn't/what problem you think there is.
I realized that the best way was to delete llama.cpp (installed with curl install script) and install a clean CUDA 12 Ubuntu version.
\\- \*\*How do I properly uninstall llama.cpp (I don't wanna mess with Ollama files)?\*\*
\\- \*\*How do I install new version from .tar.gz archive without messing with system packages?\*\*
\\-Does my error diagnosis make sense to you? Would CUDA make generation actually faster? (I am getting 8tk/s with Qwen35B Q2 on 4060Ti 8GB due to no CUDA selected)
Edit: You guys saved me! Thanks! I had to install NVIDIA toolkit and switch to CUDA 13 llama.cpp tarball build
TL;DR: I'm a solo dev who wanted a simple, private way to have local LLMs watch my screen and do simple logging/notifying. After a year of building, I released v3.0.0 and my mom was able to use it for the first time and I wanted to say thank you!
Hey r/LocalLLaMA,
It is designed to monitor anything, some use cases:
It's a micro-agent framework controlled by an MCP (Agent which I call Observer). So you type in Observer what you want monitored, and it'll control the framework to monitor it.
The desktop app uses llama.cpp as an inference engine, the webapp uses transformers.js, and they both support your v1/chat/completions endpoints :DD
You can try it out in your browser with zero setup!... running gemma-4-e2b ONNX in the browser, crazy stuff! Thanks to Xenova/HuggingFace for transformers.js c:
You guys told me that the framework was cool, but it was very manual to setup agents/workflows. So I've spent the last year slowly making it more accessible so anyone from any technical background can use it.
Every couple of months I ask my mom to use the App. And for the first time she actually was able to setup a monitoring agent with a local LLM! Which makes me think the app is ready for general public adoption (wuuuu!).
I hope this makes local LLMs useful for everyone! Tutorial/Demo Which is the whole point of the project.
The core Observer AI platform is, and will always be, free and open-source. That's non-negotiable. The code is all on GitHub for you to use, fork, and inspect.
The line in the sand which I have is "if it's free for me, it should be free for the user", that won't change ever.
This project wouldn't exist without the inspiration I've drawn from this community. You are the people I'm building this for.
I'll be hanging out here all day to answer any and all questions. Thank you again for everything!
Cheers,
Roy
I ran Swift-1.5-Qwen3.8-27b-oQ8e-mtp through the llm-bench.io a few times today: oMLX on M5 Max 64 GB, thinking on at xhigh, 262k context window.
The big difference between Swift 1.5 and the base Qwen3.8 27B is how much it writes. Per full run (agent workflow, code generation, research, role play) Swift averages 51k generated tokens and Qwen3.8 averages 77k.
| Scenario|Swift 1.5|Qwen 3.8 27B|
|:-|:-|:-|
|Code generation|28.6k|40.1k|
|Research|11.1k|21.2k|
|Agent workflow|7.5k|11.9k|
|Role play|4.0k|3.8k|
The full benchmark run duration: avg. 24 min for Swift 1.5, avg. 38 min for base Qwen 3.8 27B
Everything else is about equal:
Still 3/4 of what Swift generates is reasoning, it just does less than Qwen 3.8 27B. The output is still very usable. I'll for sure give it a try to be my daily driver for a few day.
Runs Swift 1.5:
Runs Qwen 3.8 27B:
Hi, so I was working on building a local server for an older online game client a few years ago, and the amount of data I had to synthesize was intense. I ended up shelving the project. I had managed to build a basic login server, sorted out some crypt stuff, but it was just way too slow of progress for me. I've got some decrypted packets and lots of data to work with, so the LLM is not going to be forced to do this entirely blind.
I restarted it recently with Fable, and as expected, it has been a huge help. However, I'm consistently hitting the safety guardrails now that I am further into the project. I am wondering what you all would suggest for a local setup? I have 16GB of VRAM (RTX 4060 Ti) and 128GB of DDR4. If that is entirely insufficient, I might be willing to pay for hosting a more powerful local model. I'm fine with it being slow and chugging along all day and night - I am primarily concerned with it actually figuring out the client functions, and doing things as accurately as possible.
I appreciate any insight into specific models, harnesses, and other stuff I should be looking into. Thanks!
Would increasing the number of ram channels (eg dual channel to quad channel) help improve inference speeds with freetoken in hybrid gpu/ram setups or not? Or is increasing the vram capacity and vram speed the only optjon
So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17
Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV
I’m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:
Question → many sequential tool calls → final answer/document
Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.
I’ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. I’m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.
The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.
Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?
I understand that small active parameter counts don’t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also don’t assume that more total parameters automatically means a better model. I’m interested in where that tradeoff works in practice.
There are three things I’m trying to pin down:
For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.
I’m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).
I’m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like “model + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,” along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.
Has moving from a \~30B model to a 50B+ low-active-parameter MoE actually improved your agent’s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading
I read textbooks on my PC and ideally i want a program with a normal chat environment where i can just hammer in questions about what's currently on my screen. Example: I'm working on a PDF, underline or circle things... and then just type "what does this sentence mean?", you get it.
I do NOT want to manually screenshot, navigate to the folder, drag the picture into the environment and then also have to type the question. It should also naturally be aware that the conversation is about what's on the screen, so i don't have to steer it with "make a screenshot; use your vision capabilites" etc.
I tried multiple different MCP in LM Studio, none of them were great...or worked :/
Someone said AnythingLLM has this function but i couldn't find it.
Is there a good solution?
Thank you in advance! :)
What (and if) would you buy if you had $5k ... $10k ... $20k to spend on local AI?
So, I'm a contractor and a solo dev working on some products/saas/apps. Basically I usually run up to three cline/opencode sessions in the same time, long running software engineering tasks, often run out of 128k context, so 256k ctx is prefferable. Pretty much every day for multiple hours so I could burn quite a lot of $ daily on OpenRouter. Fortunately, since Qwen3.8 i rarely need to delegate to larger models like Kimi K3 or MiniMax M3.
Apart from code I often work on confidential documents so a local AI or an approved remote AI is a must.
My current AI setup is:
\- dev server with RTX3090 running Qwen3.8 UD Q4\_K\_XL with 128k ctx q8
\- lab server with 2x RTX3090 + 128gb ram running Qwen3.8 Flash Next with 256k ctx
Both machines are fine to run up to three sessions, one on dev and 2 concurrent 128k ctx tasks on the lab server. The performance i get from the 2x RTX3090 with Qwen3.8 Flash Next is close to a single DGX Spark GB10 (according to numbers).
Soon I may need to run more agents and work with other people so started thinking of an upgrade.
Does it make sense to invest in either a single GB10 machine or two and cluster them to get more space for more context and therefore more concurrent sessions? Would you consider other options?
Any feedback appreciated.
Another day, another refusal. Apparently asking opus "what are your thoughts on this?" is an attempt at a 'distillation attack'
Our precinct is hugging face, we work at breakneck speed, we're up against art thieves, code thieves, extortionists, we're on call around the clock. The people of LocalLLaMA -- our finetunes is our job (and we write our own emdashes thank you)
WHERE ARE THE DATASETS PEOPLE. What happened to us? We used to throw pies at Dario Altman and now, what, we're penniless, downtrodden, what happened?
Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding.
https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash
Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding.
https://huggingface.co/ornith-ai/Ornith-1.5-397B-DFlash
Ornith-1.5-35B-A3B-DFlash pairs the Ornith-1.5-35B-A3B model with a DFlash draft model for speculative decoding.
#
https://huggingface.co/emperorofrome/Gmcoder
Beats Ornith 1.5 and Oxcoder on HumanEval+ Mini — and does it without the overthinking. It gets to the answer using 40–68% fewer tokens. Built as a finetuned merge.
Edit: Best in the world is just hype. The coder is comparable to Ornith 1.5 but more token efficient by about 50% on average up to 78% at times and can be faster.
I'm building a new rig to host the mission control application that should sit on its own node and I immediatly thought of a cheap single socket ddr4 solution,
does anyone have some suggestions which bundle is cheapest or gives best value for price...?
\-no need for gpus
\-no need for much ram (8gb ram sticks)
\-many threads and max core speed would be important (maybe if it doesnt spike the price)
As usauly i was scrolling around YouTubes while .... well when every man scrolls YouTube to pass the time.... anyway, came across this video - https://www.youtube.com/watch?v=doR2RhsneRA
TLDR: Guy spends $2175 USD and over 36 hours using Claude Opus 5.5 to build a pretty impressive MMORPG.
Now there is alot of caviats, is it perfect... No... is it really an MMO ... no i havent seen any code added for it.. buttt the strcuture is there, its a hell of a start.
And it got me thinking, if people with their local AI's where to do something like this What models or infra would you use to do this.
Me I think you could get a really good start with Hermes, Qwen Flash, GLM5.3 across a couple of DGX Sparks or if you had 2.5 TB's of RAM and freetoken Kimi K3.
And with the detailed level of prompting and design he laid out prior to actaully letting claude have at it, i think it would doable locally.
So to start the Discussion, what would your tech stack be to do this locally? For me:
\- 2 x DGX Sparks - GLM 5.3
\- 5090 desktop with ComfyUI and a bunch of work floors for image generation, Guassplating, Image to 3D models
and to me I think that would be all that would be needed to get started, but im intrested to see what others think up
So I put this computer together after I saw a lot of others posting what they have, and I wanted to know where this ranked and what (in your opinion) are the best local models for coding, and for always on agents.
I didn’t want to type out the whole thing, so I asked my model to give me the specs.
Creature workstation (kernel 7.0.0-28-generic) with an Intel Xeon E5-2698 v4 at 2.2 GHz (20 cores / 40 threads, 50 MiB L3), 125 GiB DDR4-2400 (90 GiB free right now), one NVIDIA RTX PRO 4500 Blackwell with 32 GB of GDDR7 / 31.9 GiB usable VRAM (896 GB/s), and 5.37 TB of raw storage — a 915 GB NVMe root (379 GB free), a 916 GB media disk, and two 1.8 TB drives.
All figures read live from lscpu, free -g, nvidia-smi and df just now.
A "token" is not a fixed chunk of text. It's a word from the model's own private dictionary. Each model ships with its own vocabulary: the list of string-slices learned during training. Analogy: two people transcribe the same sentence: one writes "New York" as one word, the other as two. Both are correct; they're just counting different things.
So tokens/sec is speed measured in steps per minute, and two models can have different stride lengths. One can takes long steps (eg, 4.83 chars each), the other short ones (eg, 3.27). A child and an adult both walking "60 steps per minute" are not walking side by side.
The rules that follow:
TL;DR: vLLM recipe for Qwen3.8 Flash on one DGX Spark / GB10. 74 tok/s peak single-stream, 60 to 70 on normal requests, 212 tok/s across 8 streams, 2x to 3.4x faster cold prefill than the recipe it's forked from, full 262K context, and quality matches the original within noise. Everything is open, including the raw per-round data and the benchmark scripts.
https://github.com/dime-online/qwen3.8-Flash-DGX-UltraFast
Most single-Spark setups I've seen posted for this model land somewhere in the 35 to 45 tok/s range, so I spent a few weeks figuring out where the time per token actually goes on a GB10 and cutting it down. If you don't have a Spark, the tricks in the middle section should still be interesting, since most of them apply to any MTP or speculative decoding setup.
What this is, in plain terms
It's a ready-made serving setup. You build the container, pull the public weights, and get an OpenAI-compatible server that answers a lot faster on one box. The speed comes from the model's own draft head guessing several tokens ahead while the full model checks all of them in one pass. Right guesses give you several tokens for the price of one step, and wrong ones get replaced by the full model's answer, so output quality doesn't change.
Decode
Peak decode speed, same workload at every point, best of 3 rounds:
|Streams|1|2|3|4|5|6|7|8|
|:-|:-|:-|:-|:-|:-|:-|:-|:-|
|tok/s|74.1|110.0|132.9|155.5|175.8|191.5|205.8|212.2|
Tokens per step stays between 3.65 and 3.94 from 1 all the way to 8 streams, so the speculation doesn't fall apart under batching. 8 is where it tops out because that's the configured max\_num\_seqs, and the gain from 7 to 8 was down to 3%.
Prefill
Cold prompt with nothing cached, three repeats each:
|Prompt|This recipe|Original recipe|Increase|
|:-|:-|:-|:-|
|16K tokens|4,016 tok/s|1,171 tok/s|\+243%|
|64K tokens|2,426 tok/s|1,071 tok/s|\+127%|
|128K tokens|2,213 tok/s|1,065 tok/s|\+108%|
With prefix caching on, a cached coding prompt starts replying in about 0.57 s, which is what makes agent loops feel fast.
What actually made the difference
The model's own MTP head, run densely, lands about 3.7 tokens per verify step. That's the single biggest lever.
I cut the draft head's vocab from 248K to 65K ids. On a GB10 the draft pass is memory-bound, and reading a full-vocab head every draft step was a real chunk of the step time. The target model still verifies against the full vocab, so this can only change speed, not output.
The quant is W4A16 AutoRound for the MoE experts, FP8 for the side layers and INT8 for the lm\_head. No 3-bit and no NVFP4, because I wanted the speed to come from the serving path and not from squeezing the weights harder.
There's a GB10-tuned low-latency GEMM for the small decode-time matmuls and a sort-free top-k in the verify step.
The prefill gain mostly comes from a faster gather path for the per-layer embedding table, which removes a pile of serial page faults during prefill.
Put together, each decode step went from 68.3 ms on the original recipe to 52.3 ms on an agent-shaped coding workload, about 1.3x more steps per second.
One thing that didn't pay off: doubling the prefill chunk to 16,384 tokens gave no prefill gain at all and ran the box low enough on memory that I rejected it.
Quality
93.1% and 93.3% on a fixed 492-question suite over two seeds, covering code with execution checks, math, knowledge, instruction following, tool calls and long-context needles. I also ran a teacher-forced check against the original checkpoint, and top-1 agreement moved by 0.06 points against a 0.15 point noise band I set before running it.
Practical stuff
The model takes about 71 GiB, the KV pool is 16 GB, and around 16 GiB stays free under load. While generating, the GPU draws about 35 to 37 W median and peaks near 70 W on long cold prefills, with no power or thermal throttling across the soak runs.
The 65K draft vocab was built from English and code, so Chinese, Japanese and Korean output drafts less well and runs slower. Quality isn't affected, because the full model still checks every token.
Built on Saren-Arterius's qwen3.8-Flash-DGX-AutoRound, so big credit there. Happy to answer questions.
...the one that will mass produce a cheap device (sub $1000) able to run current sota models at a decent speed.
for that to happen, obviously RAM has to be cheaper, models need to be more efficient and CPUs have to change. It will take time. But as in the 70s/80s computers were huge and expensive mainframes only big companies had, today we are in the same situation.
Fortunately progress happens faster now, so it won't take 30 years to get affordable "home computers". Probably 10, hopefully less.
It's encouraging that today to run the latest qwen 27B you can spend less than $3000. But still...
I bought a B70 off Amazon last week for 1300 and have been playing around with it. Today I checked and it seems like I cannot find a single one online for less than 1600 and most places dont have them in stock anymore? What happened? What drives these sudden price increases? It seems like all GPUs across the board have seen another dramatic price flux. Some 5090s I saw were going for 10k!?
No click bait baby I promise - I'm live streaming the training process kinda like MiMo.
UPDATE: \[Training is paused for an hour or two\] back to training in batches. u/FullOf_Bad_Ideas has pointed out to me I'm burning a ton of compute for nothing on sequence lengths - we'll be breaking the run up into 7/8 batches and then going again.
Original thread: https://www.reddit.com/r/LocalLLaMA/comments/1wpg4a8/im\_trying\_to\_posttrain/
If you didn't see my original post a few days ago, I'm attempting a slightly more ambitious than usual project in trying to create at least a rough AliceAI-Foundation-80B-A3B-Instruct
I spent the weekend distilling my initial instruct training dataset out of Qwen 3.8 27b, medium thinking - intentionally done because I can run it locally, and I wanted a full dataset in some reasonable amount of time. Still took my v100s running 4x instances at 25tps, like 96 hours of non stop generation to complete the dataset.
I opted not to go for a pre-existing public dataset because I wanted to practice building my own distillation engine (which was configured to work off of an OpenAI compatible endpoint, so it'll distill anything you can hook it up to). The final dataset (this time) consists of 3340 samples: 1760 of general instruct transcripts, and 1580 agentic specific work rows about SWE, harnesses, terminals, etc - I gave the distilling engine a python sandbox and got to simulate turn driven development with a user, and I trained for a bunch of different harness syntax for tool calls, which hopefully will be enough to generalize - gonna run 2 epochs at first.
My GPUs are sobbing right now - turned them on on Friday and left for a weekend vacation, got back today, waited an hour for the data to finish generating, and then immediately fired up the train.
The stuff above is the short version. I'm guessing the initial SFT train will take about 3-6 days, and I plan on working on a RL implementation after I'm satisfied that the SFT has at least worked properly. I am training a rank 16 QLoRa adapter on only q/k/v/o proj, no direct knowledge weight fine tuning.
I thought what MiMo did with their recent training was really cool to watch online, and I like sharing my work with like minded people, and frankly, there's a part of me that's hoping someone will see this and want to hire me (looking for NYC work if you know anyone looking for some passionate ML engineers!) - so I've set up my own little training stream on a cloud flare tunnel.
The stream has the live in progress status of the train, including a live view of the actual data being processed by the model. It also includes way more detail about how I actually designed and generated my training data. Happy to throw the full set on HF as well. I don't expect this model to beat any existing standards but I'll be curious to see if I can get it to operate properly in a harness so I can formally bench it.
I hope you find this interesting! The live stream is a self updating website where you can see exactly what's happening - no need to reload. To watch the training live, visit https://figure-bios-expect-cio.trycloudflare.com/ \[i am currently fixing training issues but it'll be back asap\] -- I'll be keeping it up until the initial SFT is done, at least. The stream lets you inspect the training live as well. This is just a cloudflare tunnel to the trainer.
3 hour update? Loss started at 9ish and is bouncing near 3/4
Update today: back online
The human brain is neither an LLM nor a JEV. In a crowded market you hear a lot of speech and answer almost none of it. The ongoing question is not “what should I say?” It is “you talking to me?" and "should I say anything at all?”
This live voice demo showcases three hardware tiers—the Blackwell 4500, Jetson Orin NX 16GB, and Jetson Orin Nano 8GB—solving this exact problem. By splitting the workload between a lightweight decision head for turn-taking and a Gemma 4 pipeline for text generation, the setup delivers highly responsive, low-latency vocal interaction.
EDIT: repo (the latest code yet published) https://github.com/cortexist/little-gemma
I run an Unraid server and use local models for STT, memory, a small llm (I like the new swift bonsai), and SystemOne Models. I currently have a 5070 plugged into my x570 board (with a 5600x), and then the 5060ti on a riser. The 5060 runs at x4, there is a limitation on the board setup.
Running a model big enough for both cards runs SLOW, since the connection to the 5060ti is capped.
Ive tried optimizing with ninfer and vllm, but my speed is capped at the hardware level.
I am considering scrapping the 2 card setup for a single card, and right now the best budget option is the RX7900XTX.
Can you guys argue about what would be the better setup? I dont need the newest models, and I dont need the most speed. I pay for online models. I just like to have a bit of local stuff to help with server stuff. I feel like I spend too much time managing the 2 card setup for not enough to get out of it, and I think dropping to the one card would make things easier without a loss in speed. The idea is to sell both NVIDIA cards, and just get a single card with big enough memory that isnt a pig.
Any advice would help.
If someone is interested https://preview.redd.it/wlqoavzq7ash1.png?width=1524&format=png&auto=…
On the official JevBench evaluation battery, Gevva0 scored 74.63 (#1 global rank), averaging 214ms p50 across large legal contract sets with 82.9% accuracy on the forensic hard tier.
How it works under the hood:
The repo includes the evaluation harness, raw benchmark datasets, and a local web dashboard: https://github.com/solvingSteve/Gevva0
Setup instructions and benchmarks are in the README.
Working on Demos and Use Cases now so if you have any ideas I'll try to build them next!
Hi. So....... after toturing my 750 ti 4 gb and 16 gb ram with image gen model .i want continue experiment with agent mode like hermes or opencode, using local model but before go i want ask few question. are big model = good perfomance or small model can do same stuff. what small model recommend for agent my spec what caveat of doing this ?
Qwen Flash Next speedup for Vulkan people
Seems like it would be a very nice way to get 24gb of VRAM, and it should be possible to do right? Given that 20gb 3080s exist. Just wondering
According to this post it is possible to reach very high numbers using mtp, however I have failed to reproduce the 50+ tps for 5060ti.
Am I just ignorant or how exactly do this? Or is this a fake post?
Freshly built llama fork for sm120 (Blackwell):
https://github.com/Anbeeld/beellama.cpp
"C:\\llama\\build\\bin\\llama-server.exe" \^
\-m "%MODEL%" \^
\--port 8090 --host 127.0.0.1 \^
\-ngl 99 \^
\--cache-type-k kvarn3 --cache-type-v kvarn3 \^
\--flash-attn on \^
\--load-mode mlock \^
\--jinja \^
\-c 98304 --parallel 1 \^
\--fit off \^
\--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-ubatch-size 128 \^
\-ctkd q8\_0 -ctvd q4\_0 \^
\--kv-tail-tokens auto
Runs at 20-35 t/s.
I'm setting up my first build with Open Claw. I'm new to ask of this and starting from zero knowledge.
I've done a lot of reading up, but it's a little overwhelming. So I'm biting things off in chunks.
I want everything to be local. I have a mini PC with 96GB of RAM so can fit a good sized model.
I have Open Claw set up right now with OpenRouter.
My next step is to set up my local model.
What is the current leading open source model for generic tasks and learning? I have plenty of memory to support.
Kimi K3, Opus, Quen 3.8, other?
Would you buy an AI appliance that made it easier for you to run local AI models such as Qwen 3.8 27B and Gemma 3 27B?
By easier I mean, the appliance will take care of all the infrastructure stuff for you such as installing the OS, the inference engine such as llama.cpp or vLLM, access management and perhaps include aome preconfigured agents for you start using the system.
The system is multi-user, with a role-based management system, meaning multiple people can use it, including teams.
All runs local, but with the option of using cloud based models if you need more horse power.
Is this something that has an appealing?
Or the fun is the tinkering 🤪
For me it's a pass/fail: can it fully replace a remote worker? How do you define it?
Hey guys i am planning to get a mac since the GPU and other stuff isnt currently possible for me rn so what models near or stronger than Sonnet 4.6 or somewhat nearer can I use it on mac or would i be able to use it actually?
I mainly need it for coding tasks, different file management and reports and also reasoning. SUggestions or feedbacks are welcomed for models. I want everything local for privacy concerns
Edit: Fixed the model mention
I am confused between LM Studio and Bionic.
I have LM Studio for a long time though not used frequently.
Recently I am trying to learn the agentic stuff. Watched a few videos but they were all using Bionic. The strange thing is the interface looks different than the one I just installed today. (Mine seems to be missing features, no developer mode. I am on MacOS)
On the other hand, I seem to be able to find those missing settings in LM Studio.
So what is the difference between the two? Is there anything Bionic can do but LM Studio doesn't?
Basically the idea is take your favorite model, for example qwen3.8-27b or say dsv4vision. Strip everything out that is not needed by that model so the only thing needed is just for the model. Optimize the remaining code to be fast. The idea is to have a model also do this, provide it with enough tools, prompts, docs, guidance. I reckon that if we have a llama.cpp that is optimized for just one model architecture without all the cruft needed to run and. handle other models, that it would not be surprising to easily see 2x+ performance improvement. Anyone thinking along this idea? Again, the goal will be to give this task to a smart model, let it run in a loop, and after a week or 2 you hopefully end up with llama.qwen3.8-27b or llama.glm5.3-flash that would fly.
Recently started using the MacOS menu bar app for llama.cpp available at https://llama.app When you use the UI to download a model, it seems to do an Ollama-esque hashing instead of just saving the GGUF. Even if you put a GGUF in the Model directory folder, neither the menu bar UI nor the web UI recognizes it. https://preview.redd.it/bsiw7kbn46sh1.png?width=1186&format=png&auto=…
Hello, I created a calibration free quant method called TQ. It doesn't need any data so models could be quantized as soon as they come out and the quantized models would generalize better. It performs closely to calibration-based methods. For example, for Qwen 3.8 27B the 4-bit TQ has mean KLD of 0.0282 and top-1 of 92.4% I opened sourced the quant code as Quant Factory so anyone can quantize and run open source models. Right now it only supports 4-bit but I plan to add support for 2-bit and 3-bit soon. The repo link is: https://github.com/textclf-api/quant-factory I have a collection of quantized models using TQ at: https://huggingface.co/textclf You can run these models using either using the following docker image or by following the repo's instruction. For example you can run textclf/Qwen3.8-27B-TQ-4bit like this: docker run --rm --gpus all -p 8000:8000 docker.io/textclf/tq-quant:4bit-main vllm serve textclf/Qwen3.8-27B-TQ-4bit --quantization tq The Dockerfile in the repo shows how this docker image was created. The repo's README explains the approach used for this quant method and why it is useful. You can try it and let me know what you think. Feedback appreciated. EDIT: I did KLD testing using the Wikitext-2 dataset for Qwen 3.8 37B. I also did the same test for the Unsloth-UD-Q4\_K\_XL quant. Here is what I go: |Quant|Disk Size without MTP (GB)|Mean KLD|Median KLD|99% KLD|Top 1% Agreement| |:-|:-|:-|:-|:-|:-| |TQ 4-bit|17.76|0.02823666|0.01282929|0.26897613|92.419%| |UD-Q4\_K\_XL|17.59|0.00771805|0.00318557|0.07390548|95.779%| Working on getting more tests for other models!
I had a theory that ram speed didn't really matter and the only thing that matters is your GPU capacity, vram speed, vram size, and system ram size instead of ram speed or cpu speed. I think I've been vindicated. I used unsloth/Qwen3.8-27B-GGUF:UD-IQ3\_XXS (10.9gb) https://huggingface.co/unsloth/Qwen3.8-27B-GGUF as my baseline and mtp q4\_0 1.37gb only on the dual gpu systems because mtp was too slow on the 9060xt and 5070 I used llama.cpp for all of them and tried to go for the max context I could since they are supposed to be day to day agents. I picked prism32 as my agent harness for this because it's the most compatible, fastest, and reliable one I've found. It's a custom architecture https://github.com/MegaDyneSystems/prism32 which gives it some massive advantages, especially if you want to avoid bloated frameworks eating your context or CPU. The 2007 system literally doesn't work with any other harness because they all require sse4.2 and a ton of heavy dependencies. prism32's only dependency is python 3.7 or above. It's so lightweight (uses around 5 to 10mb ram) that I actually run it bare metal on my ARM synology NAS and even my 2008 TP-link router. If you want to run any agents especially with advanced features on edge or legacy hardware without it choking your system their is no competition I tried these 5 systems: 2007 dell precision t5400 ddr2: 24gb ddr2 8 core 8 thread dual xeon x5460, rx 6700xt rx 6700 22gb vram total, 215gb ssd total system memory 46gb (pic 1) 2009 dell precision t5500 ddr3: 72gb ddr3 12 core 24 thread dual xeon x5675, rtx 5060 rtx 4060 16gb vram total, 512gb ssd, total system memory 88gb (pic 2) 2012 dell precision t3600 ddr3: 64gb ddr3 6 core 12 thread xeon e5-1650, rtx 4060, rtx 3060 12gb total vram 20gb, 512gb ssd, total system memory 84gb 2018 hp obelisk desktop 875 ddr4: 32gb ddr4 3200 i7 8700, rx 9060 xt 16gb, total vram 16gb 512gb ssd, total system memory 48gb (pic 3 ) * 2025 HP omen: 32gb ddr5 6000 RTX 5070 12gb GDDR7 vram 1tb ssd, total 44gb memory (pic 4) The speed results on each system were: 2007 dell precision ddr2 144k context, 18 tk's a second decode, 150tk's prompt processing 2009 dell precision t5500 ddr3 256k context 22 tokens a second decode, 224 tokens prompt processing 2012 dell precision t3600 ddr3 104k context 22 tokens a second short context 14 tokens a second long context, prompt processing is 315 tokens, (could optimize further but tests took me long enough) 2018 hp obelisk 875 ddr4 180k context, 16 tokens a second long decode, 477 promp processing 2025 HP omen 131k 13 tokens a second, 43 prompt processing (yes 43) conclusion The older dual xeon dual gpu setups completely destroyed the newer stuff in context length and speed even on worse GPU's which I think proves my theory, The HP omen system is at least $1500 and the 2007 system didn't cost more than 400 total, rx 6700 xt was $190, rx 6700 was $140, on ebay they are overpriced at $150 although I got it $50 second hand in 2014, and the DDR3 systems were $20 second hand and go around 80 to 150 on ebay, I'd say the ddr3 systems are the best for performance and value, next project is running qwen 3.8 flash next on the ddr3 systems
I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "*Let me look at the problem from the grader's perspective*" and referred to "*hidden tests*", "*test authors*", and "*the checker*".
I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.
\[Pictured example shows verbatim quotes from agent's reasoning\] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.
My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors:
https://joinhandshake.com/research/ai/deepswe-reward-hacking/
Not an LLM, but the ternary findings should carry over, and we hadn't seen Sherry-style 3:4 weights run in a browser before. Disclosure: this is our work at Precisit, everything is MIT.
What it is
|Model|File size|vs depth-4 bot|vs depth-6 bot|
|:-|:-|:-|:-|
|dense (fp32)|29.7 MB|0.92|0.89|
|T34, trained ternary|1.59 MB|0.93|0.91|
|T34, fine-tuned from dense|1.59 MB|0.89|0.9|
|T34, converted after training|1.59 MB|0.13|0.11|
|Base243 (TQ1\_0 style), trained|1.93 MB|0.89|0.88|
200 games each, both sides play a random move 5% of the time, a win counts 1 and a draw ½.
What we learned
Play it:
https://precisit.github.io/onepass-web/demo/c4-size/
Code, models, every result:
https://github.com/precisit/onepass-webgpu-ternary
The write-up:
https://precisit.com/en/blog/onepass-c4-size/
Has anyone gotten post-training 3:4 conversion to work on models, or does it need training?
My ninfer-ext fork of the famous ninfer now supports exl3 , and have released two models as well. https://huggingface.co/jabbatheduck/ninfer-ext-models
What FOSS/self-hosted tools or AI tools have you found genuinely helpful for students? Also interested in anything non-AI that has helped with studying, notes, organization, research, etc. Curious what you guys actually use or found useful. (written fully by human)
I'm curious, I'm a little bored of Qwen 3.8 and how slow it is even with MTP on Edit: Using Strata with my setup 55t/s, perfectly functional!
*Holo4*\-27B-GGUF is a vision-language model (VLM) for Computer Use, built on the Qwen3.8 dense architecture and developed by H Company. Used with the hai-agents harness, it can send screenshots and tool results to the model, then execute its requested clicks, typing, code, and tool calls.
https://huggingface.co/Hcompany/Holo4-27B-GGUF
*Holo4*\-35B-A3B-GGUF is a vision-language model (VLM) for Computer Use, built on the Qwen3.5 mixture-of-experts architecture and developed by H Company. Used with the hai-agents harness, it can send screenshots and tool results to the model, then execute its requested clicks, typing, code, and tool calls.
https://huggingface.co/Hcompany/Holo4-35B-A3B-GGUF
https://preview.redd.it/rqx7l4qqs8sh1.png?width=1656&format=png&auto=…
https://preview.redd.it/rzez39trs8sh1.png?width=1656&format=png&auto=…
Trying to see if I can use a local model to generate game assets, or even 3D printer models. Any suggestions? Something that would fit in 64GB of ram, speed is not an issue. Just need it to work well.
TLDR: security researcher eddie zhang used a modified+uncensored local qwen 3.8 27b to create an executable capable of dumping LSASS memory for credential harvesting while evading 2 modern EDR security products.
this makes me reflect on how cloud providers keep putting guardrails on everything to the point where even authorized testing gets blocked. local models are the only real option if we want total control, but running heavy local rigs for long agent tasks drains so much compute and management overhead.
been using claude code hooked to sumus to handle my local project workflows and orchestrate tasks in the background, while keeping full local file access on my machine. curious if anyone here is running local uncensored models as local agents for heavy automation, or if you still blend cloud models with local execution setups for your dev environment?
Here's the project. I have nothing to do with it. I'm just an amazed user.
https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BEN…
Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat.
"[6204 chunks in 119.0 s | encode: 1239 tok/s | decode: 57 tok/s]"
That's with MTP on. The PP speed in particular is just so fast. That PP speed is twice the speed of the fastest Strix Halo specific fork of llama.cpp I've ever used. Needless to say, the uplift is even greater compared to mainline llama.cpp.
It works with models other than QFN, but the current number is small. You can find the list on their project page.
I'm running Qwen 3.8 27B Q4 XS with \~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.
Build llama.cpp with Vulkan:
cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j
Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then:
llama-server \
--model Qwen3.8-27B-UD-IQ4_XS.gguf \
--mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \
--n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \
--batch-size 2048 --ubatch-size 512 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \
--load-mode none --fit off \
--cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \
--no-context-shift --jinja --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080
---
Edit: The process is documented here: https://zenodo.org/records/23088880. Feedback will be integrated into the upcoming Qwen 4 setup.