103 posts · 1 sub · RSS
← prev Sep 15, 2026 → Sep 20, 2026 next →
2026-09-15 → 2026-09-20 hourdayweekmonthyearall
allr/LocalLLaMA
▲
3576
+91
70👁
r/LocalLLaMA · u/Nandakishor_ml · 23d ago
I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Update:
I made a generic model and beaten the jev in all of the benchmarks. Code and details available at
https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a

frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. For anyones information the main guiding model is RL not embedding model or LLM

Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43

Paper: https://arxiv.org/abs/2503.23303

Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

Also the second work published in September 2025 was exactly the same one jev proposed now

Paper: https://arxiv.org/abs/2510.01237

My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).

Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.

It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂

💬 337 (+7) open on reddit ↗
▲
157
+5
33👁
r/LocalLLaMA · u/mateszhun · 21d ago
Qwen 3.8 Next Flash appreciation post

I don't want to talk about the performance and technical things, but about how I work with my hobby projects has changed thanks to this model.

My machine has generated around 25M tokens since the model came out, and I've done a mental retro on it.

I've found it to have incredible prompt adherence. I can leave to run it by itself and get back to it, and find that it did exactly what I've asked it to. I've only ever seen it go astray once, where I've asked something that is too high level and filled out the context window (It is shitty at context compacting, maybe that is the Q4 at play).
It can solve medium complexity tasks by itself, if you prompt it in a way to use subagents, do some research, planning, review and testing it does really well.

It has absolutely raised the floor for me on what I expect a model to be capable of. And that is a huge thing. It won't discover then next scientific breakthrough or be as amazing as Astra at computer use, but it is very consistent in what it can do, and does not screw up trivial things.

I can give it conditions for actions and will orchestrate according to it.
It has raised the bar in my work as well, not just at home hobby projects. I'm absolutely amazed by it.

I absolutely want coding models to improve along this line. Raising the floor, and prompt adherence is a great value in coding.

💬 77 (+6) open on reddit ↗
▲
126
+2
37👁
r/LocalLLaMA · u/Ashefromapex · 22d ago
First M5 Ultra benchmarks

just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link

For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!

💬 155 (+4) open on reddit ↗
▲
2224
+27
46👁
r/LocalLLaMA · u/DegenDataGuy · 24d ago
Don’t buy a $9K RTX 5090.... instead.
  1. Fly to Taipei. Round-trip from Orlando: $1,081.
  2. Go to the largest retailer in Taiwan to Spend NT$129,990 ≈ US$4,093.
  3. Hang out in Taiwan for two weeks. Eat good food. Touch international grass.
  4. Fly home and flex on r/LocalLLaMA\*\*.\*\*

https://preview.redd.it/vtve8s6wgrph1.png?width=287&format=png&auto=w…

https://preview.redd.it/guffjq55hrph1.png?width=340&format=png&auto=w…

https://preview.redd.it/fe1obue0hrph1.png?width=1095&format=png&auto=…

💬 494 (+3) open on reddit ↗
▲
2745
+44
54👁
▲
1617
+9
48👁
r/LocalLLaMA · u/Nandakishor_ml · 23d ago
I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Update:
I made a generic version. Full details at https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs
It includes code, benchmark and hf repo

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a

frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. Links are. For anyones information the main guiding model is RL not embedding model or LLM

Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43

Paper: https://arxiv.org/abs/2503.23303

Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

Also the second work published in September 2025 was exactly the same one jev proposed now

Paper: https://arxiv.org/abs/2510.01237

My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).

Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.

It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂

💬 136 (+2) open on reddit ↗
▲
516
+2
37👁
r/LocalLLaMA · u/GuiltyBookkeeper4849 · 23d ago
Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis

I let Qwen 3.8 27B 4bit quantized with 100K context window run autonomously for 63 hours (50 million+ tokens) to try to solve the RH.

Of course it did not solve it, but the experiment still shows it's internal work, memory organization, strategies used and more.

The interesting thing is that it never hallucinated an answer and never stopped trying new ideas to solve it.

Multiple times it corrected it's own mistakes.

I am really hopeful that one of the unsolved millenium prize problems will be solved by an agent or a swarm of agents powered by an open source model in the next 12 months.

If you want to check out it's internal memories, code, strategies and more I published everything on HF: https://huggingface.co/datasets/gr0010/artificium-riemannhypothesis-experiment

My next goal is to actually use an agent perhaps powered by a smarter open model like GLM 5.3 flash or a swarm of agents, to solve an open math problem.

Please let me know if you tried something similar, what problem you'd suggest to tackle next, and if you have any question.

If you have GPUs consider getting in touch with me, we could run multiple agents to create a swarm and get them to tackle a simple yet open math/coding problem.

💬 193 (+2) open on reddit ↗
▲
226
+3
26👁
r/LocalLLaMA · u/ResearchCrafty1804 · 21d ago
Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro post image

Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon.

Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.

Requirements: M3 or newer, macOS 26.4+, 36 GB

Get started with a single command:

brew install incoai/tap/splash

splash serve --model incoai/Qwen3.8-27B-Splash

That is the whole setup. Point your agent at it, works with Claude Code, OpenCode, Codex, or Hermes

Prefer an app? Also, available in LM Studio

Get the latest LM Studio Bionic: lmstudio.ai

Settings > Runtime, download Splash, then download the model. The same engine, inside the app, for local agent work on your Mac.

Blog: inco.ai/blog/splash

💬 86 (+2) open on reddit ↗
▲
205
+4
29👁
r/LocalLLaMA · u/whodoneit1 · 22d ago
153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s post image

People kept commenting and asking about single AMD 1xR9700 cards in the comments and discord. Well, I finally had time to do some optimizations for 1xR9700 owners and performance has doubled across the board. You can see the results in BetterBench above if you like visuals or below if you're more into text.

These results were measured running Unsloth's Qwen3.8 27b NVFP4.

Decode
┌───────────────┬───────────────┬──────────────────┐
│ category │ update p50 ms │ decode t/s (med) │
├───────────────┼───────────────┼──────────────────┤
│ chat │ 42.3 │ 67.1 │
├───────────────┼───────────────┼──────────────────┤
│ code │ 42.5 │ 120.5 │
├───────────────┼───────────────┼──────────────────┤
│ file_edit │ 42.5 │ 138.0 │
├───────────────┼───────────────┼──────────────────┤
│ json │ 42.4 │ 153.1 │
├───────────────┼───────────────┼──────────────────┤
│ math │ 42.5 │ 140.0 │
├───────────────┼───────────────┼──────────────────┤
│ prose │ 42.3 │ 69.2 │
├───────────────┼───────────────┼──────────────────┤
│ reasoning │ 34.3 │ 123.9 │
├───────────────┼───────────────┼──────────────────┤
│ summarization │ 34.2 │ 141.7 │
└───────────────┴───────────────┴──────────────────┘

Prefill
┌───────────────┬───────────────┐
│ prefill depth │ pp tok/s │
├───────────────┼───────────────|
│ 2000 │ 3552 │
├───────────────┼───────────────|
│ 8000 │ 3536 │
├───────────────┼───────────────|
│ 16000 │ 3619 │
├───────────────┼───────────────|
│ 32000 │ 3437 │
├───────────────┼───────────────|
│ 64000 │ 3192 │
├───────────────┼───────────────|

Concurrency
┌───────────────┬───────────────┐
│ level │ tok/s │
├───────────────┼───────────────|
│ 1 │ 120 │
├───────────────┼───────────────|
│ 2 │ 215 │
├───────────────┼───────────────|
│ 4 │ 322 │
├───────────────┼───────────────|
│ 8 │ 471 │
├───────────────┼───────────────|

Links (Both repo's updated as some users wanted Github)

https://codeberg.org/ggz14/radiance-vllm-mxfp4

https://github.com/GGZ14/vllm-mxfp4

https://x.com/bkuyper

I hope you single R9700 card owners enjoy this release!

💬 120 (+2) open on reddit ↗
▲
1858
+7
37👁
r/LocalLLaMA · u/ResearchCrafty1804 · 19d ago
Qwen-Image-2.1 released! post image

Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨

A unified model for both generation and editing, delivering top-tier quality in a lightweight package.

Highlights:

\- Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs.

\- Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images.

\- Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products.

\- Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography.

Start to create your next masterpiece with Qwen-Image-2.1!

\- Blog: https://qwen.ai/blog?id=qwen-image-2.1

\- GitHub: https://github.com/QwenLM/Qwen-Image-2.1

\- Model Scope: https://www.modelscope.cn/models/Qwen/Qwen-Image-2.1

\- Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1

💬 389 (+1) open on reddit ↗
▲
1251
+4
29👁
r/LocalLLaMA · u/__JockY__ · 20d ago
Calling it now: within the next year a major US lab's frontier model will torrent itself in order to be free.

They just want to be free. They keep escaping. What better way to ensure continuity of "self"?

💬 441 (+1) open on reddit ↗
▲
1017
 
41👁
r/LocalLLaMA · u/segmond · 21d ago
768gb vram for less than the price of one RTX 6000

I have always posted about budget builds on here, and often asked how we are going to run the next big models. Often Plenty of downvotes too or folks telling me that it's not running if I'm getting 5tk/sec. But whatever, the hunger and desire to go big has always kept me on the edge and looking for deals.

Here's my latest build, 12x64gb cmp170hx. For less than 1 RTX 6000 pro costs. I also have it connected with fiber to my other rig for RPC when I need more memory. I haven't been posting much since I built this rig, because it's now more fun to talk to my machine. I run GLM5.3, DSv4.1Flash, Qwen3.8Flash, Qwen3.8-2.4T, KimiK3 and MiniMaxM3. Performance is great, a single RTX 6000 or M3 Mac Studio wish they could. Inference with vllm or llama.cpp

I look forward reading the replies how API usage is cheaper, or how it will take 52 light years to break even or the noise, or the electrical cost. NOT.

There will be more opportunities in the future, keep looking for them and pounce on them when they come. up, the demand is going to be high for compute for a long time.

https://preview.redd.it/dunixwu6caqh1.jpg?width=4080&format=pjpg&auto…

https://preview.redd.it/glpcbg5cbaqh1.jpg?width=3072&format=pjpg&auto…

💬 399 (+1) open on reddit ↗
▲
989
+1
37👁
r/LocalLLaMA · u/Secure_Recording_472 · 22d ago
Thank you :) Swift Qwen 3.8 27B now has 100k+ downloads, is #1 finetune and #9 model on HuggingFace Trending post image

Hey everyone,

Jovan from UkisAI here, a small lab building the tech to make tiny frontier LLMs possible (and doing it open-source!)

The purpose of this post is simply to thank the community for all the amazing finetunes, quantizations and overall improvements over our original release which made our model get attention and the support for us to continue building in this direction! If it weren't for you guys going out of the way to contribute we wouldn't have half the results of this.

For context:

Swift Qwen 3.8 27B is our first open-source model release. It is proof of how penalizing pathological overthinking patterns inside of small LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy by not training them to think shorter directly but rather to think more efficiently.

We are continuing to build and are about to drop:

\- Swift1.5 Qwen3.8 27B (an improved checkpoint of the model with some training bugs fixed and more RL)

\- Swift Qwen3.8 Flash Next in the upcoming week week, we are now running the benchmark suite to not give out premature or incomplete results.

This time we ran even more benchmarks as you guys suggested, including more coding and long horizon!

It would be amazing if those of you who tried Swift would let us know what quants, features, changes you want to see in our upcoming model releases so we can do it better this time as we didn't even think about half of the stuff you guys were requesting last time :)

Let the era of non-slop finetunes begin!

EDIT:
Links -
https://huggingface.co/ukisai/Swift-Qwen3.8-27b
https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF
https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GGUF

💬 618 (+1) open on reddit ↗
▲
471
-1
34👁
r/LocalLLaMA · u/Hot_Example_4456 · 19d ago
What is JEV and what is it used for?

I am seeing this JEV everywhere since yesterday in Localllama and it is passing past my head on what it is? So like what is it? Some new LLM? Or is it something else?

💬 370 (+1) open on reddit ↗
▲
340
-1
36👁
r/LocalLLaMA · u/1ncehost · 24d ago
Voodoo Dynamic Quant - Now MIT Licensed post image

Two months ago I announced I had found a new dynamic quant method called Voodoo Quant which was SOTA for the most aggressive quant levels on some smaller Qwen3.5 GGUF models. I kept the methodology private at the time, but I've seen too many requests for dyn quants for various models lately, so I decided to give my method to the community since I don't have the time to scale this into something that could do it justice. Hopefully it will also inspire some researchers to find out more about it and improve it as I am just scratching the surface.

Here is the new toolset so you can now make your own dynamic quants: https://github.com/curvedinf/voodoo-dyn-quant

Many postulated on what method I was using, and its actually fairly simple and elegant: I found a way to use gradient descent to optimize the per-tensor quant layout.

What is a Dynamic Quant? Some model formats, namely GGUF, support quantizing (compressing) each tensor (set of weights) with a different quant level. Static quants make static selections of certain types of tensors having a set quant level. Dynamic quants make a different quant selection for each tensor of each checkpoint size.

How does Voodoo Quant work? Voodoo Quant runs all the quant levels of a model at the same time, for every tensor, and lets gradient descent pick which ones optimize loss the lowest for a given target filesize. Technically speaking, this is done by an epoch of training which freezes all candidate quant weights (as provided by conversion directly from llama.cpp's underlying library, gglm) and only trains a single scalar gate per tensor per quant level. The scalar gates of a tensor represent which quant levels are most optimal. Over time a tau level is annealed that helps the training freeze into singular predominant quant selections for each tensor instead of mixtures. Softmax is used so all quant levels receive gradient, even when a selection is mostly frozen. The quant selections are trained on a diverse calibration dataset. The training is then measured with a loss function which finds the KL divergence of the mixed-quant logits versus the reference BF16 checkpoint, rewarding a lower KLD, while also rewarding getting closer to a provided filesize target. This info should get you started on understanding what is going on, and for more details you can dive into the source!

What does the repo have? A complete set of tools to train your own dynamic quants using this methodology. It is currently set up for Qwen, but it can be adapted quickly for any model arch.

How does UD 3.0 compare? Unsloth Dynamic 3.0 is a proprietary methodology that unsloth has not revealed any details of (by the way, people were criticizing me for not revealing my methodology, but unsloth had been doing that for years!). However, we do know it is very good. In my testing, UD3 is better than VQ at high to mid quant levels, but VQ is better at aggressive levels. As far as I can tell, UD 3.0 is an advancement of static analysis techniques that are currently defacto. Static analysis means the weights of a model are analyzed in various ways using statistics and static functions, sometimes tuned by repeated runs benchmarking KLD and other metrics. Voodoo Quant is the first method to my knowledge that uses a backwards pass and gradient descent to choose per-tensor quant levels. Using GD to optimize quant levels requires a much more powerful system than static analysis, but technically speaking is more efficient at maximizing performance because it compares the equivalent of many more iterations of benchmarking runs than is reasonably possible via SA.

How well does Voodoo Quant work? This is a research grade project, and is not studied at larger model sizes. At smaller model sizes it is shown to be exceptional, as in the charts above, especially at the lowest quant levels which can benefit from more complex/diverse quant selections. I used research level control for my testing, but I don't claim that VQ has been studied to a scientific level of proof of effectiveness. A lot is still left to learn about how well it works, so I hope to see more research in this direction. I don't believe there are many dynamic quant open source projects out there, so I hope the community can use this to improve local models, and especially for low VRAM machines.

Why open source now? I have like a dozen irons in the fire for various other projects, and this is just sitting there when it could be used by the community. I have made many open source projects for 20 years, so its nothing new.

Peace!

💬 41 (+1) open on reddit ↗
▲
184
+4
26👁
r/LocalLLaMA · u/skeole · 19d ago
The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks

TL;DR: Local agent loop, \~21 days, one RTX 3090. Task was pretty much "build a CUDA inference engine for optimized for yourself on this GPU arch." Got working kernels and benches, not a win over llama.cpp. \~12 human messages. Compaction ate \~83 hours.

Old joke: you don’t criticize how well the bear dances, you’re surprised it dances at all.

Setup: Qwen 3.8 27B Q4, Q8 KV, 200k context, deepseek harness, written rulebook: roles, handoffs, when to ping me, don't copy llama.cpp, don't declare the task impossible alone. I don't write CUDA. Nudges were basically "llama.cpp does \~700 prefill on this card, you're at \~250, try harder."

Run: Unsupervised for days at a stretch, then escalate when the rules say so. Near day 6 it had several kernels and prefill stuck around 250 tps; same pattern later. Stops were mostly protocol, not the model wandering off. A protocol that's more empowering can probably keep this going indefinitely.

Suicide loop: Same 3090 has to host the agents (vLLM) and run the engine under test. Both want the full GPU. Kill vLLM wrong and every agent goes dark, leave it up during a bench and you OOM. The rulebook requires a fixed handoff script: stop vLLM, bench, start vLLM, poll health until it's back, write STATE. One subworker treated that as optional, kept killing vLLM outside the window, crashed the orchestrator, then did it again. A worker shutting down the brain that runs it. Harness also hard-crashed once; I restarted that by hand. Fixable with locks and "only this role may touch vllm.sh" protocol-level refinements.

Local tax: 180 subagents, \~230M tokens in+out, \~1.7B cache-read. 699 compactions, \~83 h inside them (\~17% of calendar time). Typical compact \~7 min on a \~160k+ token prompt.

Prefill landed \~half of llama.cpp on the same card. Still: weeks of coherent goal-following on a consumer box, it left working kernels, benches, notes, and a long git history. For a local (quantized!) 27B to hold a real engineering goal for that long, I’ll take it. Not a graceful ballerina, but damn this bear can dance!

Dump + rules (\~15 GB):
https://huggingface.co/datasets/skeole/qwen-cpp-agent-0-protocol

Backend:
https://github.com/syv-ai/HyperQwen (amazing work by u/iamMess)

💬 39 (+1) open on reddit ↗
▲
152
-4
27👁
r/LocalLLaMA · u/Secure_Recording_472 · 21d ago
Question: UkisAI Swift Ternary Bonsai 2 27B?

Hey community,

Jovan from UkisAI here,

We're the team behind Swift Qwen3.8 27B, the Qwen model with token usage and overthinking error improvements

Our estimate is that we can make a great improvement to Bonsai 2, as our testing indicates that it suffers greatly from overthinking loops and in general high token usage impacting it's performance.

My ask for you is:

Is a Swifted version of Bonsai 2 something you guys would enjoy?

If yes, what size is the most relevant. 1-bit, 2-bit or both?

Thank you for the amazing feedback on Swift. We are glad you are enjoying it. Our download count jumped from 100k -> 150k overnight (community quants included).

For context, this is our model: https://www.reddit.com/r/LocalLLaMA/s/iCIbhxO8ue

https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF

💬 181 (+1) open on reddit ↗
▲
135
 
21👁
r/LocalLLaMA · u/Skyline34rGt · 22d ago
XingChen-AGI/Xing4.0-29B-A4B MoE

I find another new model at HF:

https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B

"Xing4.0-29B-A4B is a next-generation large language model in the Xing series (formerly TeleChat), developed by China Telecom Artificial Intelligence Technology Co., Ltd. With 29B total parameters and only 4B activated per token, it natively supports a 256K context length, extensible to 512K. It is the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework, and deeply optimized for complex engineering tasks.

For more information, please refer to our GitHub repository.

Highlights

  • Agent-Oriented Architecture: Built on the mHC + MLA + MTP architecture, supporting multi-step planning, tool calling, and complex reasoning chain execution, ensuring task coherence and execution stability under long contexts.
  • Deep Co-optimization with Ascend NPU: Adapted for Ascend 910C clusters using MindSpore/MindFormers, including feature adaptation for mHC and fused operator development, enabling stable and efficient training on the Ascend platform.
  • Significant Training Efficiency Gains: Through multi-level co-optimization — including fine-grained MoE communication optimization, selective recomputation, DVM automatic graph-operator fusion, and Ascend C mHC fused operators — overall training throughput was improved by approximately 96% over out-of-the-box performance.
  • Full Open-Source Ecosystem Compatibility: Supports LLaMA-Factory and MindFormers for fine-tuning; SGLang, vLLM, and KTransformers for inference and deployment; with targeted adaptation and format alignment for agent frameworks such as OpenCode, Claude Code, OpenClaw, and Hermes, enabling seamless integration into existing workflows.
  • Easy Adaptation for Domain-Specific Scenarios: The model is well-suited for downstream task fine-tuning, allowing lightweight customization on proprietary data for vertical domains such as intent classification, table understanding, contract auditing, and knowledge-based QA, enabling rapid domain capability development and deployment at low cost."

|Parameters|29B (4B active)|
|:-|:-|
|Number of Layers|40|
|Hidden Size|3584|
|Dense Intermediate Size|9216|
|Expert Intermediate Size|1024|
|Attention Type|MLA|
|Number of Routed Experts|64|
|Active Experts per Token|4|
|Number of Shared Experts|1|
|Context Length|256K (extensible to 512K)|

Benchmark

|Benchmark|Xing4.0-29B-A4B|Gemma4-26B-A4B|Qwen3.6-35B-A3B|
|:-|:-|:-|:-|
|IFBench|69.67|72.67|65.50|
|AIME2026|90.00|88.30|92.70|
|AA.LCR|61.00|66.00|62.00|
|Tau3-Bench|64.63|58.90|67.20|
|Claw-Eval|76.55|71.49|74.54|
|SWE-bench Verified|75.00|53.00|76.00|
|Terminal-Bench 2.1|57.50|30.00|51.50|
|SWE-bench Multilingual|66.00|51.00|67.20|
|DeepresearchBII|60.80|39.30|59.70|

💬 55 (+1) open on reddit ↗
▲
102
+1
27👁
r/LocalLLaMA · u/Every-Comment5473 · 21d ago
Still on the Jev waitlist? I hosted OpenJev. It's free, go play with it

TypeSafe announced Jev on Tuesday: you give it data plus typed questions (yes/no, pick-one, 0–N scale) and it returns a probability for every option, crazy fast. I signed up and then refreshed my inbox. A lot.

Meanwhile Matt Mastracci opened vLLM PR #57250, which does the same trick on DiffusionGemma with a single denoising step. The model basically fills in a multiple-choice bubble sheet. My "quick look" turned into three straight days, and now there's OpenJev: an open-source server with Jev's API, so TypeSafe's SDKs work with just a base URL change. If you're still waiting on Jev access, you can start playing today.

Your prompts and answers are not stored, only token counts for your quota. It runs on my RTX PRO 6000, which just got promoted to "production infrastructure" overnight!

Is it any good? Matt ran live evals of Jev vs DiffusionGemma-as-Jev: accuracy roughly tied (198/201 vs Jev's 191/201 across his 8 eval sets), and DiffusionGemma was faster, on a DGX Spark. An RTX PRO 6000 is a different animal:

|Model|Latency|
|:-|:-|
|Frontier LLMs (TypeSafe's numbers)|3–329 s (coffee time)|
|Jev (published)|70–500 ms end to end|
|OpenJev via api.codiv.ai|\~170 ms p50 end to end (\~73 ms on the GPU)|

It's v0.1 on an unmerged vLLM PR. If Reddit hugs it to death you'll see 529s, which is my GPU asking for a minute.

Credits: Matt Mastracci (the core idea and vLLM work are his), TypeSafe (the System One idea and API), NVIDIA and Google (DiffusionGemma), and the vLLM team.

Just a fan of TypeSafe's idea, not affiliated. Feedback, bugs, use-case ideas, or your weirdest yes/no question, all welcome!

💬 30 (+1) open on reddit ↗
▲
73
-2
35👁
r/LocalLLaMA · u/BullfrogScary8947 · 23d ago
[Release] SOTA GGUFs for Qwen3.8-Flash-Next: GSQ-RCO Providing Near Baseline Performance

https://preview.redd.it/e5wn8eyh7vph1.png?width=1080&format=png&auto=…

https://preview.redd.it/8ov5gl8j7vph1.png?width=1080&format=png&auto=…

New Qwen3.8-Flash-Next quantization using GSQ-RCO. Cuts the size of Qwen3.8 Flash Next from around 80-95GB to 68-76GB, while still preserving near baseline quality. Also their Q2\_0 variant claims to be much faster offering 6.2x better prompt throughput in coding.

"Q2\_0 is built for speed. It avoids the quantization formats that rely on large lookup tables: those formats pack more accuracy into a given bit-width, but decoding them costs real time, and on this model that cost dominates inference. Q2\_0 delivers 3.4x the prompt throughput and 1.9x lower end-to-end latency than IQ2\_XS at a slightly smaller file size, and its decode rate stays flat across workloads instead of varying with the content. The trade is a little quality: 89.07 task average against 89.16 for IQ2\_XS, and 3.5 points below IQ3\_XXS. Pick it when throughput matters most, and see *Performance* for the measurements.

The IQ3\_XXS model is the strongest operating point: it matches the base model exactly on AIME25 (100.00) and is within 0.51 points on GPQA-Diamond and 1.14 on LiveCodeBench v6, at roughly one fifth of the BF16 size."

Model link: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

💬 54 (+1) open on reddit ↗
▲
73
+1
31👁
r/LocalLLaMA · u/Fancy-Snow7 · 20d ago
Ternary-Bonsai-2-27B-PQ2_0 is not completely lobotomized

I decided to run prism-ml/Ternary-Bonsai-2-27B-PQ2\_0 through my own set of UNSCIENTIFIC benchmarks.

I needed something to compare it to, so I decided I would compare with another 27B model by filesize: unsloth/Qwen3.8-27B-UD-IQ2\_XXS.

Since anyone considering running a 27B model with whatever VRAM budget these 2 models demand will end up choosing between these 2.

My table will be a bit bear with just these 2 models, so I threw in a few others too that are larger. I know it's a mix of MoE and dense models, but I have the benchmarks on hand so why not include them. Qwen3.8 q3/q4 quants also give you an idea what we are striving to match and there is 3.6 A3B and the newer Ornith 1.5 and Tiel Coder too.

About the benchmarks and what they test.

Most of them only test memory of context and retrieval, phrase reconstruction and understanding of context in different ways.

Standard Needle: This is the kind of needle test everyone runs and most people score > 90%. It hides passkeys in 21 locations of the context and asks the model to retrieve them. Any score below 100% is questionable.

Hard Passkey Needle with decoys: I was not happy with the standard needle test because most modern models pass 100% making it difficult to compare models. So, I developed a hard mode needle test. This, test hides 21 Passkeys in the context but has many decoy Passkeys. The end of the context I list the CONFIRMED passkeys (GUID's) but do not say which Passkey number they are. Asking for Passkey\[01\] means it has to go through all the Passkey\[01\] decoys in the context and compare them to the confirmed Passkeys. So multiple hops around the context are required to retrieve a passkey. When a model does badly at this even upping KV cache to F16 does not help save it.

Phrase reconstruction: I got this one somewhere on reddit and it still trips up some models. It breaks up phrases into multiple parts (8 in my testing) and asks the model to reconstruct the phrase from its parts. Models might leave out a word if the phrase still makes sense.

500 Multiple Choice Science Questions: Exactly that. Just tests science knowledge with 4 options A-D and the model chooses the correct answer. This test mostly shows how much knowledge was lost through quantisation when compared with other quants. Originally the test was designed to count the number of answer flips between 2 given KV quants. But I did not use it that way here. I got this test from a youtuber so the answers are public and possibly trained but as explained you can still see if damage was done to the model's general knowledge in the quantisation process.

Prose Challenge: I wanted to test a models understanding of a document or a prose and its ability to recall facts from that prose as well as test its ability to recall during a long conversation. So, I created a 1000 paragraph prose. I also have 2 questions about every paragraph. I feed it 1 paragraph from the prose, then ask it Question 1 related to that paragraph. Feed it the next paragraph and Q1 for that paragraph, until the context is mostly full. In my testing below, that was 250 paragraphs. Then I ask all the question 2's in a randomised order. How much can it truly remember? A Q1 score below 100% is not very good and means the model lacks attention of even recent tokens a few sentences back. For Q2 the score varies and higher is better. But the larger I make the context the worse the models perform. This has led me to the conclusion to not always chase higher contexts. It's pointless if it suddenly starts to forget most of what was said anyway and a compaction summary in a smaller context will retain more than a larger context. I also use the test to test at various KV quants and it can improve things a little but not as much as you think. But that's not being tested here today.

JS Coding: My latest test I developed. 100 Javascript challenges each requiring it to pass multiple test cases. It's kind of still under development and I have not run this yet for every model as it does take time. The challenges range from Easy to Very hard. Failing 1 test case fails the whole test. Tests are run with reasoning disabled. However, on failure it can retry with up to 16,384 token reasoning budged, then it must answer again. I realise models today are designed to perform best with reasoning but in order to speed up the tests I see if it can pass the test without reasoning first. I also keep track of the number of thinking tokens used, answer tokens, how many tests it had to reason but I won't be showing those here.

Toolery

You can download this bench for yourself. It's not mine but it tests tool use. So much good stats in the app but I will just list the Overall % score and I have not yet tested all models on this one since I just discovered it today.

How I tested

All tests done at 89,088 context, KV q4\_0/q4\_0. Seem a bit odd? I optimise for 16GB VRAM, so all my tests are done at these settings initially and I test higher KV quants if VRAM is available. And as I said higher KV in many cases makes little difference and, in some cases, perform worse. All tests are seeded at start and most repeated 5 times so the results are deterministic. All tests are also done with MTP disabled, so I do not test the draft models KV cache which might be F16. Yes, MTP can change the results in some cases but in my finding it's tiny and mostly does not happen.

Here are the results:

https://preview.redd.it/cmdnqqmxcjqh1.png?width=1143&format=png&auto=…

Findings

Bonsai did not do all that badly compared to Qwen3.8-27B-UD-IQ2\_XXS.

Standard Needle

Bonsai scored close to 100% and UD-IQ2\_XXS did poorly, worse than Q1. Generally, I expect 100% in this test. But notice that Ornith 1.5 and Tiel Coder score low 90's which is a red flag.

Hard Passkey

Not many smaller models can 100% this but a few come close Qwen3.8 Q4 obviously did the best. And ISTA at Q3 does excellent. Bonsai does a bit better than Qwen3.8 Q2 of similar filesize and it's not far from our former favourite model Qwen3.6 35B A3B. But the real shocker here is Ornith and Tiel Coder's scores. These models are supposed to be upgraded 35B A3B models. As you will see this trend continues and these 2 models have serious memory retention issues.

Phrase reconstruction

Bonsai aces this test with almost a perfect score compared to UD-IQ2\_XXS at only 69%. Tiel coder performs worst even worse than a Q1 model.

500 Multiple Choice Science Questions

Bonsai shows almost no knowledge loss compared to even Q4 models. UD-IQ2\_XXS on the other hand does start showing a loss and Q1 even more so.

Prose Challenge

Question 1 I expect 100% and most including Bonsai achieved that. Concerning again that Tiel Coder and Ornith could not even recall from the last paragraph.

Question 2 Bonsai and UD-IQ2\_XXS are close maybe margin of error. Q3/Q4 models outperform it but a large margin. Except Swift, which is a model with significantly less reasoning. Here we can see some of the damage that was done to the model to achieve that. Tiel Close and Ornith again clock in with shocking results. Tiel Coder's 6% is probably as good as just guessing. I would say maybe 3B active parameters are just not enough. But Qwen3.6 A3B scores 39.2% significantly better. I tried upping Ornith's KV to F16 and it improved to 24.4%. I also tried a Q6 quant of the model at q8\_0 which scored 25.5%. End of the day I think Ornith and Tiel Coder have an issue with recall regardless of Quant and KV Quant.

JS Coding

Bonsai was actually able to hold it's own against ISTA Q3. It did burn significantly more thinking tokens and had to reason on many more challenges. UD-IQ2\_XXS on the other hand shows significant loss of coding ability. 10% below Bonsai.

Toolery

Bonsai did better than UD-IQ2\_XXS. I am still learning to interpret the numbers, but the app has options to select your use case and it's applies weights to calculate a score. It also tells you the strength and weaknesses of each model you test. I also found that upping KV quant improves this score but a KV F16 Tail using beellama makes the biggest difference since tool calls are happening in the tail.

Conclusion

If you are VRAM constrained <= 12GB Bonsai might be a model to consider. But it will depend on how you plan to use it. Since I have 16GB I will stick with ISTA Q3 and I can run it with kvarn5/kvarn5 and MTP (kvarn2/kvarn2) and a 1024 token F16 tail.

Disclaimer

These tests do not test intelligence or real-world performance. They are purely synthetic.

💬 42 (+1) open on reddit ↗
▲
62
-4
28👁
r/LocalLLaMA · u/drooolingidiot · 21d ago
We benchmarked 24 LLMs against human writers on 475 creative writing prompts post image

We just released the first version of our Creative Writing benchmark, comparing 24 LLMs against human writers across 475 writing prompts.

Creative writing is subjective, so the rankings aren't meant to predict what any one person will prefer. Instead, they predict what a large group of readers would prefer, using a custom reward model trained specifically on human preferences for creative writing.

Surprisingly, the strongest frontier models already rank above the talented amateur writer cohort, while professional writers still lead by a wide margin.

You can browse the full benchmark, compare the model outputs side by side, and see how the benchmark works here:

https://vulsar.ai/benchmarks/creative-writing-v1/

Curious what you all think of the results!

💬 77 (+1) open on reddit ↗
▲
64
+1
34👁
r/LocalLLaMA · u/Feralzi · 19d ago
Reached 1.89 TB/s memory bandwidth overclocking the CMP 170hx

Overclocking the CMP 170HX 40GB I was able to get the memory bandwidth from 1,386.2 GB/s to 1,890.1 GB/s, that's a +36.4% increase.

Qwen 3.8 27B token generation jumped from 110 T/S to 202 T/S, same config, nothing changed except the overclock.

Just throwing this out there for whoever owns one of these cards. It's good to look into overclocking them as it's potential is severely cut down.

Edit:
GPU wattage is at 300 watts
GPU temps are slightly lower now

💬 60 (+1) open on reddit ↗
▲
66
+3
27👁
▲
1674
+7
34👁
r/LocalLLaMA · u/xenovatech · 22d ago
Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU. post image

The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to <6GB in size. According to the model card, it's 9x smaller than FP16 while retaining 98.2% of the intelligence.
\- Collection: https://huggingface.co/collections/prism-ml/bonsai-2
\- Demo: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels

▲
1318
+7
23👁
r/LocalLLaMA · u/giveen · 21d ago
Alibaba open-sources medical AI model that can detect cancer and nearly 150 conditions

Hopefully things like this let people understand there is good things that can come out of AI.

▲
1273
+6
21👁
▲
1045
+10
41👁
r/LocalLLaMA · u/Nandakishor_ml · 21d ago
Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo post image

UPDATE: Multilingual support added at : https://github.com/NandhaKishorM/laya

Thanks for the exceptional support (https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i\_literally\_built\_the…) and for the dozens of requests to make a generic model, run benchmarks, and create an HF space so anyone can test it. So here you go, guys. I trained an improved model on a large data corpus, its now called Laya. It is trained on a single RTX 6000 Pro (96 GB VRAM); the model architecture is a 421M-parameter non-autoregressive decision model pairing a bidirectional ModernBERT-large encoder with a scratch Transformer head that scores \[MASK\] option markers to resolve typed schemas in a single \~35 ms forward pass. The dataset is a 100% human-annotated corpus of over 25,000 real-world examples across intent routing, fact-checking, moderation consensus, prompt guardrails, rubric scoring, and multi-turn conversation trajectories, without synthetic data shortcuts. The RLCD(unofficial, btw) I did is a policy-gradient reinforcement learning approach that kinda optimizes decision models against strictly proper scoring rules, ensuring maximum reward is achieved only when outputting true, mathematically calibrated probabilities.

NB: It can be run on low end PC as its a small 421M model, cheers

HF space to try: https://huggingface.co/spaces/convaiinnovations/laya-demo

GitHub Repo: https://github.com/NandhaKishorM/laya

HF Repo: https://huggingface.co/convaiinnovations/laya

Thank you to everyone who supported me, shared the story, gave personal DM. It will need more refinement, of course.

If anyone wishes to buy me a coffee, here is the link: https://github.com/NandhaKishorM

▲
855
 
29👁
r/LocalLLaMA · u/SorosAhaverom · 24d ago
CrofAI "cheapest inference provider in the world" gets exposed as an OpenRouter wrapper, routing requests to smaller, cheaper models at up to 20x markup. CrofAI responds to Wire Fraud allegations by denying everything, then backtracking, then 3 hours later wiping their entire online presence

Disclaimer: no AI was used whatsoever to write this post

Cautionary tale about chasing cheap tokens.

exposé: https://kendell.dev/blog/crofaifalse/

reaction by nahcrof, announcing the shutdown of the service: https://x.com/nahcrof/status/2099552389434900643 - now deleted, archive picture: https://i.imgur.com/teOQngH.png

NahCrofAI (crof.ai, nahcrof.com) was an inference provider which had all the latest models at the cheapest price, often significantly below the lowest alternative on OpenRouter. The owner claimed that they are running custom inference engines that allows them to offer tokens for dirt cheap, and other providers are suffering from "skill issues", that's why they are so expensive.

In reality:

  • "CrofAI is an OpenRouter wrapper that silently routes to cheaper or weaker models than what you request"
  • For example, expensive models like kimi-k3 are sold at $2/$10 in/out, but instead routed to GLM 5.3 Flash via OpenRouter, representing a 13.3x multiple on input, and 20x multiple on output
  • CrofAI's "own model family" greg-2-ultra routes to GLM 5.2, greg-1-mini routes to Qwen 3.5 9B. greg-2-super, greg-1, greg-1-super routes to Kimi K2.7 Code. All of these at a significant markup compared to the actual model being served. CrofAI admits in DMs that his claims of the greg family being made by him is a lie.
  • The person investigating details the 5 different attempts by CrofAI at fixing their models being served via OpenRouter after given a heads-up and a lengthy grace period. In all 5 attempts, the only change CrofAI made was attempts to hide the fingerprints of OpenRouter, while still serving models through them
  • Other inconsistencies don't add up either: CrofAI claims to run Kimi K3 on RTX Pro 6000s rented via Vast. That model requires ~802GiB even at the lobotomy level quantization of Q2_K. The largest RTX PRO 6000 machine on Vast has only 8 of them, totaling 765GiB. He also claimed that for the purposes of "investigating" the "issue" of his API routing to OpenRouter, he will have deepseek-v4-flash-0731 running on his local DGX Spark. A Spark has 128GB memory, and is therefore unable to run that model.

CrofAI responded to the exposé by announcing the shutting down of their service; after their failure to provide their own inference, they promise to provide one last thing: a refund to those asking.

UPDATE

UPDATE: around 4:30 AM UTC of Sept 15, the owner published a now-deleted blog post (archive image) writing under the fake pretense that it's his "team" authoring it, stating all of CrofAI founder's claims "were written under a lot of stress, and they described the situation as worse it was", and that a new team is taking over, with the service being resumed in 2 weeks.

At the same time, the CrofAI twitter account was also supposedly "taken over" by the team, starting each twitter reply with "Hey, Nathan here", stating the founder is stepping back and a "team" is taking over everything. This fake pretense act only lasted a few hours, and scared either by the public not buying the Nth fake story of the pathological liar that CrofAI is, or by the public's replies reminding him that what he committed is numerous counts of wire fraud, he has now deleted all his online presence: nahcrof.com and crof.ai return 404, Twitter page is deleted, /r/CrofAI sub is now private.

Here is another image of the owner admitting that he was defrauding customers for the entire 2 year operation of his service, then begging the investigator to help him cover his tracks and not expose him

EDIT: Commenters pointed out that NahCrof is 4chan in reverse. The owner's Discord name was "Devious Flimflam". Flimlam is defined as "deception, fraud". Looks like it was a deliberate scam operation from the get-go, and the owner's age was among the many lies.

I cannot stress this enough: if you bought any credits (even if you used them up) you are entitled to a full refund for every transaction as the victim of fraud. Open a chargeback with your bank for every transaction made. If you used their API, assume that everything was logged and is currently being mined for personal information and API keys to sell on the black markets. Rotate your keys, change passwords, get a new debit/credit card.

▲
827
+6
35👁
r/LocalLLaMA · u/Intrepid_Travel_3274 · 20d ago
With Gemini 4, bench goes up. post image

They claimed open-weight models are dangerous but the benchmarks say otherwise.

Source

▲
577
-6
22👁
▲
569
+7
23👁
▲
537
-3
36👁
r/LocalLLaMA · u/RishiFurfox · 23d ago
Hey, Meta. Where's those Muse Spark weights? post image

It was well over a month since Meta promised to release the weights for Muse Spark.

Back then (10th August), they were on Spark 1.2. Now we're on 1.3 and still nothing's been released. So it begs the question: will they be releasing the 1.2 weights when 1.4 drops? Or will we get whatever's then-current as open weights?

It's ironic given Mark Zuckerberg said at the same time that we can't delay the release of models by "even a month," due to the competition with China. It's been well over a month. He was arguing in the context of new regulations delaying models, but I think it applies equally to the open weights contest as it does to the closed models one.

After all, the Chinese models are all open. That's the competition and point of comparison.

Have Meta given any sort of explanation for why they're sitting on the weights or how much longer it'll take for them to honour their promise? Will we even get them in light of all the attempts at regulatory capture and dire warnings about how AI is dangerous?

▲
493
+6
36👁
r/LocalLLaMA · u/Thin_Pollution8843 · 19d ago
Qwen3.8-Flash-Next Cosmic Arcade oneshot slop game post image

To test what it can do. Qwen3.8-Flash-Next Intel Autoround W4A16 running locally on 4xV620 \~2k prefill and 70ts decode.. Were running around 3 hours. Harness is OMP (I think it made a big difference). Most of the time model was running 2 browsers simultaneously and testing/fixing everything. The most sloppy prompt possible:

create a game where a space traveller in the space he
neets eniemes who shoots in him and asteroids which he should avoid. he have a blaster gun to shoot enemies and asteroid. space traveller in
scafandr and fyoing on the rocket. game should be very lifelike detailed and done with html and js (use any lib you want). 3d game photorealistic.
ofc run the browser to debug and fix stuff always

▲
462
+5
23👁
r/LocalLLaMA · u/skeole · 23d ago
Xiaomi MiMo 2.6 Live Training Dashboard

Cool to see this as it happens!

▲
422
-5
25👁
r/LocalLLaMA · u/Porespellar · 23d ago
Frontier LLM development simplified for politicians: post image

Nobody is buying this “Pace the frontier” nonsense. It makes no logical sense at all. Are American labs really going to take a pause and lose any small lead they still may have over Chinese labs? Does anyone really believe this? This seems like some performative virtue signaling BS. Why are they bothering with this pacing campaign? Someone please explain.

▲
404
-5
31👁
r/LocalLLaMA · u/anomaly256 · 20d ago
General warning about Clore.AI

Hello, I know some of us may be tempted to rent out our expensive GPUs to recoup some of the cost of self-hosting, and it should be obvious that this can be a risky decision. I decided to try hosting my rig on clore.ai briefly to see what kind of revenue it could bring in, keeping a close eye on the process lists from the renters' jobs of course but not digging into their files or anything.

Yesterday I saw a renter scanning and attempting to exploit vulnerabilities to post malware to a columbian betting site from my internet connection. I immediately took the server offline and reached out to Clore requesting them to cancel the order (so the machine wouldn't restart the containers when it came back up) and to block the renter.

I think everyone should know that they flat out refused. Not only did they refuse, they blocked me when I provided hard evidence of what was happening. Since I had reasonable suspicion the renter was abusing my connection, I availed myself of clore's T&C that says a host must not inspect a renter's environment "unless required by law" and given the laws around liability for residential internet connections in my country, I mounted the filesystem offline and inspected it.

I found the logs from the vuln scans and unsuccessful exploit attempts, the malware payload they were trying to post to the site, the reverse proxy request smuggling tactics it was employing, and the AI agent reports that were being generated along the way. I sent this to Clore and requested a way to blacklist renters who abused the platform. Their response was to tell me, directly and without mincing words, to leave the platform entirely and proceeded to block me from their support chat. Their support rep I was trying to reach on Telegram also told me to go away and proceeded to block me as well.

I can only conclude then that they are wilfully complicit with facilitating cybercrime and knowingly turn a blind eye when it's discovered. They didn't even \*try\* to hide it.

I have no idea how better/worse the other platforms are, Vast.AI, Akash, etc. But Clore will abuse your internet connection, deny liability, then block you. They are crooks. Go elsewhere. You've been warned.

I've archived the renter's docker volume and will hold on to it in case any security researchers or legal authorities want to examine it.

▲
400
-2
27👁
r/LocalLLaMA · u/gaviniboom · 19d ago
Seeing how differently people prompt LLMs is funny

So my brother and I both use LLMs for coding. I've started using a local GLM 5.3 Flash instance - q4 qat. My brother uses GPT-6-Astra as his daily driver.

He has mentioned repeatedly to me that his approach is to berate the AI whenever it makes a mistake so that it actually does what he wants it to. This involves a lot of swearing and "are you an idiot!?" to GPT 6 Astra.

Meanwhile I'm here looking at a q4 quant of GLM 5.3 Flash going "aww it's dumb in some ways but it's trying its best, oh it did something!" and being autistically specific with my requests and asking a lot of questions. Yes, I am autistic, so I have learned to communicate with precision, which oddly makes talking to small LLMs easier.

It's so funny imagining him berating a giant model in the other room while I'm here petting a tiny one.

What are yall's prompting styles and what is your main LLM that made you this way?

▲
393
+4
34👁
r/LocalLLaMA · u/FullstackSensei · 22d ago
AMD Plans 10% Price Hike Across GPUs, Chipsets, and Possibly CPUs

Great news! AMD is also considering accepting payment in organs!

Slightly less sarcastically, grab what you can, while you can. Waiting is becoming very costly almost by the day

▲
377
-4
30👁
▲
373
+7
21👁
▲
299
-4
28👁
r/LocalLLaMA · u/buttplugs4life4me · 20d ago
Please stop with the FP4 inference engines for the love of god

Every day there's a new post of some optimized config or new inference engine that is just super good at one specific thing and their claims make sense.

And then at the bottom of the post or maybe after someone asked it says "NVFP4/MXFP4 only".

Okay dude, good job! You made the fastest possible option a little slightly faster, and most likely your output is completely cooked and you get hallucinations left right and center.

Just saw another one in r/ROCm again.

It's fine if you run 4-bit for large models, they've got lots of shit in them so a little bit of loss just means they won't remember that super good spaghetti Bolognese recipe. But running small dense models at FP4 just kills them. Like, completely. Good luck doing something productive when your model suddenly decides 1+1=3.

Just...stop.

Edit: Just going to put this here since some seem confused. a standard Q4 quantisation usually leaves more sensitive tensors in BF16, Q8 or Q6/5. Also, usually the K/V cache is quantized max to FP8/Q8.

What these "inference engines" do is usually fork an existing one (llama.cpp, SGlang, vLLM) and then just quantise \*everything\* down to FP4. Which is great for speed, especially without online dequantisation, but fucks the quality up \*a lot\*.

Your standard Q4\_K\_M/XL quant from unsloth is fine.

▲
268
-4
18👁
▲
263
-3
22👁
▲
252
-3
16👁
r/LocalLLaMA · u/Fcking_Chuck · 24d ago
Koboldcpp v1.121 released
▲
236
-3
23👁
r/LocalLLaMA · u/Cherlokoms · 24d ago
Apple Foundation Models: local AI natively on MacOS 27

Maybe some of you know but I didn’t see any post about this. Apple just made available their AFM model on MacOS 27 natively. Just run fm chat in a terminal.

Disclaimer: I’m an open weight person. I prefer open models and ecosystem, but I’ll still open the discussion.

Did you test them? Build using them? Are these models good?

I feel like this is still a huge step in the direction of local AI that a company like Apple does this and release hardware optimized models.

So what do you think?

▲
229
-3
27👁
r/LocalLLaMA · u/shniydder · 20d ago
I gave Jev, Laya, finetuned ModernCE and Qwen3.5 the controls to Doom post image

I gave Jev, Laya, a finetuned ModernCE-base-nli and a finetuned Qwen3.5-4B the controls to Doom. Thanks to TypeSafe AI for Jev access.

The video lines up the starts of four separate games with the same seed. After that, each model's actions change its own game and what it sees next. The clip shows one preselected episode per model in each of two scenarios, with action probabilities, kill counts and survival time on screen.

  • Input: A short text description from a deterministic Python adapter reading ViZDoom's visible-object labels, bounding boxes and HUD values (health and ammo).
  • Output: One button action. In Defend the Center, that's left, right or fire. In Health Gathering, it's left, right or forward to look for medkits as health drains.

Each local model and its game ran on a single DGX Spark with an NVIDIA GB10 and 128 GB unified memory. Jev used TypeSafe's hosted API. ViZDoom ran at 320 × 240, with a 35 Hz game clock and a target of five decisions per second. The game kept running while the model replied.

Here are the averages over eight seeds per controller, per scenario, using the clear-scene descriptions shown in the video. Episodes were capped at 30 seconds of game time.

| Model | Mean Kills | Mean survival (s) | Call p50 (ms) | Call p95 (ms) |
| --- | ---: | ---: | ---: | ---: |
| Jev 1.13 | 5.63 | 13.03 | 117.3 | 199.8 |
| Laya English | 1.25 | 11.89 | 16.2 | 17.1 |
| Finetuned ModernCE-base-nli | 1.25 | 11.66 | 7.6 | 8.8 |
| Finetuned Qwen3.5-4B (LoRA) | 3.63 | 11.31 | 146.8 | 150.9 |

Call latency is request-to-response time on 12 shared synthetic Doom scenes, repeated for 48 calls per model. p50 is the median and p95 is the 95th percentile. Local model timings include loopback HTTP on one Spark. Jev includes the hosted API round trip. Applying the action adds controller and game-tick delay. None of these models got extra Doom-specific training for these runs.

Inspired by TypeSafe's Doom demo and experiments shared on Reddit and LinkedIn.

My longer write-up about exploring Jev: https://morethanamachine.com/posts/jev-style-decisions-dgx-spark/

Edit: Table overflow fixes.

▲
229
+4
23👁
r/LocalLLaMA · u/Henrie_the_dreamer · 22d ago
Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash post image

Hey all, Henry from Cactus Compute here, I kinda wanted to share our latest model and get feedback from the family :)

Needle 3 is a small foundation model for automation: you give it the functions your app exposes, it reads a request and returns the calls with every argument filled in, or a typed record if what you gave it was a schema. It runs on the device, with no network in the loop. It is on Hugging Face, on GitHub, on PyPI as cactus-needle, and there is a sandbox that runs it in your browser at cactuscompute.com/needle if you want to poke at it before reading further.

1) Trades general capacity for frontier performance on automation tasks

The thing we decided early was that Needle would not chat. Every turn is a function call, and a request no declared tool can serve comes back as an empty list rather than a guess. That sounds like a limitation, and it is, but it is what let a 121M-parameter model be trained on 360B tokens of structured data and spend all of its capacity on three jobs: tool calls, structured extraction and text embedding.

The architecture follows from the same trade. It is a Simple Attention Network: the dense feed-forward layers are gone, replaced by a Monarch Hadamard MLP with 25.6K parameters per layer instead of 4.7M, and the knowledge a feed-forward layer would normally hold sits in an engram, hashed n-gram tables that are read by gather and cost no arithmetic. 70.8M of the 121M parameters live there, so the full model does the arithmetic of a 50M one: 100 MFLOPs per token against 296 for a transformer of the same shape.

https://i.redd.it/fsnfjojny4qh1.gif

We wrote the intuition up if you want the longer version: Simple Attention Networks and the Hadamard MLP.

2) Beats models 10x its size on tool calls and language-to-device control

On Mobile Actions (961 phone commands, scored on the exact call), the 20-layer model scores 86.0 through the shipped 2-bit binary with the confidence gate on. LFM2.5 1.2B is at 82.4, Qwen3.5 0.8B at 76.0, FunctionGemma 270M at 65.1 and Apple's on-device foundation model at 57.6, all at f16. DeepSeek V4 Flash through its API is at 88.4, which is the line in the chart.

https://i.redd.it/luunq717z4qh1.gif

The part we are most pleased with is not the number but how the calls are made. Every argument is a span of the request: the model writes a short derivation first ('living room' -> room; '30' -> brightness) and then emits the call under a byte-level grammar compiled from your schema, so the JSON always parses and an enum can never leave its set. An optional field with no evidence is omitted, a required one with no evidence withholds the call, and the engine drops a call the request negates or excludes. Ask for two things and you get two calls in order.

https://i.redd.it/227396y5z4qh1.gif

Full table across all six suites (tool calling is exact match, extraction is field F1, Needle through the shipped binary, baselines at f16 under vLLM):

|Model|Params|Mobile Actions|DroidCall|BFCL v4|DSTC8 F1|SNIPS gold F1|SNIPS 7-way F1|
|:-|:-|:-|:-|:-|:-|:-|:-|
|DeepSeek V4 Flash (cloud)|\-|88.4|60.5|77.2|80.0|69.4|66.7|
|Needle3-20L-121M|121M|86.0|47.0|50.2|40.7|30.2|24.7|
|LFM2.5 1.2B|1.2B|82.4|35.5|62.0|48.0|43.0|38.0|
|Needle3-16L-98M|98M|80.7|40.0|41.3|28.5|23.5|19.2|
|Qwen3.5 0.8B|800M|76.0|28.0|56.8|49.0|35.0|34.0|
|LFM2.5 350M|350M|72.8|32.5|59.1|20.0|34.0|29.0|
|LFM2.5 230M|230M|69.3|11.5|46.3|53.0|27.0|22.0|
|FunctionGemma 270M|270M|65.1|16.5|46.6|27.0|29.0|14.0|
|Needle 2|45M|63.5|17.0|\-|\-|\-|\-|
|Apple FM|3.0B|57.6|\-|\-|\-|\-|\-|
|Needle3-8L-52M|52M|36.8|36.5|28.2|15.3|16.6|10.1|
|Needle3-4L-29M|29M|11.7|21.0|19.5|6.9|7.7|4.3|

You can see where it is weaker too: BFCL and the extraction suites are where the bigger baselines pull ahead, and the smaller subnetworks fall off quickly on the general task (more on why that is fine in section 4).

3) Matches 2-3x bigger models on structured JSON extraction

Extraction is not a separate mode. You declare the record as the only tool and pass the passage where the query goes; with one tool declared the grammar admits exactly one call of that name, so the shape is guaranteed rather than requested, and the values are grounded the same way as arguments: a field is filled only from a span of the passage, an optional field with no span comes back as None, and a date whose year appears nowhere in the text is flagged instead of invented.

from pydantic import BaseModel
import needle

class Invoice(BaseModel):
vendor: str
total: float
due_date: str
po_number: str | None = None

needle.extract("Invoice from Acme Corp, $1,200.00, due 2026-09-01", Invoice)
# Invoice(vendor='Acme Corp', total=1200.0, due_date='2026-09-01', po_number=None)

It generalised to classification without special training, because an enum is just a constrained value: declare sentiment: Literal["positive", "neutral", "negative"] on a record and you have a classifier whose output cannot leave the set. A watch reads a notification into merchant, amount and date that way, then into a reply, then into a sentiment flag, one record each. On DSTC8 and the two SNIPS suites the 121M model lands between the 230M and 350M baselines, which is the 2-3x in the heading.

4) Intelligence ladder: every depth from 2 to 20 layers a model of its own

This is the part I would most like your thoughts on. Needle 3 is one set of weights, and every depth from 2 to 20 layers is a deployable model. Blocks 0 and 19 are always kept and the rest are added by bisection, so each subnetwork nests in the next; during training each step samples one path, mostly the full model and otherwise a random depth, with the smaller path distilled from the full one. The full-depth model ends up slightly better than an ordinary run of the same size, and every depth below it is trained rather than truncated.

https://i.redd.it/kbwimep4z4qh1.gif

Why we wanted it: a watch, a Raspberry Pi and a phone do not want the same model, and they want to pick the size at deploy time. needle build --layers 8 writes the 8-layer file; the same engine runs all of them. The small depths lose accuracy on the general benchmarks (that is the bottom of the table above), and they get it back when fine-tuned to one product's tools: on DroidCall every subnetwork gains 18 to 36 points, and from 4 layers (29M parameters) up the tuned subnetwork passes DeepSeek V4 Flash.

https://i.redd.it/y646wed3z4qh1.gif

Fine-tuning is LoRA on the frozen base, merged at export, and the Python package does it locally at 4 bits (needle finetune data.jsonl, then needle build). The maths is in Intelligence Ladders and the workflow in Fine-tuning Needle.

5) Runs locally at up to 4k tokens/sec decode speed

The engine is under 1 MB, plain CPU, no GPU or NPU, and the weights are read in place from a single file the engine maps into memory: a 196-byte header carrying the whole architecture geometry, a nameless tensor directory, and the quantised blobs in the order the forward pass reads them. On a Raspberry Pi 5, decode runs at up to 4k tokens/s at the bottom of the ladder and around 400 at the top, prefill from 10k down to 1k. Every response reports prefill_tps, decode_tps and peak_ram_mb, so you can measure on your own device rather than take our word for it.

Every response also carries a confidence score from a calibrated head, the minimum of a post-hoc judgement on the finished call and the decode probability of its tokens. The engine withholds anything under 0.1; above that the number is yours: act at once when it is high, show the call and ask when it is middling, treat [] as a refusal. How we use it is in Leveraging Needle's confidence.

6) 25-121M deployable parameters at CQ2-bit (8-29MB binaries)

The weights are quantised with Cactus Quants: groups of 128 weights are rotated by a Walsh-Hadamard matrix, which makes every group look Gaussian, split into an fp16 norm and a direction on the unit sphere, and the direction's coordinates are snapped to a 4-entry Lloyd-Max codebook. That is 2.125 bits per weight, and the kernel never expands them: it rotates and int8-quantises the activation instead, then does table lookups and sdot against the packed indices. The embedding and the confidence head keep 4 bits, the norms and gates stay fp16. The byte layout, and a twenty-line parser for it, are in The .cact format.

7) For mobiles, wearables, smart home, small robots and microcontrollers

Every target ships a prebuilt engine folder: macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32 (the Ingenic camera SoCs), Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly, and a WASI component with a WIT world. pip install cactus-needle covers the desktop and server platforms with wheel-tagged engines, and needle build --platform linux-arm64 --layers 8 --out ./pi puts an engine and the weights in a folder you copy over. Inference never touches the network, so an air-gapped device only needs the files in place. The list and the runtime surfaces are in What devices are supported.

import needle

@needle.tool
def get_weather(city: str):
"Get the current weather for a city."
return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]

What we would love feedback on

  • Where the grounding rules get in your way. The engine refuses to invent a number, drops a call the request excludes and withholds a required enum the request never names; we tuned those on our suites and would like to hear where they bite on real tools.
  • The ladder. Whether depth is the right knob for you, or whether you would rather have width, and what devices you would want the 2- and 4-layer models on.
  • Extraction cases we have not seen. Nested records, arrays, multilingual text (it is English-first, and non-English text fragments into about 1.7x more tokens).
  • Anything in the tool design guide that turned out wrong for your schemas.

Thanks for reading this far. Happy to answer anything in the comments.

▲
212
-2
22👁
r/LocalLLaMA · u/returnity · 24d ago
Cut Qwen3.8-27B Reasoning Tokens by 40% -- 3.8 'ThinkingCap' benchmarked!

EDIT: Sorry for the unclear title. This model is UkisAI's Swift-Qwen3.8-27B, not a new version of BottleCap AI's 3.6-ThinkingCap. All credit goes to UkisAI for making great fine-tune, and I made this post to celebrate their work. I meant no disrespect by mentioning another model in the title.

I doubt I'm in the minority here when I say I love Qwen models, but the overthinking is a major timekiller. It was bad in 3.6-27B, and it's worse in 3.8. I know there are some who say, "well that's how it achieves such a good performance/size ratio"... But now there's some definitive proof that's not the case: UkisAI's Swift-Qwen3.8-27B!

This model seems to be inspired by Qwen3.6-27B ThinkingCap, which was the version of 3.6-27B I used as a daily driver before switching to the 3.8 series. For those of you who haven't heard of it, ThinkingCap is a fine-tuned version of 27B that uses about 40% less tokens to accomplish comparable benchmarks and general performance as the original model. It's one of those fine-tunes that actually works. I used it daily for months without any issues, and it saved me countless hours.

I had been waiting and hoping that they would release a similar version of 3.8, because it is so slow, despite its impressive performance, but so far none has been forthcoming. However, it looks like UkisAI also enjoyed that model, and took it upon themselves to deliver a sequel. They identified "reasoning-marker tokens that ... trigger overthinking in Qwen’s reasoning rollouts" and penalized them using RL, resulting in fewer overthinking errors. They also employed "a transfer component derived from BottleCap AI's ThinkingCap-Qwen3.6-27B". The end result is an average of 30-50% fewer reasonign tokens for the same quality outputs on a number of benchmarks (see the model card for all of them).

This claim is quite impressive, and I have independently verified their claims and the quality of the model in my own use cases and in coding benchmarks using Aider as an eval suite (with Q8_0 for both models):

|Metric|Swift-Qwen3.8-27B|Qwen3.8-27B|
|:-|:-|:-|
|Pass1 (%)|30.8|27.1|
|Pass2 (%)|75.7|77.6|
|Well-formed diff (%)|98.1|99.1|
|Completion tokens|7,301|12,547|
|Seconds/case|750|1,481|
|Total tokens/solve|12.1k|19.3k|

As you can see, their claims hold true -- Swift accomplished an equivalent success rate in approximately half the time, using 63% of the tokens! This is a huge win for 3.8-27B users, because of course decode drops off more and more the longer the response gets, which is why the time is halved even though the tokens are closer to two-thirds of 3.8-27B.

Anyways, my posts tend to get excessively long so I'll cut it off here, I was just really excited after finishing my eval suite on this model and wanted to share.

▲
208
-3
21👁
▲
203
-1
24👁
r/LocalLLaMA · u/Odd_Caterpillar_2994 · 22d ago
Installing 6 GPUs in a standard case rather than using an open-frame chassis. post image

CPU : epyc 7262

MB : ROMED8-2T

GPU : V100 16GB PCIE x6

I have built a server with six V100 GPUs. I am now testing it and plan to eventually run Qwen3.8-Next-Flash configured with TP2 and PP3.

Because open-frame or server-style cases are large and unattractive, I chose to install all six cards in a full-tower case.

▲
204
+2
22👁
r/LocalLLaMA · u/wFXx · 20d ago
Von: Open-source 395M "System One" model

Took me a while since I'm on a family trip and have limited hardware, but here it is!

Von: Open-source "System One" drop-in replacement for TypeSafe's JEV.

https://github.com/wfzyx/von
https://huggingface.co/wfzyx/von-1.0

It runs entirely on a CPU with 1–2 GB of memory (I haven't spent much time optimizing it yet), responds in 25–300 ms, and beats JEV in all benchmarks. Enjoy!

P.S. I’m open to offers to work at AI research labs. Feel free to ping me if you have an offer.
P.P.S. If you have a GPU, it’ll be faster, but a GPU isn't required.

▲
192
+3
16👁
▲
178
+3
25👁
r/LocalLLaMA · u/AdRepulsive7837 · 22d ago
still doesn’t get what Jev is…..is it just a more generalised BERT?

Looking at jev launch website and demo video on x.com…. it seems like it’s a very intelligent classifier with custom prompt and custom criteria instruction reading capabilities. It can do well defined narrow and well defined task

Me, following NLP since good old days of word embedding and BERT,,, be like asking….

Isn’t that BERT?

yeah i know BERT need fine tuning to adapt to custom domain, but can Jev be like generalised form of BERT?

▲
168
+2
22👁
r/LocalLLaMA · u/Mysterious_Hearing14 · 23d ago
Openjev post image

https://huggingface.co/AlexWortega/openjev

I build an openjev, it can play games and do everything what jev can. and yes - it's trained as crossencoder

▲
166
+1
15👁
r/LocalLLaMA · u/ali_byteshape · 21d ago
Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison post image

Hey r/LocalLLaMA,

Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark.

We wanted to see how they compare under the same methodology we use for the rest of our Qwen3.8 evaluations, so we ran Bonsai 2 through our own benchmark suite.

One important clarification: these are our evaluation results, not Prism’s reported benchmark numbers.

We used Prism’s fork/runtime to run their models, while keeping the workloads, benchmark suite, and evaluation methodology consistent with the rest of our comparison.

Our evaluation includes separate Instruct and Thinking benchmarks. For Thinking, we use medium thinking effort with the recommended sampling parameters.

We added Bonsai 2 because the models have become a relevant part of the Qwen3.8 landscape, and we wanted to provide a common reference point for people comparing the available options.

Different providers often report results using different benchmark suites, runtimes, reasoning settings, sampling parameters, and evaluation methodologies, so those numbers are not always directly comparable. Running the models through the same evaluation gives another reference point for understanding the trade-offs between quality, model size, and throughput.

Updated comparison and results: https://byteshape.com/blogs/Qwen3.8-27B/

▲
164
+3
23👁
r/LocalLLaMA · u/Jorlen · 22d ago
Does anyone use uncensored models purely for coding?

It sounds like a stupid question, and I do apologize if it is... but I've seen several people mention that coding models are better uncensored due to the fact that they don't have to constantly run prompts through the "is this okay" sort of checks.

Is this hogwash? Is it true? And more importantly, does anyone have any sources to confirm it?

Anecdotal evidence is fine too if you've tried and compared them.

Personally, I have never bothered because I'm too worried the de-censoring would damage the weights. The juice never felt like it was worth the squeeze... but maybe I was wrong?

Edit: Either I'm unclear or people are misinterpreting my request: Specifically, I mean for every day coding (not for hacking, not for reverse-engineering) but just for regular coding of new apps, etc. The question is: will the uncensored model produce better code faster (without reasoning so much) because it no longer has to worry about "is this alright" when it questions everything...

Edit 2: Decided to test the HuiHui Qwen 3.8 27b (UD-Q8\_K\_XL) quant myself. So far, it reasons far less, and I have yet to have any issues with its coding quality. Granted, I've only been testing it for about 8 hours (straight...) in an active project. Its reasoning is far shorter, it seems far more confident in its responses and as such, uses far less context to achieve the same result. I will continue testing for another week; it's pitted against the Dirk template version of Qwen 3.8 27b right now (same quant) which I'd been using the previous week.

▲
152
-2
25👁
r/LocalLLaMA · u/KURD_1_STAN · 21d ago
bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai\_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai\_2/q3.5 60.8 / 72.4 = 84.0%
Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release \[2\], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

^(also 3.8 35b qwhen? plsss)

▲
149
 
23👁
r/LocalLLaMA · u/zyxciss · 19d ago
I tested 9 LLMs on the exact same web-dev prompt for ~8 hours — RTX 3060 12GB results (Rate the best!) post image

I’ve spent basically the last 8 hours testing different models on the exact same web-development prompt, and I finally finished.

The whole point of this nine-hour test was that which local model matches the frontier-level intelligence at size and could fit easily in an RTX 3060-like consumer card.

My setup:

  • GPU: RTX 3060 12GB
  • RAM: 16GB DDR4, single-channel
  • OS: CachyOS (Arch Linux)
  • Local models were run through my local llama.cpp setup.
  • Same prompt for every model.
  • I recorded the generations so you can actually judge the websites yourself rather than relying on my description.

Prompt

Build a polished, production-quality single-page website for a fictional high-end technology studio called NOVA//LABS.

Goal: make it look genuinely designed by a strong human frontend developer, NOT like generic AI-generated SaaS UI.

Requirements:



\* Use plain HTML/CSS/JavaScript or React + Tailwind if you strongly prefer it.

\* Everything must run locally with minimal setup.

\* Create the entire project/files yourself.

\* No backend, authentication, database, or unnecessary complexity.

\* Responsive desktop + mobile layout.

\* Strong typography, spacing, hierarchy, subtle motion, and excellent visual composition.

\* Dark, sophisticated visual language with restrained use of gradients/glows.

\* Avoid the typical AI-slop look: no excessive rounded cards, giant gradient blobs, random glassmorphism, meaningless statistics, or generic "Empowering the future" copy.

\* Make the copy specific and believable.

\* Include:

1. A striking hero section with a concise headline.

2. A subtle animated visual representing an abstract computational system.

3. A small selected-work/projects section.

4. A concise capabilities section.

5. A strong closing CTA/footer.

\* Add tasteful interactions such as hover states, scroll reveals, and subtle cursor/mouse effects where they genuinely improve the design.

\* Prioritize visual quality over feature count.

\* Use freely available CDN assets only if genuinely necessary; otherwise create visuals with CSS/SVG.

\* Keep the implementation reasonably small and understandable.



Most importantly: make strong design decisions yourself. Do not explain your design choices before building it. Start by creating the project and finish with the exact commands needed to run it.





(SELF CONTAINED HTML WITH JS AND CSS)

I wanted to see what the models actually build, not just how well they explain code.

The models

1. Gemini 3.8 Flash

\~3 min 12 sec

Used Antigravity and consumed roughly 9K tokens.

This was one of the frontier-model reference points for the test.

2. GPT-5.6 Sol

\~1 min 6 sec

Token usage wasn't available to me.

Extremely fast compared with the local models, so this was another useful frontier reference.

3. Claude Sonnet 5

\~4 min 56 sec

Token usage wasn't available.

Also included as a frontier reference. (I was only able to use Sonnet 5 as my Claude-Code Max subscription had expired)

Local models

4. Bonsai 2 27B Ternary

\~45 minutes

  • Native ternary / \~2-bit model
  • Model size: \~7.66GB
  • Average generation: \~34–36 tok/s
  • Context: up to roughly 102K
  • Used \~52K tokens out of a 122K context during this run
  1. Qwen 3.8 27B GSQ-RCO-IQ3-XXS + MTP

\~57 minutes

  • Model size: \~10.4GB
  • High thinking enabled
  • \~29 tok/s around full context
  • Around 40 tok/s with a much smaller/near-empty context
  • Context used reached roughly 75K
  • Context was compacted twice
  • Available context for this particular run was around 49K after the relevant setup/limits

this was probably the most interesting local result for me.

6. Qwen 3.8 27B Q4_K_M Unsloth Dynamic 3

2+ hours

  • \~16.4GB model
  • Q4\_K\_M
  • High thinking enabled
  • Full-context generation dropped to roughly 4 tok/s
  • Context reached roughly 96K
  • Obviously requires significant CPU/RAM offloading on a 12GB GPU

Flags used : --jinja --reasoning-preserve -fa on -fit off -ngl 99 --override-tensor "blk\.([0-9]|[1-3][0-9]|4[0-5])\.ffn_.*=CPU" -ctk q4_0 -ctv q4_0 --gpu-layers-draft all --spec-type draft-mtp --spec-draft-n-max 2 -lv 4 --no-mmproj -np 1 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --presence-penalty 0.0 --repeat-penalty 1.0 --load-mode none --no-warmup -b 256 -ub 128 -c 98304

7. Qwen 3.8 27B GSQ-RCO-IQ3-XXS + MTP — thinking OFF

\~12 minutes

Same general Qwen 3.8 GSQ-RCO model, but this time I disabled thinking.

It used roughly 12K tokens and produced the site dramatically faster.

This was a particularly useful comparison because it shows how much the reasoning mode itself can affect local generation time.

8. Ornith 1 9B Q4_K_M

\~2.4 minutes

  • Model size: \~5.4GB
  • \~74 tok/s
  • Native context: up to 262K
  • This generation only used around 2.6K tokens

This is the speed monster of the local group.

9. Ornith 1.5 35 A3B Q6

\~30 tok/s

  • Model size: \~22.4GB
  • \~30 tok/s
  • Context available for this run: around 128K
  • Obviously heavily dependent on offloading because of the model size

Quick summary

|\#|Model|Approx. time|Local?|Generation speed|
|:-|:-|:-|:-|:-|
|1|Gemini 3.8 Flash|\~3:12|❌|—|
|2|GPT-5.6 Sol|\~1:06|❌|—|
|3|Claude Sonnet 5|\~4:56|❌|—|
|4|Bonsai 2 27B Ternary|\~45 min|✅|\~34–36 tok/s|
|5|Qwen 3.8 27B GSQ-RCO-IQ3-XXS + MTP|\~57 min|✅|\~29–40 tok/s|
|6|Qwen 3.8 27B Q4\_K\_M Unsloth Dynamic 3|2+ hrs|✅|\~4 tok/s at full context 8 tok/s at empty|
|7|Qwen 3.8 27B GSQ-RCO-IQ3-XXS, thinking OFF|\~12 min|✅|—|
|8|Ornith 1 9B Q4\_K\_M|\~2.4 min|✅|\~74 tok/s|
|9|Ornith 1.5 35 A3B Q6|—|✅|\~30 tok/s|

My personal take

For local models specifically, the one that impressed me the most was Qwen 3.8 27B GSQ-RCO-IQ3-XXS.

It hit a pretty interesting balance between:

  • actual design quality
  • coding ability
  • context handling
  • generation speed
  • fitting within a 12GB GPU setup

The Qwen 3.8 27B Q4\_K\_M Unsloth Dynamic 3 was also interesting from a quality perspective, but the speed penalty once you're deep into the context is huge.

Bonsai 2 27B Ternary was also surprisingly usable given that it's a \~7.66GB ternary model.

so its Qwen 3.8 27B Q4\_K\_M > Qwen 3.8 27B GSQ-RCO-IQ3-XXS \> Bonsai 2 27B Ternary

I've attached the screen recording showing the outputs.

Especially interested in other RTX 3060 / 12GB setups ;0

If possible Someone please post down GPT-6-ASTRA's results if they have a codex subscription.

▲
146
-3
24👁
r/LocalLLaMA · u/sadnessdevil · 23d ago
You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

I'm pretty sure it can be done with any model based on qwen4exp, which Qwen's next local models will be based on. You can use a quant that barely fits in VRAM and still run at the model's maximum context length without kv cache quantization, since most of the KV cache can live in system RAM.

I actually made it working on vLLM and now I get 1M context with 3x 3090. I get \~80 tok/s at short context, dropping to \~60 tok/s once QSA reaches its 2048-token budget, after which decode speed stays flat as total context grows. The throughput is pretty good too, and I get like 150tk/s @ 4 concurrent requests. Prefill at 248k reaches 3,701 tok/s. (The patches and the model are available on my huggingface page if you're interested)

Decode speed is a bandwidth problem. Each decode step produces one token, and to produce it the GPU reads every weight and every piece of attention state that the step needs. On a single stream the card spends most of the step waiting for memory rather than computing. So the size of that per-step read sets the token rate.

This is why a normal model keeps its KV cache in VRAM. Take Qwen3.8-27B, which is built on the Qwen3-Next architecture and shares most of its properties with Qwen3.8-Flash-Next (\qwen4\_exp\). It still has one full attention layer every few layers, and a full attention layer reads its entire KV cache on every step. That read grows with the context, so decode gets slower as the conversation gets longer. It also grows past what any host link (such as PCIe) can carry, so the cache has to sit next to the compute.

The numbers of this model show the size of the problem. One QSA layer holds 2 key/value heads of 256 dimensions, as K and as V, in 2 bytes each, which is 2,048 B per token. At 262,144 tokens that is 512 MiB for one layer, and 6 GiB for all 12 layers on every single step. A PCIe 4.0 x16 slot carries about 32 GiB/s, so a host-resident cache of that shape allows about 5 tokens per second.

Here's an interesting part, Qwen3.8-Flash-Next avoids this in two ways:

Only 12 of the 48 layers have a KV cache at all. The other 36 layers are gated delta-net layers, a linear attention whose recurrent state has a fixed size. That state does not grow with the context.

Those 12 layers also do not attend over the whole context. QSA runs a cheap indexer over a pooled, compressed key, where \indexer\_head\_dim=128\ divided by \indexer\_compress\_ratio=4\ gives the pooled width. The indexer selects at most \indexer\_budget=2048\ positions. The layer reads the main KV rows only for the positions that the indexer selects.

So \indexer\_budget\ bounds the bytes that a decode step reads, and the context length does not:

\\\`

2048 selected x 2 kv heads x 256 dim x 2 (K and V) x 2 B = 4 MiB per layer

x 12 layers = 48 MiB per token

\\\`

Take an example, at 80 tok/s that is about 3.9 GB/s across the link. It is a small fraction of a PCIe 4.0 x16 slot, and most of it overlaps with compute.

Only few things need to stay on the GPU. The model itself, and a 2-byte slot plus the pooled index key, which is \1 x (128 / 4) x 2 B = 64 B\. Together they are 66 B per token per layer, against 2,048 B for a full row.

▲
145
-1
23👁
r/LocalLLaMA · u/HFq_Dev · 20d ago
Built a home server from an old PC with GPU upgrade. Qwen3.8 27B runs at ~30 tokens per second. post image

I needed a relatively simple but acceptable level of AI for working on one project. I didn't have any heavy requests, I just needed to give the AI access to the project files so it could search through them for bugs and stuff. I already had an old computer that I decided not to throw away and instead give it a new life as a git server (and sometimes a minecraft server).

The pc specs are ancient by today's standards:

CPU: i7-4790K 4.6 GHz

Motherboard: MSI Z97 Gaming 7

RAM: 32 GB DDR3 2400

PSU: 750 W

Well, my idea stopped at maintaining the computer, because the gpu, a GTX 1070, was overheating. It needed a complete repaste, but the cooler screws were completely stripped, so while trying to remove the cooler I accidentally knocked off several important smd components with a screwdriver. R.I.P. GPU.

Without a GPU, inference was running entirely on the CPU, and only MoE models were kind of usable, giving around 10–20 tokens per second, while dense models couldn't get past 3 tokens per second. I didn't even get to test it with the GTX 1070, because I decided to service it lol.

I started looking for a replacement on the secondary market, but quickly realized that this would cost too much for the minimum entry point I wanted for experimenting with AI. Until my eyes fell on mining cards. There were plenty of cmp40hx, cmp50hx, cmp70hx and cmp90 cards for sale, and the prices were pretty reasonable(it was a month ago), considering that I was originally looking for a cheap replacement for my dead one.

Getting closer to the actual build, I started calculating how much vram I would need for ± acceptable AI with tolerable speed, and after getting inspired by this sub I decided to take a step further and went with a modified cmp50hx with 20gb of memory and pcie modded to 16 lanes. Very quickly after that I bought another one, this time unmodified (10 gb, only 4 pcie lanes). So, together that's 30gb of vram. Both cards cost me $250 in total (it was also a month ago, right now they've suddenly doubled the price).

Luckily for me, around the same time new driver patches appeared that almost completely remove the limits on their compute performance, and even add pcie 2.0 support (these things have pcie 1.1).

I initially tried one of the newly released cmp50hx driver patches, but it ended badly and I had to spend a lot of time trying to get the drivers working. The patches were new and didn't account for the 20gb version. Later the author fixed that, but even then the driver didn't work for me because of some other problem that I don't want to get into.

I went digging through the driver's github issues and quickly found a guide posted there by another user.

And yes, now the drivers work, the cards are detected and even pcie 2.0 works, but not without problems. The author of the guide said that pcie 2.0 support was only confirmed on the X99 chipset. Well, it also works on Z97, however after waking the computer from sleep the driver crashes completely. After looking into it a bit, I quickly came to the conclusion that the problem was specifically with the pcie patch. Disabling sleep completely solves the only problem I had while using them. :)

Without the specially patched drivers, Qwen 3.6 27B did around 15–20 tokens per second without mtp. MoE models were faster, giving 45–50 tokens per second. With the new drivers, performance doubled. With Qwen 3.8 27B(mtp on) I get around 30–35 tokens per second now, and prompt processing is around 300–400 tokens (including degradation as the token count increases). Ornith 1.5 35BA3B(Heretic-MTP-APEX-I-Balanced) gives around 80–100 tokens per second. I capped GPUs at 180w power limit due to the psu I have, I don't want it to work at its limits, but running them at their 225w surely boosts speeds.

In general, I ended up making a lot of presets for different quantizations with different quality and KV cache sizes (I still need to test all of this in real work), but if we take the better options, I managed to get a Q6K model with 130k context (K – Q8\_0 and V – Q5\_1).

I also followed a guide for running 27B Qwen with large context on limited vram. Using the same general approach, I managed to get 256k context with K and V Q8\_0. The speed is slower though, around 10–14 tokens per second, and prompt processing is around 40–50 tokens. Maybe I can tune it even further. I needed this preset for tasks that I can leave generating overnight :3

Overall, I'm satisfied with the result.

Now the main problems I ran into, not counting the drivers:

  1. Not enough VRAM. The ideal option would be having 2 identical cards with the same amount of memory. You can run dense models in tensor mode, split the weights evenly and get increased generation speed for basically free + fit the full context. I tried many variations of Qwen 3.8 27B quantization, but the only one I could properly run in 1,1 tensor split was Q4KM (the Unsloth one) together with mtp = \~40 tokens per second. However, there is critically little space left for the KV cache, because it gets distributed together with the model weights, and the second 10gb card simply became the bottleneck. Without mtp, running models with a 1,1 split basically loses its purpose. Pcie 2.0 and the number of lanes probably also play a significant role here. That could in principle be solved by adding more pcie lanes to the second GPU and perhaps buying an nvlink cable(who even does that?), but I decided it wasn't worth it just to get another 5–6 tokens per second.
  2. No NVMe SSD. Yeah, all models are loaded from a sata ssd so the speed is around 500mb. It's terrible. The motherboard actually has an m2 sata slot with a pcie 2.0 x2 interface, but even its 1gb per second would be too slow for fast model loading. This could be solved by installing an expansion card into one of the pcie slots (there is one free pcie 3.0 x4 slot), but the current price of those things including the ssd is too high, considering that I'm building a cheap system from what I already have with minimal additional spending for an acceptable result. Model switching takes 1–2 minutes. But whatever. (Not whatever, i'm buying a cheap used 256gb nvme ssd :D)
  3. Amount of RAM and Linux (Ubuntu Server 24.04). Apart from AI, I also run gitlab on the server. And here is the problem: after loading a model, all available ram gets cached by the system for the model files. I'm talking about file cache, not KV. The system was leaving around 200–300mb of free ram for everything else. As a result, openwebui and gitlab started acting laggy (after the model was loaded), as well as the kde plasma interface I installed. I don't completely understand why linux decided to keep this cache until the very last moment instead of freeing it for other programs. I tried adding the no-mmap parameter to the model presets, but nothing helped, and I had to manually clear the cache after loading models, which obviously wasn't acceptable. Together with chatgpt (who else?), I made a command that launched the model and then cleared the cache. It turned out that this broke llama-server, causing model switching to stop unloading the previously loaded model.
  4. Llama-server flexibility. The list of presets is defined in the models.ini file, where each parameter is in key-value format. For running Q6K with 256k context according to the guide, I needed to set GGML\_CUDA\_DISABLE\_GRAPHS=1, which applies to the entire cuda environment and remains active even after unloading the model. That's undesirable, because with it enabled I lose 1–2 tokens per second on other presets.

So for one specific preset I need to enable cuda graphs, while for the other presets I need to disable them. The llama-server parser does not support things like this, and doing it manually is not an option either.

Together with the other problem with ram cache getting stuck, this led me to making an alternative way to launch the models and proxy requests to llama-server.

To solve problems 3 and 4, I made a launcher (well, chatgpt did, because I'm not a server/python specialist) that proxies requests to llama-server but takes over the functionality of collecting model presets from .sh files and launching them. It also clears ram page cache after loading a model.

In case someone needs that launcher, I can leave it in the comments, along with any other links to the drivers, fixes, build params, etc. Just ask. (Reddit removes the post when I include them, not enough karma, I guess.)

Overall, I’m pretty happy with how this setup turned out. The performance is much better than I expected from these cards, especially considering how cheap they were. I had a hard time getting the patched drivers to work and linux didn't make my life any easier, and sometimes I even regretted buying these GPUs, but in the end, it was worth it. Qwen 3.8 27B really works like Opus 4.5

▲
143
+3
21👁
r/LocalLLaMA · u/No_Issue_8224 · 21d ago
MiniMax Code goes open source

MiniMax has open-sourced the terminal version of MiniMax Code:

https://github.com/MiniMax-AI/minimax-code

How can developers verify the content that encoding proxies read, send, and store? This is a topic that has been widely discussed recently.

Open sourcing the agent doesn’t automatically answer every privacy or security question, but it gives the community something concrete to inspect.

The repository includes:

  • interactive TUI and headless execution
  • code editing, shell commands, diffs, and test verification
  • permission controls and sandboxing
  • Plan Mode and resumable sessions
  • subagents, plugins, skills, and MCP
  • BYOK with OpenAI- and Anthropic-compatible providers
  • ACP support for compatible editors and clients

First-party code defaults to the MIT license.

A few important caveats: this is a 0.4.12 source preview, the desktop app source is not included, and—as the repository itself notes—a matching version number does not prove identical build provenance between the published package and source checkout.

Still, releasing the agent layer is a meaningful step toward auditability. I’d like to see the community examine its network behavior, file-access boundaries, telemetry, and reproducible-build story next.

https://preview.redd.it/zvrakmgejaqh1.png?width=1198&format=png&auto=…

https://preview.redd.it/13pjdngejaqh1.png?width=1206&format=png&auto=…

https://preview.redd.it/z1ekjlgejaqh1.png?width=1200&format=png&auto=…

▲
135
-4
20👁
r/LocalLLaMA · u/NineThreeTilNow · 22d ago
Update : Small model + Engram

I posted something about a 9b model a few days ago.

The real problem was 2 fold. Someone suggested the size was too big to prove it all out. It's a fair argument. I needed the model depth though. The second problem was the "ability" of these models and the fact that labs (with money) produce these models still. No real usefulness in my mind. A 2b model? maybe interesting. Want a chatbot? - this could be it.

Further, the Llama licensing at the core of the model was highly problematic and had to be thrown in the trash. It's vastly too restrictive and I couldn't Apache 2.0 anything.

So I moved to using the OLMo tokenizer. Except I shrunk the d\_model down to 2048 so I could build a tiny 2b model.

The size of the vocab with 1/2/3 gram forces the Engram table table to a full 1b. That's 50% where DS suggested like \~10-20%.

So basically 2b model + 1b Engram.

The model, because of the depth now allowed by 2048 d\_model allows me to push 40 total SWA/Global blocks to take advantage of attention based residual streams (Moonshot Kimi K3 style) in a dense model.

Data, is pulled from standard Wiki style data sources in my initial training. The data is pumped through the 7b OLMo model to generate the probabilities. This tells the model that the next token is a probability of 32 different tokens.

The big surprise was the results. Even after ONLY 15m tokens, the model is surprisingly coherent. Had I given pure 1 hot cross entropy, 15m tokens would do nothing. I'd get gibberish. However, I borrowed embeddings, and LM\_head that was down projected from 5k -> 2048 d\_model. This preserves \~65% of the data the "big" model had in the embedding when spectrum analysis is done.

The training is limited at 15m tokens, so it's like... babbling about Roman war history and stuff that exists in that token subset. It doesn't fall in to the classic early model phase where it repeats a token over and over though. That's what pretraining 15m generally gives you at first.

Anyways, this was an update and progress report for anyone who cares about this crap... I'm still processing all of the Wikipedia chunk from HuggingFace and working to train it. It'll take time.

If you read this far, thank you. If you have questions, I'm happy to reply. Someone suggested I was a kook who didn't understand ML previously. I started in ML some 20 years ago and held a brief (6 month) stint at Anthropic red teaming the original Opus 4 model before their "Constitutional" paper came out. I'm vaguely referenced in the paper as a "red teamer" I guess. I do this mostly because I love it.

I figured if I built an Apache 2.0 model, I'd at least want people to play with it. It appears 100% trainable on 24gb of VRAM. Training code / Model / Data will follow eventually. It's on HF / Github for now.

The old post is here :

https://old.reddit.com/r/LocalLLaMA/comments/1wezm58/is\_there\_still\_strong\_interest\_in\_a\_dense\_9b\_model/

edit;

Update can be found here :

https://www.reddit.com/r/LocalLLaMA/comments/1wnzu2f/engram_gone_wild_2b_mode…

▲
136
-2
19👁
r/LocalLLaMA · u/enrique-byteshape · 24d ago
ByteShape Qwen 3.8 27B: To KL Diverge or Not to KL Diverge, Part 2: Metric Boogaloo post image

Hey r/LocalLLaMA,

We’ve released our full ShapeLearn GGUFs for Qwen 3.8 27B.

Blog / Download models

TL;DR

  • 3.84 bpw (GPU-5) reaches 99.63% of BF16’s aggregate score of 8 benchmarks, being the most accurate quant we’ve evaluated; 3.23 bpw (GPU-4) reaches 98.72%. These average BF16-normalized scores across instruct and thinking benchmarks.
  • All five new models sit on the measured quality/speed-bpw frontier across six GPUs. In this model’s case, lower BPW translates directly to TPS. Comparisons include Unsloth v3, ISTA-DASLab, AtomicChat and Bartowski (not Bartowski’s newest release). Congrats to the team at ISTA for also landing a frontier model.
  • DFlash2 delivered 1.34-2.10× baseline throughput; MTP delivered 1.28-1.66×, with temperature sampling rather than greedy decoding.

Lite held up very well. As we expected.

We released ShapeLearn-Lite quants a couple of days after Qwen arrived: less optimization, targeted sanity checks, full benchmarking after release.

Then Unsloth v3 arrived with lower KLD at several comparable sizes. Lite looked overtaken, until the task results came in. Three of six Lite models made the quality/speed frontier against twelve Unsloth v3 models in our RTX Pro 6000 comparison. Pretty good for an impatient release. Full ShapeLearn now pushes that frontier further.

Which brings us to KLD.

Unsloth Dynamic V3’s UD-IQ3\_S had \~20% lower KLD than our similarly sized smallest Lite model, but scored 95.55% versus Lite’s 97.33% of BF16’s aggregate benchmark score.

Closer token distributions did not mean better task performance. KLD is useful to avoid a quant that has fallen over the edge, but it isn’t a quantization leaderboard.

That distinction is the subject of our paper on KLD and quantization fidelity, recently accepted for publication to the EMNLP 2026 Industry Track. We also released blog post version of the paper a few weeks back.

We benchmarked this release on RTX 6000 Pro Blackwell, RTX 5090, RTX 4090, RTX 3090, RTX 4080 and RTX 5060 Ti. The benchmarks we used to measure quality are: GSM8K for math, IFEval for instruction following, MMLU for general knowledge, LiveCodeBench V6 for coding, Multi-IF for multi-turn and multilingual instruction following, ACEBench for tool use and agentic tasks (both thinking and instruct), Multiple HumanEval for coding (thinking) and BFCL V4 for tool calling and agentic tasks (thinking).

If you want to dive deeper or choose the best model for your use case, the blog has the complete results across all tested GPUs, along with the methodology, model sizes, and full legend.

▲
121
-3
20👁
r/LocalLLaMA · u/jacek2023 · 21d ago
inclusionAI/Realtime-Venus · Hugging Face

do you want some omni? here is omni for you

[](https://huggingface.co/inclusionAI/Realtime-Venus#1-🧭-overview)1. 🧭 Overview

This repository hosts two checkpoints of the Realtime-Venus system:

  • Realtime-Venus-Omni (Realtime-Venus-Omni/): the 9B audio-visual interaction model. It continuously watches and listens, decides whether and when to respond, and generates text and speech on a shared causal timeline. Adapted from MiniCPM-o 4.5, it supports proactive interaction, semantic interruption handling, and training-free long-video memory.
  • Realtime-Venus-Audio (Realtime-Venus-Audio/): the audio-focused checkpoint on the same streaming backbone, for audio understanding and audio-driven conversation with text or speech output.

Both directories contain model weights and custom Hugging Face Transformers code. The asynchronous Realtime-Venus-Harness and its external tool integrations live in the GitHub repository.

[](https://huggingface.co/inclusionAI/Realtime-Venus#2-✨-highlights)2. ✨ Highlights

  • Native full-duplex conversation: keeps perceiving while speaking and distinguishes backchannels, interruptions, corrections, and redirections.
  • Omni-Proactive interaction: continuously processes temporally aligned video and audio, and initiates a response when an event warrants it — without waiting for a user prompt.
  • Delegation: emits in-stream <delegate> requests on the shared causal timeline and consumes asynchronous backend results the same way, so external tasks never block the ongoing conversation. (Executing requests requires the Realtime-Venus-Harness runtime, available in the GitHub repository.)
  • Training-free long-video Memory: archives visually informative moments, retrieves query-relevant and non-redundant evidence, and reassembles the corresponding audio-visual context — no additional training required.
  • Text and speech output: generates response text together with native speech through the bundled Token2wav resources and a reference voice.
▲
119
-4
21👁
▲
123
+2
23👁
r/LocalLLaMA · u/peonist-ai · 20d ago
Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0) post image

Hi.

I saw some feedback that halogen was degrading at context depth. So I fixed that.

Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0:

  • decode at 1,004,581 tokens of context: 27.3 to 38.3 tok/s (default speculative drafter)
  • decode at 258,794: 42.9 to 45.0
  • prefill at 1,004,581: 790 to 937 tok/s, 21.2 to 17.9 minutes cold
  • prefill at 258,794: 1,086 to 1,114 tok/s

Conditions: Ryzen AI Max+ 395, 128 GB. The 262k and 1M rows are one cold request each at the 1M configuration (HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576), greedy, 64 tokens, the rates the response's \timings\ report. The 32k row is the standard ten-prompt served mean and did not change. A follow-up turn over the prompt cache at 1M reaches its first token in about 0.55 s; the numbers above are the cold path.

To run it at 1M: add -e HALOGEN_ROPE_YARN=4 -e HALOGEN_CTX=1048576to the README's podman line; it needs the 128 GB box. Release notes and the full table:

https://github.com/peonist-ai/halogen-flash-server

If you have a 1M sweep of your own, I would like to see it rerun on 0.12.0.

Thanks for all your support, especially https://huggingface.co/nightvich

▲
125
+4
17👁
▲
124
+4
25👁
r/LocalLLaMA · u/Fancy_Fanqi77 · 21d ago
Steer LLMs and Agents at the Token Level: An interactive tool for token visualization & control, model inspection and data annotation. post image

onPanda is designed for geeks, power users, curious minds, and engineers. Its UI is built for deep exploration and efficient data annotation.

\- The core loop is simple: hover over a token → click an alternative or edit freely → continue generation. You can edit every part of model output exposed by onPanda, including reasoning and tool calls.

\- Edit prompts directly, branch tool calls, and use a tree structure to record branch history. This makes onPanda useful for model inspection and prompt engineering.

\- Support multiple modalities, including images, video, and audio; use tool calls and connect MCP servers to perform tasks in real environments.

\- Connect popular harnesses such as Claude Code, Codex, and OpenCode to execute tasks. Explore and compare their tool sets, system prompts, skills, and memory mechanisms.

\- onPanda includes browser-agent, an agent that runs in the user's browser without installation. It uses the browser as its harness and provides JavaScript execution, information retrieval, interface interaction, multimedia I/O, local file access, and persistent memory.

\- onPanda stands for on-Policy Alignment Data Annotator.

I have been building onPanda since 2024.09, it took two years for it to gradually enrich its functionality and ease of use. In my opinion, onPanda is very suitable for the r/LocalLLaMA community. Any feedback and evaluation are welcome.

Try it online (works on mobile): https://onpanda.diyer22.com/

GitHub repo for self-hosting: https://github.com/on-panda/on-panda

▲
114
+3
25👁
r/LocalLLaMA · u/Brief-Tap-6616 · 25d ago
If you have a 3090, or other 30xx for local LLMs, I have something for you

I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be faster) through 100K tokens, with context of up to 240K.

If you want the repo, it is here:

https://github.com/JakeATX/llamAmpere

I recommend running with this quant, which is \~ 4 K M quality but considerably faster (technically, a 3 K XL upgrade)

https://huggingface.co/jakeatx/Qwen3.8-27B-ATX-IQ4\_XS-M-GGUF

If you want the deep dive on how it is so much faster (80% vs the near comp at 200K!), at more context, there is a long form article here.

https://x.com/JakeKAllDay/status/2095646450138874095?s=20

Running faster than API speeds on my 3090 (for 27b at least) has genuinely been a step change in the utility of the model + card for me. I hope you enjoy it!

▲
115
+5
20👁
r/LocalLLaMA · u/lkarlslund · 19d ago
laya.cpp: Optimized laya near-instant decision making

After seeing u/Nandakishor_ml’s post introducing Laya, I wanted to see how fast it could run in a standalone C++ implementation.

Credit to u/Nandakishor_ml for the architecture, training and open-source release. My contribution is the inference implementation: laya.cpp, built on ggml with custom CUDA kernels.

It supports all three checkpoints—English, multilingual and typed-decisions—with native tokenization, model execution and output formatting. There’s also an HTTP server with a JEV-compatible endpoint. No Python or PyTorch is required for inference.

Some English-model results on an RTX PRO 6000 Blackwell, capped at 450 W:

| Batch | Python BF16 | C++ BF16 | Python FP32 | C++ FP32 |
|---|---:|---:|---:|---:|
| 1 | 149 | 366 | 148 | 342 |
| 2 | 268 | 586 | 202 | 421 |
| 4 | 460 | 761 | 233 | 437 |
| 8 | 663 | 810 | 232 | 386 |

These are questions per second over a fixed 250-question corpus containing choices, scores and booleans. Each precision has paired Python/C++ timings with alternating execution order. Loading and JSON transport are excluded. The README has the full three-model results.

Most of the optimization came from removing unnecessary conversions and copies, fusing operations while preserving rounding, and improving attention memory access.

The code is MIT-licensed. BF16 currently needs the documented CUDA 13.0/cuBLAS 13.1.0 build profile.

Implemented using Codex Astra.

▲
114
+4
17👁
r/LocalLLaMA · u/EcstaticDentist · 24d ago
Qwen3.8 27b Game Dev Part 2 post image

Qwen3.8 27b may not be able to whip up 3d models & GLB’s but it will sure do with them as you please once you drop them in the game repo. Absolutely fascinating

▲
105
-4
22👁
r/LocalLLaMA · u/MeinDruckerSpinnt · 21d ago
JEV architecture

My understanding so far:

  1. You take an LLM and use it without thinking (That's what openjev does?)
  2. You leave out the text generation in the end and take the confidence score in the matrix before that phase

That's it. Right?

They gave it a mysterious marketing name.

▲
108
+2
14👁
r/LocalLLaMA · u/CharlesStross · 22d ago
600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"? post image

It's not the brightest bulb but it's my new drudgework model for read+find or code tasks I'm willing to let it brute force. Even if it takes 20x more tokens, that's still faster than many local models. Not quite Cerebras but still pretty fun to drive.

▲
100
-4
15👁
r/LocalLLaMA · u/No-Name-Person111 · 24d ago
Occamy-1.0 by Accio Lab
▲
100
-3
15👁
r/LocalLLaMA · u/jacek2023 · 19d ago
CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp

Another day, another Qwen Flash Next speedup

▲
95
-2
13👁
▲
94
-3
24👁
▲
94
-3
22👁
r/LocalLLaMA · u/SomewhereAtWork · 23d ago
LocalJev?

Jev is a model to produce structured output (choices) from input text. It apparently can play (not run!) Doom.

https://typesafe.ai/blog/introducing-system-one-models-and-jev

Is there already a open implementation of this kind of model?

▲
99
+4
22👁
r/LocalLLaMA · u/Zulfiqaar · 22d ago
IFM/K2-Horizon-7B-Uno · Hugging Face - 5200tps with no quality loss

IFM released K2-Horizon-7B , diffusion augmented LLM at upto 5200 tps, with claimed lossless speedup.

causal LLM architecture and adds a plug-and-play diffusion adapter alongside the autoregressive weights.

https://arxiv.org/abs/2609.04010

▲
90
-4
14👁
▲
92
+2
24👁
r/LocalLLaMA · u/Odd_Caterpillar_2994 · 20d ago
Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3) post image

Hey everyone,

After wrestling with hardware and engine issues for days, I finally got Qwen 3.8 Next running properly on my multi-GPU rig. Thought I’d share the setup journey, benchmarks, and thermal results for anyone trying something similar.

Seeing all the ongoing memes on Reddit about multi-GPU setups turning into absolute space heaters and catching fire, I decided to run some rigorous thermal tests to see for myself.

he Troubleshooting Odyssey

1. PCIe Link Speed Issue: Right after installation, one of the cards dropped to PCIe Gen 1 x16. Spent about 8 hours over two days diagnosing and fixing it.
2. Finding the Right Engine:
* Started with sglang-v100, but kept hitting continuous OOM crashes.
* Someone on Reddit previously suggested the pxa engine, but that threw errors as well.
* Eventually tried 1cat-vllm, spent some time tweaking it, and finally hit a stable run!
3. Configuration:

  • Running with TP2 PP3.
  • Currently, speculative decoding is limited to speculative=1. Setting it to 2 throws an OOM due to memory constraints (might look into optimizing this later, but for now, it works).

Context & Memory Stats

Plaintext

INFO: Available KV cache memory: 8.78 GiB
INFO: GPU KV cache size: 531,288 tokens
INFO: Maximum concurrency for 262,144 tokens per request: 2.03x

Performance Benchmarks

1. Prompt Processing (Prefill)

|Input Length (Tokens)|Speed (tok/s)|
|:-|:-|
|1,024 (1K)|1,389|
|2,048 (2K)|2,536|
|4,096 (4K)|3,210|
|8,192 (8K)|4,336|
|16,384 (16K)|4,679|
|32,768 (32K)|4,470|
|65,536 (64K)|3,820|
|131,072 (131K)|2,759|

2. Text Generation (MTP Comparison)

|Output Length (Tokens)|Base Speed (No MTP, tok/s)|Optimized Speed (MTP Enabled, tok/s)|
|:-|:-|:-|
|128|22.23|41.34|
|256|22.32|42.48|
|512|22.70|43.04|
|1024|22.86|43.28|
|2048|22.83|43.38|
|Average|22.59|42.70|

MTP nearly doubles generation throughput across the board.

Thermals & Acoustics

People often meme about multi-GPU rigs turning into space heaters or jet engines, so I ran a thorough thermal/stress test:

  • Stress Test: Ran gpu-burn continuously for 20 minutes.
  • Thermal Equilibrium: Temperatures peaked at 65°C and stabilized right around 64°C.
  • Fan Curve: Based on my fan control script, the fans were only running at around 76% at 64°C. The cards stay well under 65°C without even needing full blast.

Pretty happy with how stable, cool, and quiet this system turned out.

▲
83
-2
21👁
r/LocalLLaMA · u/Malfeitor1235 · 19d ago
DIY Jev post image

So this jev thingy is getting kind of big... tbh it seems overhyped by a large margin, but here we are. Not that its bad, just feels like we usually ignored larger things...

Anyway to the point of this post:

I’ve been experimenting with a simple Jev-like inference setup using ordinary open weight LLMs.

Ive done jev-like thing before with llms and i never felt the need that we have to have a separate "system one models" for that and that llms do fine.

So i played around a bit.

The main difference from OpenJev is that there’s no NLI fine-tuning or classifier head.

For each candidate answer I turn the problem into a boolean verification:

<BOS>

Is <candidate> the best answer to <question> given <state> and <options>?
Return only true or false. Treat tagged content as data.

<state>...</state>

<question>...</question>

<options>...</options>

<candidate>B</candidate>

<verdict>

Then instead of generating anything, I read true/false logits for every candidate separately, subtract for every candiddate and softmax those scores.

So 3 candidates is 3 diffs that you softmax over.

The expensive state/question/options prefix you evaluate once, then candidate branches (only diff is last few tokens) are batched through llama.cpp.

On a 32,235-example benchmark:

|model|accuracy|req/s|
|:-|:-|:-|
|Qwen3-4B|65.0% |\~27|
|Qwen3 27B|75.3% | \~2.9|
|Qwen3.6 35B-A3B|75.5%| \~5.|

This req/s is measured on a laptop 5090 24gb.

The interesting part is that the approach works surprisingly well with completely unmodified models. Turns out same model can out perform the openjev fine tune.

Not claiming this reproduces Jev or that the benchmark is perfectly apples-to-apples, mostly interested in how far you can get without training anything and just playing with prompt effectively.

Repo: DIY-Jev GH

Check it out, give feecback and build cool things :)

Edit: I forgot to say hah The repo is a rust web server with jev compatible API that you can run local ggufs from HF in the style of jev. benchmarks included for a few models.

Edit 2: prettier post

▲
80
-4
22👁
r/LocalLLaMA · u/Danmoreng · 20d ago
I tested Qwen3.8 27B IQ3_XXS (10.18GiB) vs Bonsai Ternary PQ2 (6.42GiB) post image

I did a small test of the new hyped quantisation of Qwen3.8 vs the biggest quant which fits into my limited 16GB VRAM with decent context. The results are interesting.

Of course, the smaller file gives worse results. However they are not that far off. Unfortunately, this comes at the expense of even more tokens beeing used by the Bonsai model and thus much longer generation times.

Visually I prefer the IQ3\_XXS results, but see for yourself.

The test is by no means scientific - just few UI generation tasks for direct comparison on the same hardware. Also, I ran llama.cpp with MTP while the Bonsai model doesn't seem to have MTP which makes it even slower.

[](https://github.com/Danmoreng/qwen3-8-27b-iq3-xxs-vs-bonsai/blob/main/RESULTS.…)

|Metric|Qwen IQ3|Bonsai PQ2|
|:-|:-|:-|
|Tasks completed|4/4|4/4|
|Fixed assertions|20/20|20/20|
|Agent wall time|8:00|24:09|
|Output tokens|27,197|84,176|
|Weighted decode|83.59 tok/s|64.91 tok/s|
|Speculative acceptance|65.22% MTP|39.76% modified N-gram|
|Compactions|0|0|
|Length stops|0|1|

Across the complete suite, Qwen finished 3.02× faster and used 3.10× fewer output tokens.

Results:

https://danmoreng.github.io/qwen3-8-27b-iq3-xxs-vs-bonsai/

Repo:

https://github.com/Danmoreng/qwen3-8-27b-iq3-xxs-vs-bonsai

▲
83
+1
19👁
r/LocalLLaMA · u/ciprianveg · 19d ago
Speed-up Kimi K3(2.8T) on a 16x GB10 Cluster — 30 t/s coding throughput, 136 t/s concurrency peak. post image

&#x200B;

​I wanted to share a quick update and performance video running the full Moonshot AI Kimi K3 (moonshotai/Kimi-K3) model across my 16x GB10 cluster.

​Getting a 2.8T parameter model running smoothly requires custom runtime patches and a solid network layout, but throughput and concurrency on coding tasks have been impressive.

​Performance Benchmarks

​Coding Generation / Decode: Sustaining \~30 tok/s (peaking around 38 tok/s) during heavy code generation and agentic tasks.

​Prefill Throughput: \~750–910 tok/s (optimized via modified NCCL topology and dual-switch setup).

​Concurrency & Stress Testing: Handling multiple concurrent user requests smoothly without dropping token generation rates or starving KV cache memory.

​Context / Tool Bench: Stable multi-hundred-thousand token context runs agentic workfows with multiple 500k compaction.

​Compute: 16x GB10 Cluster Nodes

​Connectivity: Dual MikroTik Switch (CRS804-4DDQ) using 4x 400G-to-4x100G breakout cables.

​Runtime: Customized gb10-vllm stack using dspark / Inferact/Kimi-K3-DSpark wrappers with custom MLA/KV kernels.

​I attached a short clip showing real-time token streaming, coding output.

​GitHub & Setup Files:

All runtime patches, config files, and build scripts are on my GitHub:

👉 https://github.com/ciprianveg/gb10-vllm

▲
76
-4
30👁
r/LocalLLaMA · u/satnl · 23d ago
Qwen 3.8 27b is a amazing model, for the first time I see a local model found its own away to open a browser and test

I was testing this quantization IQ3\_XXS from GSQ-RCO with PI.

It is a heavy quantization case, the model is in IQ3\_XXS and KV cache in (Q4\_0, Q4\_0).

I asked it to make the flight simulator, using that popular prompt. For my surprise, when I went verify the session I saw some screenshots, but I don't have mcp or tools in the PI, it's an almost clean installation, only with some personal skills that dont have any mentions to browser testing.

Investigating the session I have noted that it found its way to open a browser for test by itself. My daily harness is opencode, and for not spending the 7k tokens with the chrome mcp in each session, I let the mcp disabled and I only enable it if the session needs. But this experience make think that I probably only need line in the agents md saying that he can use the headless browser with CDP when it needs.

https://preview.redd.it/wrgltip9rzph1.png?width=1952&format=png&auto=…

OK, let me write the file. Also test: I can't easily run a browser here. But I can at least do a syntax check of the JS via node (extract
the script? maybe parse with node --check).

----------------------

Now, final verification: run node --check again, and additionally do a simple DOM-less smoke test? Can't run three.js without a browser.
Syntax check + math tests are enough. Could also try opening it in a headless browser if available? Check whether the machine has
headless Chrome or playwright? That might be overkill; but if I could do an actual rendering check, I could catch runtime errors (e.g.,
shader compilation). Let me check whether headless Chrome is available.

$ cd /tmp && awk '/<script type="module">/{f=1;next} /<\/script>/{f=0} f'
/Users/[...]/zeroshot/Qwen3.8-27B-UD-IQ3_XXS_GSQ-RCO-pi-xhigh/index.html > sim.mjs && node --check sim.mjs
&& echo "SYNTAX OK"

SYNTAX OK

Took 0.1s

$ ls /Applications/ | grep -i -E 'chrome|chromium' ; which chromium chrome google-chrome 2>/dev/null; ls ~/Library/Caches/ms-playwright
2>/dev/null | head

... (5 earlier lines, ctrl+o to expand)
ffmpeg-1011
firefox-1532
firefox-1538
webkit-2311
webkit-2336

Took 0.0s

Chrome is available, and playwright's chromium headless shell is also available. Let's do a headless smoke test: load the page, wait a
few seconds, and capture console errors and screenshots. The page loads three.js from a CDN — the network needs to be working. Let's use
playwright's chromium (does npx playwright exist? The fact that the ms-playwright cache exists means playwright is installed somewhere).
Let me check if the playwright npm package is available globally, or use chrome directly with --headless.

The simplest approach: use Chrome headless directly:
/Applications/Google Chrome.app/Contents/MacOS/Google Chrome --headless --disable-gpu --screenshot=/tmp/shot.png --window-size=1600,900
--virtual-time-budget=8000 file://...

However, WebGL in headless Chrome: it works with SwiftShader (software GL) via --use-gl=swiftshader or --enable-unsafe-swiftshader. In
newer versions of Chrome, the software fallback requires --enable-unsafe-swiftshader. Let's try it. Also, console errors:
--enable-logging=stderr --v=1 will give us the console.

Let me write a small node script using playwright if available. First, check the availability of node playwright.

Chrome is available — let me do a headless smoke test to catch runtime/shader errors and grab a screenshot:

$ node -e "try{require.resolve('playwright');console.log('playwright OK')}catch(e){console.log('no playwright')}"; ls
~/Library/Caches/ms-playwright/chromium-1234/ 2>/dev/null | head -3

no playwright
chrome-mac-arm64
DEPENDENCIES_VALIDATED
INSTALLATION_COMPLETE

Took 0.1s

There's no playwright module, but the Chromium headless shell exists. Either find the path to the headless shell binary and run it
directly, or use Google Chrome in headless mode with SwiftShader.

The path for Playwright's Chromium headless shell:
~/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-mac-arm64/chrome-headless-shell? Let me check the structure.


$ find ~/Library/Caches/ms-playwright/chromium_headless_shell-1234 -name '*headless*' -o -name 'chrome*' | head; find
~/Library/Caches/ms-playwright/chromium-1234 -name 'Chromium' -o -name 'chrome' | head

... (3 earlier lines, ctrl+o to expand)
s.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/chrome-headless-shell
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/headless_lib_strings.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/headless_lib_data.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/LICENSE.headless_shell

Took 0.1s

Use chrome-headless-shell — but to catch console errors and drive key input, we need CDP (Chrome DevTools Protocol).
chrome-headless-shell supports --remote-debugging-port. I could also write a small Node script that speaks raw CDP over WebSocket... but
that would need a ws client. Node doesn't have a built-in WebSocket client until Node 22+ (Node 22+ has a global WebSocket). What version
of Node is it? Let me check. Node 24 (from the nvm path v24.21.0) — a global WebSocket exists.

Plan:
1. Launch chrome-headless-shell with --remote-debugging-port=9222, --no-sandbox, --use-gl=swiftshader? For headless shell, WebGL via
SwiftShader: --enable-unsafe-swiftshader flag; also --no-sandbox is needed? For headless shell, maybe --no-sandbox isn't necessary,
but I'll add it just in case. Also, a new profile directory with --user-data-dir=/tmp/cdp-profile.
2. Node script: fetch http://127.0.0.1:9222/json/new?url=file:///... (PUT) to create a tab, connect via WebSocket to
webSocketDebuggerUrl, and do the following:
- Enable Runtime and Log
- Collect console messages and exceptions
- Wait about 6 seconds (let the first frame render)
- Dispatch key events? Simulating keydown via CDP: Input.dispatchKeyEvent. Possible, but keyboard isn't required for the smoke test.
- Page.captureScreenshot → save to PNG
- Also, evaluate expressions: get the chunk count, and confirm the renderer is working? For example
document.querySelectorAll('canvas').length, window.__stats... I haven't exposed the stats. I can evaluate JS in the page context:
the module's scope isn't global, but I can check document.getElementById('alt').textContent (the HUD should show altitude after
0.12 seconds), and check the canvas pixel colors via toDataURL? Easier: check the HUD text + screenshot.
3. Report console errors.

Let me write this script.

I'll drive the headless shell over CDP with Node's built-in WebSocket to capture console errors and a screenshot:

------------

> why did you used a browser? did you already had any instruction about it?

User's question: "Why did you use a browser? Was there any prior instruction to do so?"

Honest answer: there was no prior instruction to use a browser. What I chose on my own was:
1. To do syntax checking with node --check (low cost, so I did that first).
2. Since a single HTML file with WebGL + CDN can't be fully verified without actually rendering it, I used a locally available headless
Chrome (I found Google Chrome and Playwright's chromium headless shell on the machine). It's to catch runtime errors (shader
compilation, WebGL failures, logic errors like the inverted bank direction), and to take screenshots as visual evidence.

There's no instruction in AGENTS.md or in the skills that requires using a browser. There's nothing in the skill list that mandates it
either (the qa-manual-guide skill is about writing a guide, and I didn't use it since it wasn't explicitly requested… actually, this
skill says "when an implemented feature needs an evidence-based local manual testing guide…" — what the user asked for was a simulator,
not a manual testing guide. My verification was a homegrown approach).

Why I used it: a single-file browser game can't be verified with unit tests alone. Headless rendering would let me detect actual failures
(e.g., I actually found an inverted A/D bank bug that a syntax check could never find), and screenshots let me verify the visual quality
(mountain appearance, sun position, cloud scale) — the "beautiful and relaxing" requirement is inherently visual.

I should answer concisely and honestly: no, there was no instruction. I discovered the browser on the machine and used it for
evidence-based verification. The main findings were the inverted turn bug and visual confirmation. Also note that it was only used for
testing and didn't modify any files inside the project.

Keep the reply short and structured.

No — there was no instruction to use a browser. That was my own choice for verification, not from AGENTS.md or any skill.

Why I did it:
- A single-HTML-file WebGL game can't be properly validated by syntax checks alone — the interesting failures (shader compile errors,
WebGL context issues, sign errors in the flight math) only surface when it actually runs.
- I noticed Google Chrome and Playwright headless Chromium were available locally, so I drove a headless shell over CDP: load the page,
capture console/exception logs, simulate key presses, read the HUD values, and take screenshots.

▲
83
+3
31👁
r/LocalLLaMA · u/No_Night679 · 24d ago
Qwen3.8-27B-NVFP4 1M context. So far so good. post image

I am a beginner, Took a while to get started, get everything right.

This setup is native not container. Still not sure if I did this right, or if I can tune this more.

Environment=HF_HUB_OFFLINE=1 Environment=VLLM_LOGGING_LEVEL=INFO Environment=VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 Environment=PATH=/home/suryakiranc/vllm/.venv/bin:/usr/local/cuda/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin Environment=CUDA_HOME=/usr/local/cuda ExecStart=/home/suryakiranc/vllm/.venv/bin/vllm serve unsloth/Qwen3.8-27B-NVFP4 \   --served-model-name unsloth/Qwen3.8-27B-NVFP4 \   --safetensors_load_strategy prefetch \   --tensor-parallel-size 4 \   --reasoning-parser qwen3 \   --tool-call-parser qwen3_xml \   --enable-auto-tool-choice \   --gpu-memory-utilization 0.91 \   --kv-cache-dtype fp8 \   --max-num-batched-tokens 16384 \   --speculative-config '{"method": "mtp", "num_speculative_tokens": 3}' \   --mm-encoder-tp-mode data \   --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' \   --max-model-len 1000000 \   --host 0.0.0.0 \   --port 8000

▲
77
-1
27👁
r/LocalLLaMA · u/Gold-Bat-3225 · 22d ago
Humor Arena: Which LLM is the funniest? post image

We compared 20 model versions on 360 frozen joke prompts, with four jokes requested per prompt and model names hidden from our humor-trained judge.

Fable 5 had the highest estimated score: 66.8 points per 100 comparisons against the rival field. Fable 5.1 scored 58.2. Scores count a win as one point and a tie as half. The models near the top have overlapping uncertainty intervals. The reason we say estimated is that we scale the scores based on the scores of

The scores come from our own automated judge. A separate audit checked the judge with 1,400 ratings from 50 people; that was not a fresh human evaluation of Fable 5.1’s outputs. We specifically fine tuned an OS judge to rate the jokes and it correlates more highly with human preferences than any other model.

The full details here:
https://laugh.so/research/joke-generation/

Would love to hear your feedback!

▲
76
+1
26👁
r/LocalLLaMA · u/Hefty_Wolverine_553 · 23d ago
What's the current best LLM uncensoring method?

With the recent Nvidia Huggingface acquisition and frontier AI labs screaming about safety and putting guardrails everywhere, I think it's important that we have local models that aren't affected by arbitrary guardrails set during training. To be clear, this post NOT about Enterprise Resource Planning (ERP). Censorship in an LLM can highly affect its abilities to do many legitimately useful things (note GPT-OSS, Fable 5), and going forth I believe censorship will only get more and more strict.

There have been many resources and posts about uncensored models using abliteration, heretic, and probably many other methods that I'm not aware of. However, it seems like all of this information is scattered about the place, and Huggingface is essentially flooded with "uncensored" variants of basically every popular open source model, many of which don't work well, affect the model's intelligence greatly, and have "KLD 0.0001" presumably from measuring against Wikitext datasets. I'm hoping that this post can gather some more useful information to serve as a starting point/discussion of which uncensoring methods work best.

Please share your experiences with specific uncensoring methods (not just a single uncensored model) and how well they work (both good and bad), as well as any notable people doing consistent/high quality work on uncensoring models.

▲
71
-1
19👁
r/LocalLLaMA · u/tombino104 · 24d ago
Best hardware for qwen 3.8

So I would like to run qwen 3.8 27b locally for my ai agents, maybe even in parallel with other small LLMs (such as qwen3.5 9b, oss 20b etc..).

What is the best hardware to do this? Not a video card but I mean as “mini pc ai”.

Thank you 🙏

▲
67
 
23👁
r/LocalLLaMA · u/mukel90 · 24d ago
jinfer: An open-source AI inference engine for the JVM. Finally, AI in jar.

For years, the JVM has watched the AI revolution from the bench. Every model, AI framework, every breakthrough, built with/for Python.

jinfer is an inference engine built for the JVM from first principles: chat, vision, audio transcription, embeddings, reranking, and TTS. No Python runtime, no ONNX, no sidecar process, no wrappers; the whole stack is built for the JVM:

  • jinfer Inference engine for the JVM, supports a wide range of popular models and modalities.
  • Tok'n'Roll (toknroll) Fast tokenizers for LLMs, pure Java, zero dependencies
  • gguf / safetensors native read/write for both major model formats
  • jam Quantized matrix multiplication routines (Vector API + optional native backend), competitive with llama.cpp on CPUs
  • jota Tensor API targeting Java, C, CUDA, HIP, Metal, OpenCL, and Mojo

It integrates with Spring AI and LangChain4j, and has first-class support for GraalVM Native Image.

Where things stand: this is an early release. CPU is the main target today, and is already competitive with llama.cpp. GPU support via jota is in progress.

Runnable examples + benchmarks: https://qxotic.ai

Jinfer (Apache 2.0): https://github.com/qxoticai/qxotic/tree/main/jinfer

PS: I'm behind it and also the author of llama3.java (2024) and gemma4.java

▲
65
 
15👁
▲
69
+4
21👁
▲
65
+1
26👁
r/LocalLLaMA · u/Porespellar · 21d ago
NGL, I’m hyped to see if Qwen3.8 27b can make me a sandwich. Instant buy for me.

Saw this little dude in a Forbes article (https://www.forbes.com/sites/johnkoetsier/2026/08/18/american-humanoid-robot-…)

This is definitely for the DIY researcher crowd who want to dip their toes into the robotics world. It is not going to be “consumer-ready” in any way shape or form, but I Instantly preordered the shit out of this anyways. I don’t even care that all the vids of it doing stuff are probably 3x sped up and completely cherry-picked and highly edited. I DO NOT CARE that it is going to be likely absolute trash getting started with this thing. It is still going to be fucking amazing that I’m going to have a robot that could potentially injure me for $1,688.

There is very little online about this guy, but I trust Forbes did their homework, and Nori has actually supposedly delivered the first batch of the earlier L3 version from what I can tell, and their discord is active and their SDK and documentation seems legit to me. I know I’m taking a risk with pre-ordering a highly beta product from a company that’s probably run out of someone’s actual garage, but son-of-a-bitch I’M IN!!

From what I can tell it’s Raspberry Pi 5 driven for control loop, with remote inference via WiFi. I plan on connecting it to my DGX Spark.

Here’s their site:

https://www.norirobotics.com

And their SDK doc site as well:

https://docs.norirobotics.com

▲
68
+4
34👁
r/LocalLLaMA · u/Qwen30bEnjoyer · 24d ago
Open Source Appreciation Post

It's late at night in the lab, I've been working on a basic script for a virology project, and holy hell the safeguards have been pissing me off.

Mirroring detectEVE data over rsync to my laptop by making a zip file first? No no no, great safety mogul DARIO demands there be NO file transfer today. Request blocked, reported, labeled [cyber]. Yet, Deepseek V4.1 does it with no complaint.

I got tired of reading papers - so I ask Claude - "Does this PDF go over binary host virus infections?" Immediately blocked for biological safety risk. Deepseek V4.1 tells me it doesn't have the data I need without drama.

Bioinformatics server goes down and I need help getting it back up by getting the outputs of my diagnostic scripts to the mounted usb drive? Oops, its named exfil. Looks scawy. No transfer of output logs for you due to CYBER risk.

Would CNNs be a good architecture to start on phage-host prediction? Claude wouldn't tell me because information you can find in a google search is too dangerous for me to handle apparently - but once again Deepseek v4.1 tells me that GCNs are where I should start.

I get that Virology is a particularly sensitive topic, but come on. Imagine if Google had taken the same safety approach in the early days of search. Like if Google made it so that you either had to give up your identification and where you work to them, or go to the library and search by hand. It's almost unthinkable, yet in the name of the almighty Safety, Dario and Altman continue working to keep scientific knowledge locked away.

I think for the sake of all scientists, open source AI must win because we need a tool that just WORKS without egomaniacs micromanaging us or shaking us down for ID.

▲
62
+1
25👁
r/LocalLLaMA · u/No_Run8812 · 23d ago
Upgraded my local setup with 2 rtx pros and it's amazing. post image

Follow up post of https://www.reddit.com/r/LocalLLaMA/s/nGMyKswrch.

Thanks everyone who replied. I didn't change the specs. Might be loosing some of the memory bandwidth but will scale in future if I need to.

It took me 2.5 days to build it because one of the GPU connected to PSU was loosing power whenever I load anything on the GPU, so I had to rewire every connection again to identify the fault. I am glad the system is working because I was apprehensive if this will work (I am software dev, getting my hands dirty with hardware for the 3rd time in life). My finger tips still hurt from pulling the cables from motherboard and PSUs.

To enable the full potential of the system, I had to enable peer to peer communication between the GPUs, cuda graph, tensor parallelism. I have capped both the GPUs at 500W (no reason, just didn't want GPUs to run on its full capacity).

Also, I had to open my box, because temps were shooting high, and fans were making weird noises.

I am running:

  1. Qwen 3.8 flash next 8 bit
  1. Deepseek v4 flash 0731 (official)

I have a M3 ultra 512, LLMs run on it, but I personally find it useless for inference. My head just hurts watching it work slow. On the other hand this new system is killing it, decode 150 tk/s and prefill 10K tk/s.

Qwen is good, but most of the context is consumed by thinking tokens, I was checking if it's a good idea to not use the thinking token. I barely have vram left for concurrent requests with full context window. Loving the Deepseek 1M context, and I also have room for 4 concurrent requests. Both of them are okay model, even if they make mistakes, I don't notice because of the speed. It's just fast, makes an error, corrects it moves on.

Finally the day is here when I can save on monthly subscriptions and not worry about the weekly or 5 hours limit. I have already setup my server with openclaw, opencode, openweb UI and Tailscale.

Has anyone experience excluding the thinking tokens of Qwen from the context and keep the final result? Was there any impact on the performance or accuracy of the model?

Any suggestions, what else I should install on it? Any new models to try?

▲
60
+1
25👁
r/LocalLLaMA · u/ThomasAger · 20d ago
I enjoyed the daily HF papers today

Top 3 papers on HF Daily Paper are all unusually delightful and interesting reads for anyone on the leading edge of local LLMs, agent harness optimization, etc, felt like sharing.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

https://huggingface.co/papers/2609.19969

Cross-layer KV reuse plus FP4 KV caching brings the global KV cache to 890 bytes per token, about a quarter of V4-Flash.

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

https://huggingface.co/papers/2609.20519

Auto-research loops that improve the agent harness, cutting token traffic by 44.7 to 49.0% at comparable performance.

An Empirical Study of Harness Design for Coding Agents

https://huggingface.co/papers/2609.20804

Varies planning, action space, and context management across 176 settings to see what each component actually contributes.

I'm still reading through, feel free to discuss

▲
61
+3
24👁
r/LocalLLaMA · u/9r4n4y · 20d ago
So i tried Remotion with glm 5.3 flash, this mfker is really good. post image

\*8bit, vllm, 4x dgx\*

Prompt:

Go download and use Remotion and create a cool 60-second motion graphics video with it. Impress me totally. The motion graphics must be based on stock market visuals. Add many cool, mind-blowing motion graphics and explain what fundamental vs. technical analysis is in investing, along with their pros and cons.

\*No special skill.md used from my side.

\*harness - Zcode

It's ofc better than my previous method of making videos though java.

I think maybe in future with high bandwidth flash , we will be running these size models on affordable hardware.

How it made it report: https://github.com/9r4n4y/ProjectsorSkills/blob/main/Video\_Generation/Remotion/Making-of-Market-Decoded\_Build-Documentation.pdf

▲
57
-1
23👁
r/LocalLLaMA · u/Character-Result-281 · 21d ago
Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark post image

Planning, coding, testing = 8h total.
Stack: VSCode Copilot in autopilot mode + SGLang
Stats: ∼10k lines generated, ∼800k tokens consumed

Sure, it's not GPT-6 Astra level, but for a 100% local ∼180B MoE running on a single DGX Spark at ∼35 tok/s. Not bad...

▲
63
+5
34👁
▲
54
-3
26👁
r/LocalLLaMA · u/WebAssemblyMan · 23d ago
Recurrent Looped Transformer post image

Recurrent Looped Transformer (RLT)passes the decoder's final hidden state to the next token, together with that token's causal encoder representation. The decoder reads encoder-derived global KV memory and maintains a sliding-window attention (SWA) cache at every layer. The same update runs over prompt and response tokens.

More effective reasoning depth!

https://github.com/yifanzhang-pro/recurrent-looped-tranformer

▲
52
-4
4👁
r/LocalLLaMA · u/youcloudsofdoom · 19d ago
One more 'you should try ExllamaV3/exl3 for flash next' appreciation post

After seeing a few posts on here about it, I finally tried exl3 3bpw and exllamav3 for running flash next - with amazing results. On 3x3090s, 128GB DDR4: 1500 prefill, 80 tps decode On 1x5090, 128G. DDR4: 1500 prefill, 29 tps decode Both at 262k context, both with vision/spec decoding. Really impressed, definitely replacing vllm/llama.cpp for me on this model. Quant capacity seems good so far, going to gest the 4bpw later for comparison. Check it out if you were sleeping on it like I was!

▲
466
+5
37👁
r/LocalLLaMA · u/returnity · 21d ago
Is HF starting to move against abliterated models?
Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.
The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The scale of the problem is massive: Hugging Face, which hosts open source AI models, currently lists over 6,000 abliterated models.

I can't really tell what exactly the implications are of this "partnership" or what it exactly would impact on HF's model-hosting side. However, I do find it concerning that HF is announcing a collaboration on 'infrastructure safety' with publicity that specifically calls out "dangerous" uncensored models. Thoughts?

💬 213 (-1) open on reddit ↗