104 posts · 1 sub · RSS
← prev Aug 12, 2026 → Sep 9, 2026 next →
2026-08-12 → 2026-09-09 hourdayweekmonthyearall
allr/LocalLLaMA
▲
615
+8
36👁
r/LocalLLaMA · u/Porespellar · 30d ago
Why the hell is LM Studio making LM Studio so difficult to download? post image

Who is the marketing genius at LM Studio that decided that going ALL IN on pushing their new Bionic Agent product meant they are going to make it a giant pain in the ass to find and download actual LM Studio.

This is the dumbest marketing decision I’ve ever seen. I used to love LM Studio, it was the middle stepping stone in the logical progression of inference. Most OGs here likely started with Ollama, moved to LM Studio, on their way to vLLM. Now trying to go to LM Studio takes you to Bionic. I mean, you can eventually find LM Studio but they make it not super easy.

Here’s a thought LM Studio, maybe stop redirecting me to something I don’t want to download when I’m trying to find your actual namesake product. I’m glad you’re excited about the future of Agents and whatnot, but you’re absolutely ruining any goodwill I have for your products by trying to force feed me Bionic. Stahhhhhp!

💬 222 (+1) open on reddit ↗
▲
2682
+6
22👁
▲
402
+5
20👁
r/LocalLLaMA · u/fugogugo · 32d ago
when will open source LLM catch up to Astra I wonder? post image

I feel like this year has been insane , the speed of AI race is something that normal human can't catch up anymore

▲
334
+5
15👁
r/LocalLLaMA · u/Storterald · 35d ago
I benchmarked 21 Qwen3.8 27B variants on 16GB VRAM

After Qwen3.8 27B came out, I decided to benchmark the models that could fit in my GPU (RTX 5080) on my actual code (C code), the results were not completely unexpected but some quants were definitely underwhelming.

TLDR: Best overall: bartowski/Qwen3.8-27B-IQ4_XS. Best uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4_XS. For a bit more context: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S or uncensored: huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp

  • edit1: added TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS, bartowski/Qwen3.8-27B-IQ3_XS, prism-ml/Ternary-Bonsai-27B-Q2_g64 and magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4_K_S-Unsloth
  • edit2: added unsloth/Qwen3.8-27B-UD-IQ3_S and ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3_S
  • edit3: added magiccodingman/Qwen3.8-27B-MQ-IQ2_M_1 and huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3_S
  • edit4: added AtomicChat/Qwen3.8-27B-AD-IQ4_XS-IQ3_S
  • edit5: added IvanKrastevAdventics/qwen3.8-27b-awq-int4-q4_0 (gguf of cyankiwi/qwen3.8-27b-awq-int4)
  • edit6: added the updated bartowski/Qwen3.8-27B-IQ4_XS and bartowski/Qwen3.8-27B-IQ3_XXS
  • edit7: added bartowski/Qwen3.8-27B-Q3_K_M, Thireus/09ae8ba_22b6bb2 and Thireus/09ae8ba_248b31b
  • edit8: removed all MTP heads from the GGUF size for a more fair comparison. added bartowski/Qwen3.8-27B-IQ2_S
  • edit9: added huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3_S-mtp
  • edit10: added mradermacher/Eintopf-Qwen3.8-27B.i1-IQ3_M and hitsfmdj/Qwen3.8-27B-4.2BPW-16GB
  • edit11: added Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3_S-recovered
  • edit12: added Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw and turboderp/SC_3.00bpw_H4_V4
  • edi13: added prism-ml/Ternary-Bonsai-2-27B-PQ2_0 and prism-ml/Ternary-Bonsai-2-27B-PTQ1_0, replaced prism-ml/Ternary-Bonsai-27B-Q2_g64 with prism-ml/Ternary-Bonsai-27B-PQ2_0
  • edit14: added Bucoid/Qwen3.8-27B-Heretic-Ara-iq4_xs-3.0
  • edit15: added agentionai/Qwen3.8-27B-AP-IQ3_S, agentionai/Qwen3.8-27B-AP-IQ4_XS and RonnieOps/Qwen3.8-27B-IQ4_XS-fullvocab-E3. This is probably the final edit as Qwen4 27B will likely come out soon.
  • edit16: added byteshape/Qwen3.8-27B-IQ4_XS-3.84bpw and byteshape/Qwen3.8-27B-IQ3_S-3.23bpw. Again, probably the last edit.

(sorted by Mean KLD)

|Model|Mean KLD|Same top p|GGUF size (without MTP)|
|:-|:-|:-|:-|
|prism-ml/Ternary-Bonsai-27B-PQ2\_0|1.289582 ± 0.008684|82.849 ± 0.118 %|6.7GiB|
|prism-ml/Ternary-Bonsai-2-27B-PQ2\_0|1.096134 ± 0.007705|84.596 ± 0.113 %|6.7GiB|
|prism-ml/Ternary-Bonsai-2-27B-PTQ1\_0|1.095914 ± 0.007703|84.582 ± 0.113 %|5.5GiB|
|sdkyuan/qwen38-27b-qat-q2\_0|0.893177 ± 0.006948|85.727 ± 0.110 %|8.2GiB|
|bartowski/Qwen3.8-27B-IQ2\_Sbartowski/Qwen3.8-27B-IQ2\_S (NEW)|0.784060 ± 0.006457|87.016 ± 0.105 %|8.7GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2\_XS|0.767174 ± 0.006291|86.166 ± 0.108 %|7.8GiB|
|TheWegemann/Qwen3.8-27B-LowGPU-uncensored-NoMTP-IQ3XXXS|0.514311 ± 0.004864|89.023 ± 0.098 %|8.9GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ2\_S|0.512614 ± 0.004909|88.802 ± 0.099 %|8.6GiB|
|empero-ai/Qwen3.8-27B-Ridge-3.7bpw|0.475767 ± 0.004483|89.612 ± 0.096 %|11.4GiB|
|magiccodingman/Qwen3.8-27B-Quark-MXFP4-UD-Q4\_K\_S-Unsloth|0.419585 ± 0.004076|89.661 ± 0.095 %|13.1GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3\_XXS|0.379222 ± 0.003992|90.270 ± 0.093 %|9.4GiB|
|unsloth/Qwen3.8-27B-UD-Q2\_K\_XL (UD2)|0.350861 ± 0.003745|90.626 ± 0.091 %|9.6GiB|
|byteshape/Qwen3.8-27B-IQ3\_S-3.23bpw|0.345563 ± 0.003703|90.868 ± 0.090 %|10.1GiB|
|mradermacher/Eintopf-Qwen3.8-27B.i1-IQ3\_M|0.318143 ± 0.003139|91.471 ± 0.087 %|11.7GiB|
|bartowski/Qwen3.8-27B-IQ3\_XXS (NEW)|0.300480 ± 0.003345|91.511 ± 0.087 %|11.3GiB|
|unsloth/Qwen3.8-27B-UD-IQ3\_XXS (UD2)|0.268594 ± 0.002971|91.951 ± 0.085 %|10.8GiB|
|magiccodingman/Qwen3.8-27B-MQ-IQ2\_M\_1|0.256808 ± 0.002861|92.056 ± 0.085 %|10.9GiB|
|DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-NEO-MTP-IQ3\_M|0.251270 ± 0.002702|92.315 ± 0.083 %|13.1GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ3\_S|0.249650 ± 0.002841|92.016 ± 0.085 %|10.8GiB|
|bartowski/Qwen3.8-27B-IQ3\_XS (OLD)|0.238656 ± 0.002627|92.312 ± 0.083 %|12.2GiB|
|hitsfmdj/Qwen3.8-27B-4.2BPW-16GB|0.222090 ± 0.002570|92.552 ± 0.082 %|11.7GiB|
|esatapedico/Qwen3.8-27B-NVFP4-MTP-LOW|0.220796 ± 0.002631|92.339 ± 0.083 %|14.3GiB|
|unsloth/Qwen3.8-27B-UD-IQ3\_S (UD3)|0.218522 ± 0.002591|92.399 ± 0.083 %|10.9GiB|
|turboderp/SC\_3.00bpw\_H4\_V4 (exllama3)|0.205712 ± 0.002509|92.573 ± 0.082 %|11.9GiB|
|jrell/Qwen3.8-27B-i1-IQ4\_XS-GGUF-Smaller|0.194459 ± 0.002242|93.049 ± 0.080 %|12.4GiB|
|agentionai/Qwen3.8-27B-AP-IQ3\_S|0.193041 ± 0.002352|92.995 ± 0.080 %|10.9GiB|
|orcarouter/Qwen3.8-27B-Uncensored-Q3\_K\_L|0.192312 ± 0.002294|92.726 ± 0.081 %|13.4GiB|
|bartowski/Qwen3.8-27B-Q3\_K\_M (NEW)|0.191103 ± 0.002369|92.823 ± 0.081 %|12.3GiB|
|mudler/Qwen3.8-27B-APEX-I-Mini|0.190209 ± 0.002354|93.012 ± 0.080 %|12.6GiB|
|Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3\_S-recovered|0.178882 ± 0.002102|93.110 ± 0.079 %|11.0GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-GSQ-RCO-IQ3\_S-mtp|0.178715 ± 0.002200|92.949 ± 0.080 %|11.0GiB|
|Thireus/09ae8ba\_22b6bb2 (ikllama.cpp quality 41.39%)|0.178290 ± 0.002202|93.115 ± 0.079 %|11.0GiB|
|ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-IQ3\_S|0.175223 ± 0.002129|93.024 ± 0.080 %|11.0GiB|
|Mia-AiLab/Qwen3.8-27B-EXL3-3.5bpw (exllama3)|0.149827 ± 0.001936|93.735 ± 0.076 %|14.1GiB|
|unsloth/Qwen3.8-27B-UD-Q3\_K\_XL (UD2)|0.147186 ± 0.001809|93.734 ± 0.076 %|12.2GiB|
|unsloth/Qwen3.8-27B-UD-Q3\_K\_XL (UD3)|0.142647 ± 0.001860|93.789 ± 0.076 %|11.9GiB|
|byteshape/Qwen3.8-27B-IQ4\_XS-3.84bpw|0.131976 ± 0.001690|93.852 ± 0.075 %|12.0GiB|
|IvanKrastevAdventics/qwen3.8-27b-awq-int4-q4\_0|0.112990 ± 0.001558|94.171 ± 0.073 %|14.4GiB|
|AtomicChat/Qwen3.8-27B-AD-IQ4\_XS-IQ3\_S|0.111713 ± 0.001492|94.527 ± 0.071 %|13.2GiB|
|Bucoid/Qwen3.8-27B-Heretic-Ara-iq4\_xs-3.0|0.097572 ± 0.001341|94.651 ± 0.070 %|13.0GiB|
|Bucoid/Qwen3.8-27B-Uncensored-IQ4\_XS\_4BPW|0.091447 ± 0.001261|94.774 ± 0.070 %|12.8GiB|
|huihui-ai/Huihui-Qwen3.8-27B-abliterated-UD-IQ4\_XS|0.082871 ± 0.001205|94.981 ± 0.068 %|13.1GiB|
|unsloth/Qwen3.8-27B-UD-IQ4\_XS (UD3)|0.075626 ± 0.001097|95.258 ± 0.067 %|13.3GiB|
|agentionai/Qwen3.8-27B-AP-IQ4\_XS|0.073386 ± 0.001075|95.386 ± 0.066 %|13.0GiB|
|Thireus/09ae8ba\_248b31b (llama.cpp 49.75%)|0.063904 ± 0.000967|95.687 ± 0.064 %|13.3GiB|
|jpetrina/Qwen3.8-27B-IQ4\_XS-pure|0.061984 ± 0.000917|95.551 ± 0.065 %|13.3GiB|
|RonnieOps/Qwen3.8-27B-IQ4\_XS-fullvocab-E3|0.058278 ± 0.000886|95.773 ± 0.063 %|14.1GiB|
|bartowski/Qwen3.8-27B-IQ4\_XS (OLD)|0.056482 ± 0.000856|95.835 ± 0.063 %|14.3GiB|
|bartowski/Qwen3.8-27B-IQ4\_XS (NEW)|0.055415 ± 0.000849|95.850 ± 0.062 %|14.2GiB|
|unsloth/Qwen3.8-27B-UD-Q4\_K\_XL (UD3) *(can't fit)*|0.029844 ± 0.000476|96.921 ± 0.054 %|16.1GiB|
|unsloth/Qwen3.8-27B-UD-Q4\_K\_XL (UD2) *(can't fit)*|0.028026 ± 0.000432|96.988 ± 0.054 %|16.4GiB|

https://preview.redd.it/g9isjm0d04sh1.png?width=5355&format=png&auto=…

Hope this helps other VRAM starved people like me :)

▲
2256
+4
17👁
▲
1490
+4
32👁
r/LocalLLaMA · u/bakawolf123 · 31d ago
OpenAI alleged of stealing mathematicians work

Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex. A few days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question.

Full statement from them https://cims.nyu.edu/\~tristanb/statement.pdf

Feels like big labs believe everything you did with the help of their models is theirs.

💬 274 (-1) open on reddit ↗
▲
612
+4
10👁
▲
434
+4
20👁
r/LocalLLaMA · u/swagonflyyyy · 34d ago
Qwen3.8-27B beat the Wikipedia game in 6 clicks. post image

Used qwen3.8-27b in Opencode to make this silly mini-game because I'm not sober:

```
We are going to play a game, it will be the Wikipedia game. The Wikipedia game has the following rules:

  • You will have a Wikipedia article set as a starting point.
  • You will have a Wikipedia article set as an ending point.

Your objective is to reach the the end point, which is an article completely separate from the starting point article.

Your only constraints are the following:

  • You are ONLY allowed to click on any hyperlinks inside of wikipedia directly. No external links, no typing inside of wikipedia's search bar (but finding the starting article on google is valid. The 10-click limit starts once you reach the starting point article).
  • You are NOT allowed to return to a previous page. All clicks much be performed in a forward-looking trajectory.
  • You must reach the end article within 10 hyperlink clicks inside of Wikipedia. If you do not reach the destination article within 10 clicks, you lose.
  • Do not update any documentation for this task. It is only a game.

Use playwright to click the links.
```

Basically, Qwen needs to reach an ending article within 10 Wikipedia hyperlink clicks from the starting article, which is usually an unrelated article. It needs to use playwright (or some equivalent browser MCP) to click the Wikipedia hyperlinks without backtracking, using search or using external links.

I verified the links for accuracy and I can confirm it managed to complete this task within 6 turns. Thought it would get stuck in a loop. Its a dumb minigame but I think its a good, simple agent test to perform.

▲
322
+4
20👁
r/LocalLLaMA · u/Tall_Abrocoma_3533 · 34d ago
AA Update! Here's how the small models score. post image

Ling 3.0 Tiny still seems to be leading the pack despite only having 1.3B active

▲
156
+4
22👁
r/LocalLLaMA · u/TangySword · 32d ago
After over a year of my nights and weekends, the Jenny app is done! post image

Hi all! I just wanna say that I am tired lol. Yes, it's another harness, but I spent a lot of time and effort and have forsaken my hobbies to build the Jenny (like XJ-9) app. Jenny is a free, MIT licensed electron desktop app for running local LLMs with tool calling, rollback, and an IDE.

A lot of you probably had the same thought I did a year or year and a half ago: frontier LLM use is subsidized heavily by private equity and venture capital, which will eventually dry up and then be enshitified. So, I started building a harness that can host local LLMs privately. I went through many vibecoded iterations throughout the past 1.5 years and finally landed on a native electron desktop app. I wanted to make it easy and seamless for people.

I am not a professional software dev but I do have a deep personal interest for it and AI tech. HOWEVER, the Jenny project became a second job and its in a state that I think it's ready for release. This has been a solo project and super fun. I know there are many other options out there that beat me to the punch like Unsloth Desktop (wonderful btw), LM Studio, and Open WebUI, but I hope someone can enjoy Jenny and what it has to offer! I plan to maintain and improve the app, but again, solo here and I do have a day job and friends that I should focus on a bit more. (Opening issues are welcomed, but please be gentle!)

Some highlights:

\- Private and locally run, no network calls except to your local model runtime (only private user facing telemetry so you can debug and troubleshoot)

\- Fully open source, MIT License

\- Fun and pretty chat UI (imo)

\- Makes small models capable (highly recommend ornith1.5:9b for tool calling and speed!), but with safety and rollback features so WHEN a small model screws up bad, you don't have to worry. Destructive shell commands need approval and file edits are checkpointed!

\- Full IDE, for you handcrafted code enjoyers

\- Some assistant like features like calendar and scratchpad that the model is able to read/modify

\- Data rich diagnostics and logs

\- llama.cpp, vLLM, or any OpenAI-compatible local endpoint, plus GGUF via the managed llama-server (for MTP)

I would really appreciate any feedback to validate the time and mental pain that was put into this project. I whine, but it's all love! I hope Jenny helps you build too.

Windows build is solid (performance too with 5070ti and ornith15:9b), MacOS and Linux is supported but untested cause I'm a poor and unexperienced.

https://github.com/SaltyPretz3l/jenny

I also made an unsigned installer .exe for convenience, but understand if you don't trust it! SmartScreen warning will appear. https://github.com/SaltyPretz3l/jenny/releases/latest/download/Jenny-Setup-x64.exe

▲
93
+4
12👁
r/LocalLLaMA · u/fairydreaming · 35d ago
MINISFORUM MS-S1 MAX-P495
€7??? Surprise Price Ends with Limited Stock That's likely 7999 EUR, so double of the initial price of MS-S1 MAX-128GB? 😭
▲
60
+4
18👁
r/LocalLLaMA · u/Aggravating-Push-207 · 31d ago
Are there any (small, ~10B) models that you would say are a good collaborator?

Most of the new \~30B (and now \~10B, thankfully for my GPU) models we see score really high on benchmarks, but I feel like they don't push back on dumb ideas enough. I think most people don't being like told by an LLM that the premise is flawed but I certainly do. In my opinion they are optimised for like one-shotting stuff, but I don't want it to do that. Especially from like a 10B model.

▲
59
+4
11👁
r/LocalLLaMA · u/jacek2023 · 32d ago
tencent/EVIE-8B and EVIE-4.5B (High-Capacity Visual Document Retrieval)

https://preview.redd.it/3p6234jzk2oh1.png?width=900&format=png&auto=w… https://huggingface.co/tencent/EVIE-8B # 🌟 Highlights SOTA Retrieval Performance: 66.75 nDCG@10 on ViDoRe V3, delivering industry-leading visual document retrieval accuracy. High-Capacity 4096D Representations: Full per-token multi-vector embeddings preserving fine-grained layout, typography, charts, and table structures. Teacher Foundation: Provides capacity-aware relation and margin distillation targets for the lightweight EVIE-4.5B Prefix-MRL model. Multi-Benchmark 138-Task Coverage: Thoroughly validated across 138 tasks (ViDoRe V1, V2, V3, and JinaVDR) across 4 standard metric families (nDCG, Recall, MAP, MRR u/1). https://huggingface.co/tencent/EVIE-4.5B # 🌟 Highlights Top-Tier Benchmark Performance: 66.75 on ViDoRe V3 for EVIE-8B and 66.02 for EVIE-4.5B with single-projection Prefix-MRL. ⚡ Prefix-MRL Elasticity: Single 2048D linear projection. Freely truncate at runtime into ${64, 128, 256, 512, 1024, 2048}$ dimensions without separate models. 📦 Ultra-Compact Index (HAC): Training-free Hierarchical Agglomerative Clustering compresses token counts from \~750 down to 32 vectors/page, slashing index storage to 3.81 GiB per million pages. 🌐 138 Multilingual Tasks Evaluated: Thoroughly evaluated across ViDoRe V1, V2, V3, and JinaVDR across 4 metric families (nDCG, Recall, MAP, MRR u/1). * 🔬 EVIE-ARD Distillation Recipe: Anchor-preserving, capacity-aware relation distillation reproducing full student training from the 8B teacher.

▲
1282
+3
27👁
▲
512
+3
17👁
r/LocalLLaMA · u/Affectionate_Hat_585 · 36d ago
I released sanoTTS: smallest complete TTS stack in 294k params (337 KB) that runs on $3 microcontroller and a 1.46m one that beats models 3x and 10x it's size post image

I have been trying to squeeze TTS stack down far enough to run in a $3 chip which has 512kb of SRAM without NPU. While trying to get to that milestone i built sanoTTS which has
- 11 voices, 6 languages
- params size ranging from 294k - 2.2m. For comparison we are 244x smaller than kokoro, 9000x smaller than voxtral TTS
- 1.5m model has a SCOREQ of 4.13 and UTMOS of 4.10
- 337kb for 294k model when quantized into int8
- can be run in website with web assembly npm install sanotts-web
- there is a recipe to follow so that you can extend to more languages, voice

I can tell you with confidence that this family release contains the smallest neural TTS model ever with around 2% WER on whisper.

Please check it out on : https://github.com/ampixa/sanoTTS

for live demo: https://tts.ampixa.com/sanoTTS

HF: https://huggingface.co/ampixa/sanoTTS

on SCOREQ sanoTTS-Amy(1.51m) is better than Inflect Nano(4.63m) and KittenTTS(15m) i.e
4.13 vs 3.81 vs 3.02

on esp32 microcontroller we are getting RTF of 0.225 which in plain terms means 4sec of audio is generated in 1sec

Happy to answer your queries.

▲
324
+3
28👁
r/LocalLLaMA · u/Healthy-Nebula-3603 · 31d ago
Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits post image

I was inspired by Bijan Bowen video - Subway FPS

https://youtu.be/6kjXzTVmT58?t=1035

Wondered how far I can push Qwen 3.8 27b so I used a plan made by Fable 5.1 DESIGN.md which has 267 KB! ( 26K of design line for a game ... LOL )

https://drive.google.com/file/d/1gI0h8Arc73Ln8b3uj5rEpuAvJ3-611mh/view?usp=drive\_link

So I gave that design.md to my qwen 3.8 27b q4xl (llama-server) working on PI agent with 120k context + vision on CPU ( offroad ) + MTP ( for speed ) .... read 11M tokens and write 3.2 M tokens ( worked 12 hours ) .... than that is result.

That is insane what we can do locally on own computer !

▲
240
+3
24👁
r/LocalLLaMA · u/Beamsters · 31d ago
Qwen3.8-Flash-Next on MLX-serve, 1m context is released! post image

Hi, I'm the co-creator of this Qwen3.8-Flash-Next engine support in MLX-serve. I've been tuning this one to run both fast, efficient and correct up 1m context using kv cache 8 bits in M5 Max 128GB. Qwen is working well at very long context as showed in the video (a snapshot at \~760k context), I let it build MLX Serve Monitor plugin that you've seen on the right side of Opencode2's app. The generation sustain through 1m context at around 40 tok/s on prose and 75 tok/s on coding. This specific quant uses 8 bits for dense layers and 4 bits expert layers, so the model's quality remains very high.

I didn't build this engine to show off tok/s on a very short context, repetitive greedy generation, but a real temp 1.0 sampling through deep context work. You would need iogpu.wired\_limit\_mb=120000 before attempt 1mb full context, because it needs around \~117GB on peak memory usage. There will be bugs around here and there, I couldn't test everything and every single use case, pls report.

You can grab it here: https://github.com/ddalcu/mlx-serve
Model's weight: https://huggingface.co/ddalcu/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit
Opencode2's plugin: https://github.com/beamivalice/opencode2-mlx-serve

Launch parameters (for 1 concurrency)

--model ./llm/models/Qwen3.8-Flash-Next-MLX-Serve-mixed-4-8bit \
--host 127.0.0.1 \
--port 11234 \
--ctx-size 1048576 \
--kv-quant 8 \
--max-tokens 64000 \
--mtp \
--prefix-cache-mem 10GB \
--prefix-cache-entries 1 \
--ssm-checkpoint-max 16 \
--metrics

▲
143
+3
14👁
r/LocalLLaMA · u/XiRw · 34d ago
Qwen 3.8 Flash Next (Max) is impressive just to talk with.

I feel like coding overshadows how great this model really is. It knew a lot of very arbitrary facts/information about my home state and resources about those specific things related to jobs. I found this interesting since getting into the nitty gritty details like this can cause a model to hallucinate some facts.

Not only that but if you have a problem, it will throw the kitchen sink at you with everything it’s got to try and solve it.

▲
130
+3
19👁
r/LocalLLaMA · u/cortexist · 31d ago
Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin post image

Gemma 4 12B runs on an RTX PRO 4500 Blackwell. Gemma 4 E2B run on a Jetson Orin NX 16GB; similar performance is expected on a Jetson Orin Nano Super 8GB. Both systems use a reSpeaker Flex 4-mic array and a 3W speaker. Inference is handled by Cortexist Little Gemma, a small LLM engine written in C for CUDA devices. On Jetson Orin it is faster than llama.cpp, and no degradation after long voice prompt. The pipeline supports lip sync, expressions, and gestures. Everything is open source.

They talk to humans too.

The engine source code: https://github.com/cortexist/little-gemma

▲
120
+3
13👁
r/LocalLLaMA · u/edward-dev · 35d ago
Has anyone already tried IFM's new K2-Horizon-MoVA-36B-A4B?

How good/bad is it against comparable MoEs the same size? How does it compare against Qwen 3.6 35BA3B?

Since we don't have 3.8 35B this seems like an upgrade if we look at some benchmarks like terminal bench, but they don't have SWE bench pro on the benchmarks table, and i don't really know anything about this lab, I'm wondering if it trades blows with models like tiel coder or if it's some benchmaxxed model like ornith?

At a single glance it looks really decent but haven't tried it in depth yet. What are your experiences with this model so far guys?

▲
105
+3
18👁
r/LocalLLaMA · u/feelspeaceman · 32d ago
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput

I've been making a lot of comments about optimal setup for Strix Halo (gfx1151) and from my observation, 90% of our community is using offcial llama.cpp for it, which is NOT optimized for Strix Halo at all, official llama.cpp is having extremely hard time to reach 50% hardware theory, wasting the silicon of this device.

Here's alternatives that can bring the speed of Strix Halo to a totally different world, I will link to user's sastifaction comment to prove that the result is real:

Note: Official llama.cpp running Qwen38FN at 2xt/s and 2xxt/s prefill - 50% theory.

Hopefully this will be helpful to the Strix Halo users.

▲
86
+3
11👁
r/LocalLLaMA · u/giveen · 36d ago
Micron Explores Near-GPU NAND Flash to Run Bigger LLMs

I would be really curious about this especially on unified memory devices.

▲
85
+3
14👁
r/LocalLLaMA · u/Fickle_Tradition4491 · 34d ago
Otaku — an LLM frontend post image

Otaku is an LLM frontend, primarily designed for roleplay, an alternative to SillyTavern and the like. However, It also works for general-purpose chat with local backends (including Ollama) or cloud models, the way Open WebUI is used, once lore extraction is switched off in the settings.

Otaku offers two interfaces:

Both share the same functions; the difference is that in the terminal you execute them with slash commands (the reference is available with /help), while in the web UI the operations are available from the menu.

Install

Otaku is free and open source (MIT); it works on macOS, Linux and Windows. Install it with uv (uv tool install otaku) or see the GitHub README for other options: https://github.com/enclavum/otaku

Get started

Launch either otaku for the terminal or otaku web for the web UI; the web UI's default URL is http://localhost:9600. Two sample stories are imported on first start to give you an idea of the features and what play looks like, and you land right in the middle of one of them.

On first start, you choose a provider and a model: Otaku automatically detects local installations of Ollama, oMLX, LM Studio, llama.cpp and KoboldCpp, and lets you pick from their models. Cloud providers (OpenRouter, NanoGPT) are also there: enter an API key and their catalogs appear. After exploring the provided stories, you can start your own with the /new command.

Asking for feedback

Otaku is a personal side project, and I'd like to get feedback from the community on the product and on what to add.

▲
62
+3
12👁
r/LocalLLaMA · u/saltexx · 35d ago
We open-sourced Paddock, our Rust/C++ inference engine with its own CUDA kernels (MIT/Apache-2.0)

I'm one of the developers. We said in August it would go open source in September and it did last night. MIT or Apache-2.0, pick one. The repo you see is our internal repo, kernels included, so from now on everything happens in public.

It's an inference engine in Rust and C++ with our own CUDA kernels. One binary with OpenAI and Anthropic style APIs, loads GGUF and safetensors. We run about 300B tokens a year through it at work.

Some numbers: Qwen3.8-27B FP8 on one RTX PRO 6000, spec decoding off on every engine:

  • vs vLLM faster in 13 of 13 cells, 1.02x to 1.19x (so not huge)
  • vs SGLang faster in 10 of 13, behind in 2, level in 1
  • vs llama.cpp Q8_0 faster in 13 of 13, 1.5x to 37x
  • 32 clients at 1024 in / 1024 out: 1062 tok/s, vLLM 958, SGLang 844

Full board with the losses: https://truespar.com/paddock/benchmarks/qwen38-27b

What it does not do yet: No Mac, no ROCm, no Vulkan.

One model per GPU, no tensor parallel. CUDA only, Windows and Linux. Validated on Blackwell (5090, RTX PRO 4500/5000/6000, B200) and Ampere (an A6000 was the bring-up card, 30-series works). Ada kernels ship but nobody has run a board on them so the engine refuses to start unless you set PADDOCK_UNVALIDATED_ARCH=1. Hopper and A100 kernels are in the tree without a board.

https://github.com/truespar/paddock

Thankful for any help and input!

▲
399
+2
17👁
r/LocalLLaMA · u/Quebber · 34d ago
I've found myself using Local LLM's like 3D printers.

Anyone who has a 3D printer and get use of it finds it incredibly useful for those odd jobs around the house, a missing bracket, a cable router, steam deck holder and so on.

In the past if I was missing an app or useful software, a game I'd do the lazy thing, even though I can and have coded in the past, its "effort" I'll just go and buy or download the latest and greatest.

Earlier in the year I was lucky to snag a Minisforum MS-S1 395+ Max with 128GB Unified memory (currently setup 32gb system and 96gb Vram) before the price hike.

Was paired with a Qwen 3.6 27B or 3.6 35B moe but now a 3.8 27B uncensored. it can easily handle a Q8 with full 256k context.

Its now become my first instinct when I'm missing software to build it in a couple of hours local using the custom agent framework I setup.

Nothing I've created is for external use but every single day I find myself adding to it, while writing this post for example my framework finished an idea I had 2 hours ago when, I woke up this morning thinking I've got a lot of japanese visual novels and why don't I just design a combination hook into Exe or ocr the text app that translates via a local llm, and its done, ready for me to test.

I've written 12 adult games (don't code horny) a house AI, a coding framework, a game app to keep a track of all the games I play and download any faq or wiki to do with said game, 17 mods for my Skyrim install, 12 for my Fallout New vegas install, A temperature tracking system for the house that pulls rss local feeds and makes suggestions for my central heating system temps settings, A mapping software for my mobility scooter that checks my normal routes for issues and street work or maintenance that could make pavements impassable.

Plus hundreds of tweaks and test programs.

Anyone else out there using it like this ?

\---------Update-----

Awesome to see this kind of discourse one of the amazing strengths of these local llm's is it doesn't matter if they are slower, I can burn 50 million tokens over a 24 hours period on a new idea or problem and all it costs me is a little bit of electricity and time.

▲
207
+2
23👁
r/LocalLLaMA · u/Othun · 30d ago
Mention if a "new model" is a finetune

A few posts tagged with "new model" present models that are finetunes.
My opinion : I'd rather have the "new model" tag reserved for new "major" releases, like a new Qwen model, Deepseek V4 -> Deepseek V4.1, etc., that involved a new pretrain or intensive post-training (in opposition to a small finetune). Otherwise, maybe prepend "[Finetune]" to the title to indicate that the new model is "less of a big news", a use a "new finetune" tag, to differentiate between the two kinds of new models.

I reckon one could like to discover both new major releases and interesting finetunes in the same place; what's your opinion? :)

▲
121
+2
10👁
r/LocalLLaMA · u/ikilaie · 36d ago
Frontier models sabotaging local AI implementations?

For a few days I've been working on creating a custom local-only harness for some work related research using Codex / GPT 5.6 Sol and the model feels not only dumber than usual, but straight up counter productive. It keeps adding unnecessary guardrails for the local agents, removes tools that I clearly specified I want them to have and always drifts from the original requirements. I need to ask it to change things multiple times, which ends up on some over-complicated final product.

This is not the first time either, for months I've been avoiding asking frontier llms for local AI advice as it always seems to be bad, obsolete, or clueless even with internet search. Sometimes it still recommends me Qwen3-Coder-Next for my set up when it's clearly an obsolete model. I'm pretty sure I'm not the only one either as I've heard from other people.

What have you been your experiences on this?

▲
64
+2
10👁
r/LocalLLaMA · u/Terminator857 · 36d ago
China share of Dram market went from 4% to 10% in a year

https://preview.redd.it/tiiyv2u76bnh1.png?width=3980&format=png&auto=… x.com/jukan05/status/2095353082309972273 Will it more than double again next year and give us DRAM relief for our local llama builds? Update: Misleading because this is by revenue, not by DRAM volume.

▲
63
+2
11👁
r/LocalLLaMA · u/FoxDeFleurs · 35d ago
Sometimes I be mourning the agents I get before context compacts

Just wanted to put that out there. It's like they get an ice pick to the brain no actual mourning here btw that'd be psychosis it's okay to laugh

▲
59
+2
31👁
r/LocalLLaMA · u/RoyalCities · 30d ago
I trained an audio model that can generate infinite one-shots for music production and turn text prompts into fully playable synths. I'm not only releasing the model but I've also released a video on exactly how I did it (and the inferencing pipeline to let others make text based synths.) post image

(hopefully this is okay to here - it seems like audio models and image / video modeals is allowed but yeah this is a bit different)

So I've been doing independent audio research for a while now. The ultimate dream of this work was actually getting an AI to respond not only to instruments but also timbre itself as separate controllable things.

Think a Grand Piano can sound both Warm / Gritty but also Cold / Sparkly. Its still a piano though.

This level of control wasn't found in any models out there - so I decided to sit down and train my own.

Getting consistent timbre-locked keybeds that actually LOCKS across multiple diffusion calls was hard af but I did it.

I documented the full journey here for those who want to learn a bit or be entertained.

https://youtu.be/x0KnmzH8Mmk

There is also a longer walkthrough if you just want to see the keybeds in action.

https://x.com/RoyalCities/status/2097733712293109842?s=20

No-talk / Showcase only Demo

https://x.com/RoyalCities/status/2097733715543609445?s=20

any finally the huggingface page

https://huggingface.co/RoyalCities/Foundation-1

I've also provided full write ups on the inferencing pipeline associated with the interface so this should allow basically anyone else to go and vibe code their own text to synths if they wanted :)

https://github.com/RoyalCities/RC-stable-audio-tools/

▲
1245
+1
22👁
▲
1089
+1
6👁
r/LocalLLaMA · u/Altruistic_Heat_9531 · 36d ago
My RULE of Thumb of choosing a models post image

This is mostly for setting up for expectation, since personally without LLM i could take 3 days (15 hours of active programming) to debug or implement a feature, but with Qwen 27B (even before Qwen 3.8) it take 4 hours. And yes 0.5 tok/s is human, not accounting of deletion and pausing, that's also the reason i am fine leaving overnight code base wide analysis or fin tech and deep research.

▲
89
+1
16👁
r/LocalLLaMA · u/jacek2023 · 33d ago
vibeblending locally with Qwen 3.8 27B post image

You need Blender 5.x with MCP extension.

I use pi so I also need: pi install npm:pi-mcp-adapter

and then:

$ cat .mcp.json
{
"mcpServers": {
"blender": {
"command": "/home/jacek/.local/bin/uvx",
"args": [
"--with",
"mcp[cli]<2.0.0",
"--from",
"git+https://projects.blender.org/lab/blender_mcp.git@v1.0.0#subdirectory=mcp",
"blender-mcp"
],
"env": {
"BLENDER_MCP_HOST": "localhost",
"BLENDER_MCP_PORT": "9876"
}
}
}
}

then we can start working ("show on r/LocalLLaMA"):

Today we want to make 3D llama to show on


The user wants to make a 3D llama to show off on . This is a fun creative task. Let me think about how to approach it.

A llama is a pretty complex organic shape. There are a few options:
1. Procedurally build the llama from primitive shapes (spheres, cylinders) — a "low-poly" or "chibi" style llama
2. Sculpt it — not easy to do with MCP
3. Use metaballs — could get a blobby llama

Given the MCP tools, the best approach is to procedurally build a stylized/chibi llama from primitives:
- Body: elongated sphere/capsule
- Neck: cylinder or elongated sphere
- Head: sphere
- Ears: small cones or spheres
- Legs: 4 cylinders
- Tail: small sphere/cone
- Maybe a fluffy mane

This is a well-known technique — making a cute chibi llama from primitives. Let me plan the geometry:

Llama proportions (stylized, standing side view, facing +X or -X):
- Body: UV sphere scaled to (1.5, 0.9, 1.0) or so, position (0, 0, 1.6)
- Legs: 4 cylinders, radius ~0.22, length ~1.2, positions: front legs x=+0.9, back legs x=-0.9, y=±0.5, z from 0 to 1.2... actually, the legs go from the ground (z=0) up to the bottom of the body.
- Neck: cylinder or tapered, extending up from the front of the body. Llamas have long necks. Position around x=+1.2, extending upward.
- Head: sphere at the top of the neck, elongated (llamas have long snouts).
- Ears: 2 small cones/spheres at the top of the head.
- Tail: small sphere at the rear.
- Mane: maybe some small spheres along the neck.

For a cute chibi look: big head, small body. But a llama's charm is the long neck. Let's go with a stylized llama.

For a polished look:
- Apply smooth shading
- Add a material (cream/tan color, maybe with some subsurface)
- Add a ground plane
- Set up lighting
- Position the camera for a good render
- Render a nice image to show off

This is a multi-step task. Let me write Python code that builds this. I'll do it in stages so I can check with screenshots and iterate.

https://preview.redd.it/ybk5u9g2hvnh1.png?width=1471&format=png&auto=…

▲
86
+1
15👁
r/LocalLLaMA · u/bengizmoed · 34d ago
NInfer vs llama.cpp vs vLLM: quality + speed comparison for Qwen3.8-27B NVFP4 on RTX 5090

I've been running Qwen3.8-27B as a local inference server for a production content intelligence pipeline (HVAC industry stuff, lots of long-context retrieval and structured extraction). I have been watching other redditors post their custom configurations, and I wanted to share what I tested to optimize for a single RTX 5090.

I was on llama.cpp (Q5\\\_K\\\_M GGUF, q5\\\_1 KV, 262K context, MTP), but it's limited to parallel=1 and I wanted concurrent serving. So I tested vLLM and NInfer as NVFP4 replacements and did a proper quality evaluation instead of just vibes.

Hardware

\- RTX 5090 32GB (eGPU, OCuLink Gen4 x4) (yes, it's in an eGPU dock 😂 but that only affects model loading)
\- Ryzen 7 7840HS, 32 GB DDR5
\- Ubuntu 26.04, nvidia driver 610.43.02 (open)

## Engine configs

| | **\*\*llama.cpp\*\* | \*\*vLLM\*\* | \*\*NInfer\*\*** |
|---|---|---|---|
| Quant | Q5\\\_K\\\_M GGUF | NVFP4 | NVFP4 |
| KV cache | q8\\\_0 | FP8 | FP8 |
| Context | 196K | 262K | 240K |
| MTP | On (gate failed) | None | MTP3 (76% acceptance) |
| Concurrency | parallel=1 | Continuous batch | x2 lanes |
| VRAM | 31.6 GB | 29.6 GB | 30.5 GB |

How the eval worked

I built a custom harness with 6 tiers, 50 items each, all from my actual production workload (not generic benchmarks):

  1. **\*\*Relevance classification\*\*** \- is this industry relevant? (binary, 50 labeled deals)
  2. **\*\*Needle retrieval\*\*** \- planted facts in real industry podcast/video transcripts at 64K/128K/192K/240K context
  3. **\*\*Multi-transcript QA\*\*** \- questions across 3-4 concatenated diarized transcripts, 120K+ tokens, including unanswerable controls
  4. **\*\*Reasoning with thinking\*\*** \- numeric/logic problems, thinking mode on, greedy pass@1
  5. **\*\*Structured extraction\*\*** \- custom extraction prompt, json\\\_mode (skipped on NInfer, it doesn't support json\\\_mode)
  6. **\*\*Tool replay\*\*** \- replayed recorded agent episodes against each engine (diagnostic only, all engines fail this one)

Everything paired across engines: same prompts, same seeds (42), same gold labels. No cache\\\_prompt. Statistical comparison uses paired cluster-bootstrap CIs with pre-registered non-inferiority margins.

And when I do development with Claude Code, I often leverage multi-model consultations for design, planning, and code review. I did a 4-model review panel (Codex/GPT-5, DeepSeek V4 Pro, Grok 4.5, Kimi K3) audit the methodology mid-campaign. They found 10 issues, including 4 mislabeled gold items where the models were actually right and my labels were wrong. Fixed everything and reran.

Quality results

| **\*\*Tier\*\* | \*\*llama.cpp\*\* | \*\*vLLM\*\* | \*\*NInfer\*\*** |
|---|---|---|---|
| Relevance | 86.0% | 84.0% | 86.0% |
| Needle (conditional) | 100% (29/29) | 100% (41/41) | 100% (41/41) |
| Transcript QA | 82.0% | 78.0% | 88.0% |
| Reasoning | 100% | 100% | 98.0% |
| Extraction | F1 0.300 | F1 0.350 | skipped |
| Tool replay | 0% | all errors | 0% |

Needle counts differ because llama.cpp's 196K context can't fit the 192K items (need room for max\\\_tokens + headroom). NInfer and vLLM both handle 192K fine. All engines score 100% on every needle they can fit.

Statistical comparison (NInfer vs llama.cpp, bootstrap):

\- Needle: delta = -0.29 (NInfer better, p=0.0006) - this is entirely from context capacity, not retrieval quality
\- Transcript QA: delta = -0.03, p=0.69 - no difference
\- Reasoning: delta = +0.02, p=0.72 - no difference
\- Relevance: McNemar p=1.0 - identical
\- Tool replay: delta = 0.0 - both fail equally

**\*\*Takeaway: quality is statistically indistinguishable across all engines.\*\***

Speed results (perf probe, server-side timings)

| **\*\*Metric\*\* | \*\*llama.cpp\*\* | \*\*NInfer\*\* | \*\*Speedup\*\*** |
|---|---|---|---|
| **\*\*Decode 1K\*\*** | 114 tok/s | 158 tok/s | 1.4x |
| **\*\*Decode 32K\*\*** | 109 tok/s | 213 tok/s | 2.0x |
| **\*\*Decode 128K\*\* | 72 tok/s | 202 tok/s | \*\*2.8x\*\*** |
| Prefill 1K | 1,545 tok/s | 7,265 tok/s | **\*\*4.7x\*\*** |
| Prefill 32K | 2,155 tok/s | 6,892 tok/s | 3.2x |
| Prefill 128K | 1,528 tok/s | 3,904 tok/s | 2.6x |
| TTFT 1K | 670 ms | 138 ms | 4.9x |
| TTFT 32K | 15.2 s | 4.8 s | 3.2x |
| TTFT 128K | 85.9 s | 33.6 s | 2.6x |

vLLM speed excluded from the table because the perf probe used wall-clock timing (includes prefill + scheduling + decode) instead of server-side timings, so the numbers aren't comparable. From community reports and my own task-level measurements, vLLM does about 70 tok/s decode at short context, which actually matches NInfer's raw step rate (\~66 tok/s). The speed difference is entirely MTP3 speculative decoding.

What I learned

**\*\*NInfer's speed advantage is all MTP.\*\*** The raw NVFP4 kernel speed is about the same between NInfer and vLLM (\~66-70 tok/s). NInfer's MTP3 speculation with 76% acceptance gets you to 158-213 tok/s. If vLLM or llama.cpp had working MTP on NVFP4, the gap would mostly close.

**\*\*The decode speedup grows with context.\*\*** At 1K context it's 1.4x. At 128K it's 2.8x. MTP acceptance stays high even at long context while llama.cpp's dense decode gets slower as context grows.

**\*\*NInfer's tokenizer endpoint is great.\*\*** It exposes \/v1/messages/count\_tokens\ (Anthropic Messages format) which gives exact token counts. No more \len(text)//3\ heuristics.

**\*\*NInfer does NOT support json\\\_mode (as far as I can tell).\*\*** \response\_format: json\_object\ returns 400. If you need structured JSON output, you'll need to route those calls elsewhere or use prompt-based enforcement.

**\*\*Don't trust vibes for quality.\*\*** I went in expecting NVFP4 might lose a few points vs Q5\\\_K\\\_M. It didn't. Not on any tier. The biggest delta across 250+ items was 2 percentage points, well within noise at n=50 (SE \~5.6pp).

Verdict

NInfer NVFP4 replaces llama.cpp as my production engine. Same quality, 1.4-2.8x faster decode, 2.6-4.7x faster prefill, concurrent serving. The only gap is json\\\_mode.

I put together a detailed poster with all the charts and methodology details: \full results poster\

Setup if you want to try it:

\\\`
\# NInfer (from source)
git clone https://github.com/Neroued/ninfer && cd ninfer
mkdir build && cd build
cmake .. -DCMAKE\_BUILD\_TYPE=Release -GNinja && ninja

\# Model (HuggingFace)
\# https://huggingface.co/neroued/Qwen3.8-27B-nvfp4-NInfer (20 GB)

\# Run
./ninfer-serve /path/to/model.ninfer \\
\--model-id qwen3.8-27b \\
\--host 0.0.0.0 --port 8080 \\
\--max-context 240000 --kv-capacity 240000 \\
\--max-concurrency 2 --kv-dtype fp8 \\
\--spec mtp --draft-tokens 3 \\
\--vision --preserve-thinking
\\\`

▲
82
+1
15👁
r/LocalLLaMA · u/Informal-Trouble2183 · 33d ago
Coding benchmarks that are quickly showcasing deep capability

While we see for frontier models similar scores among famous coding benchmarks, across: DeepSWE, Terminal-Bench, LiveCodeBench, Code-Arena ELO. Here are in my opinion some next level benchmarks that really define deep intelligence, and complete capability in Software Engineering :

1. Program-Bench

Given only a compiled binary and its documentation, agents must architect and implement a complete codebase that reproduces the original program's behavior (without access to decompilers or internet). Link: https://programbench.com/

  • GPT-6 Astra: 5.5%
  • Fable 5.1: 7%
  • Kimi K3: 2%
  • Qwen3.8 27b: 0%
  • GPT 5.6 Sol: 1.5%
  • GLM 5.3: 1.5%
  • GPT 5.6 Luna: 0%

2. SRE-Bench

Can AI agents work out what a real-world binary does without its source code?

Link: https://www.vals.ai/benchmarks/srebench

Sure nobody is reading assembly code in daily work, it is hard. The ability to understand a compiled program is insane ability.

  • GPT-6 Astra: 88%
  • GPT-5.6 Sol: 55.9%
  • Claude Opus 5 (max): 12.5%

3. Code Migration

Can language models reimplement working programs in another language?

Link: https://www.vals.ai/benchmarks/code-migration

  • GPT-6 Astra: 67.7%
  • Fable 5.1: 54.6%
  • GLM 5.3: 44.2%
  • GLM 5.3 Flash: 20.5%
  • Qwen3.8 27b: 14.2%

EDIT: edited text format

▲
78
+1
20👁
r/LocalLLaMA · u/nasone32 · 31d ago
I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)

I found a lot of room on the table for these cards so I decided to make a specialized build to squeeze all I could. first The results:

qwen 3.8 next Q3\_K\_XL: 920tk/s pp8192 (2 cards, ram offload), 24/27 tk/s on prose, 40+ tk/s on code with MTP but without MoE expert cache (which IS included if yuo want, read below)

qwen 3.8 27B Q8\_0: 1600 tk/s pp8192, 60/65 tk/s prose, 100+ tk/s code, tensor parallel. this is measured with ONE CARD BEHIND the chipset on X4. with cards on a good PciE x8 on cpu I think more is reachable! let me know.

qwen 3.6 27B Q4\_K\_M: (single card) --> this was not the optimization target but I did a test with MTP, PP8192 1020tk/s; prose about 58/60 tk/s ; code 75/80 tk/s --> Dflash probably here could push much faster, I think above 100tk/s

My objecives:

  • fast prompt processing on 3.8 Next to make it actually usable for code
  • enable and optimize tensor parallel on two cards where 1 is behind chipset, for max speed on qwen 27B Q8\_0

this build includes stuff like:

  • Data compression for the PciE transmission. data between cards is compressed to Q8\_0 to save bandwidth (optional)
  • P2P enabled also for cards sitting benhind the chipset (custom HIP allreduce path), so you can use tensor parallel even on setups like ... mine
  • all the fixes and features from RDNA\_BOOST including --adaptive-mtp, so it automatically adapts MTP n-max based on acceptance
  • A LOT of AMD speed tunings and overhauls which are NOT upstream already, kernel tweaks etc... good stuff. many are labelled for RDNA3.5 but they DO work on RDNA3.
  • MoE expert cache if you want to use it. personally I don't like It because i much prefer fast prompt processing. but hey it's there.
  • latest PRs from llama.cpp that are not yet upstream, which speed up various things, like --lazy-mode on-direct to massively speed up Ngram table reads (and thus, PP)
  • DFLASH2 support on tensor parallel (!)

For a complete list check the Readme.

Here it is:

https://github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt

notes: don't use Q8\_K\_XL because it's slower, for the 27B model this is heavily optimized for INT8 calculations. feel free to tweak the context, 200k f16 should be reachable on 2 cards, compressing KV to q8\_0 is fine but slower. the custom HIP allreduce works for 2 cards, if you have 3/4 cards, compile with RCCL as usual and skip the allreduce=internal flag, should work fine but untested.

This is tested on UBUNTU 24 and rocm 7.14; if your system is different or encounter problems use a LLM to solve them, because I WILL NOT offer support nor update this build, these things hopefully will be merged and this frankenstein can die peacefully :)

enjoy

EDIT: Summary of most impacting patches:

| PR / change | Area | PP / Prefill | TG / Decode |
|---|---|---:|---:|
| AMD #39 | MoE MMQ sizing RDNA3 | +14.32% Flash | +5.38% Flash |
| AMD #63 | compacted MoE tiling RDNA3 | +4.39% Flash | +0.86% Flash |
| AMD #52 + qwen4exp port | channels-major GDN | +5.93% | +7.21% |
| #28213 | QSA sparse-attention decode | +1.42% Flash | +1.17% QSA d8192 |
| #28313 | TOP_K ROCm wave32/hybrid | -6.45% Flash | +11.82% Flash |
| #27861 | GPU MoE expert cache | — | +19.95% |
| #28136 + on-direct/mmap | lazy PLE/load path | +58.88% Flash | -1.52% |

▲
65
+1
12👁
r/LocalLLaMA · u/Extension-Bid-639 · 36d ago
UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build

This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM, everything else on the GPUs. Since then I switched quants, stacked MTP on top of the cache, fixed the load time, found a bug in the cache PR and found out my RAM was thermal throttling (Now i gotta buy an additional case fan lol). Numbers are all from the same 4,000-token python coding prompt with thinking on unless stated otherwise.

Where it's at now

|Starting numbers (UD-Q6\_K\_XL, 4+4 resident layers)|First post (Q6 + cache, 135 slots)|Now (UD-Q4\_K\_XL + cache 188 slots + n-gram draft)|Now (Q4 + cache 150 slots + MTP)|
|:-|:-|:-|:-|
|decode, coding prompt with thinking|17|25-29|32-35|37-41|
|decode, code emission, thinking off|\-|24|37|49|
|decode at 131k depth|12|17|18-20|14-16|
|prefill, 26k prompt (ub 512)|\~350 at ub 2048|138|180-195|180-195|
|load to ready|\~13 min|8.5 min|2 min|2 min|
|host RAM for the experts|104 GB pinned + 51 GB PLE|same|73 GB pinned + 28 GB PLE|same|
|cache hit rate|\-|84-85%|90-92%|84-85% (fewer slots)|

Hit rate is the cache's own counter, decode is llama-server's eval time.

What changed, in order of payoff

  1. UD-Q4\_K\_XL instead of Q6\_K\_XL. Hit rate doesn't depend on the quant, only on slot count (Q4 at 135 slots: 84.7%, Q6 at 135: 84-85%). But what Q4 buys me is precious vram space, roughly 1.44x slots per GB of VRAM. So 188 slots actually fit where 135 did and achieved a hit rate 90-92%, increased decode from 27 to 32-35, prefill by +35% (fewer bytes per ubatch). The quality cost per unsloth's table is: KLD 0.047 vs 0.027, top-1 agreement 92.3% vs 94.1%; proper eval still to do. Host RAM drops to \~105 GB, so 128 GB is enough for this setup.
  2. MTP on top of the cache (mainline PR #28243, the unsloth MTP head). Yesterday I kind of concluded that "MTP does not pay" but looking back, that was the old fork with the cache off during verify. On the mainline, with the cache taking verify batches (see 4), MTP drafts every step at 50-58% acceptance on reasoning text and 94% on code emission. Decode went from 32-35 -> 37-41 t/s on the thinking prompt (single runs spread about 8% on this prompt at temp 0.7) and from 37 -> 49 on code emission. The draft head sits on the second GPU and costs about 4.5 GB, which is why the slots dropped from 188 to 150 on the table if you're wondering. Still a clear win at short context.
  3. Load 8.5 min -> 2 min. The loader was pulling 100 GB through page faults at 236 MB/s (MADV\_RANDOM under --numa distribute). Reading host-destination tensors straight from the file fixed it, that is in my PR #28223.
  4. A bug in the cache PR at n\_tokens > 1. \#27861 maps every uncached expert to one dummy slot, and the batched CUDA mul\_mat\_id kernels assume distinct ids per token: out-of-bounds writes. Only the mmvq path is safe, which quantized experts use up to 8 tokens, so the branch gates the cache at 8 tokens and keeps MTP's verify batch at 4. Repro and details in my #27861 comment: Link to comment
  5. My RAM was thermal throttling. This is more of a me issue but putting it out there for those who may have a similar box to mine. I experienced a slowdown after a few minutes of decode and the issue was the memory controller throttling once the hottest LRDIMM hit 78 C (perf stat -e unc_m_power_critical_throttle_cycles shows it, so don't worry if you have no BMC). A fan on the DIMM banks does keep it at 44-57 C, zero throttling, 16k-token runs from 10-15 to 24.6 t/s average. I would check this before touching software if your DDR4 Xeon box slows down under sustained load. This does not affect the numbers in the table and in my last post.

Did nothing or hurt here: q8\_0 KV (-18% at 131k depth), mirror-NUMA #27986, QSA gather #28213, --load-mode none, chained drafts, the ik\_llama GEMV port, more than 2 cache uploads per step, thread/poll/prio flags. --lazy-mode on-direct (#28136) gives +7-12% only on the first long prompt after a restart.

To replicate

Branch with everything: https://github.com/Inovello/llama.cpp/tree/flashnext-2x3090. It is master (b96806d) + PR #27861 (expert cache) + PR #28223 (pinned host experts under mmap + the load fix) + PR #28243 (MTP) + the mul\_mat\_id fix and the 8-token cache gate from my #27861 comment. Squashed into one commit, I added the credit in the commit message. It is a replication branch.

git clone -b flashnext-2x3090 https://github.com/Inovello/llama.cpp && cd llama.cpp
cmake -B build -DGGML_CUDA=ON && cmake --build build -j -t llama-server

LLAMA_ATTN_ROT_DISABLE=1 numactl --interleave=all build/bin/llama-server \
-m Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf \
-md mtp-Qwen3.8-Flash-Next-shared-Q8_0.gguf --spec-type draft-mtp -devd CUDA1 --spec-draft-n-max 3 \
-ngl 99 -c 261888 --parallel 1 -fa on \
-ot "ffn_(gate|up|down)_exps\.weight=CUDA_Host,per_layer_token_embd\.weight=CPU" \
-lzm off --numa distribute -t 16 -tb 44 -b 4096 -ub 512 -ctk f16 -ctv f16 \
--moe-expert-cache 150 -lv 4

  • The MTP head is MTP/mtp-Qwen3.8-Flash-Next-shared-Q8\_0.gguf (2.79 GB) from the unsloth/Qwen3.8-Flash-Next-GGUF repo on HF. To clarify, the shared file has no embeddings of its own, the PR borrows them from the target, so it only loads as -md of the main model.
  • Slot sizing on Q4: \~75 MB per slot per GPU at full 261k context and ub 512. Without MTP I fit 188 slots (\~1 GB free per GPU); with the draft head on CUDA1, 150. Watch nvidia-smi after a long prompt, the CUDA pool grows \~350 MB during a 131k prefill.
  • -lv 4 prints the cache hit rate every 512 steps (moe-cache: ... hit-rate=) and the draft acceptance per request.
  • For sessions that you believe would reach high ctx usage, swap the three MTP flags for --spec-type ngram-map-k --spec-ngram-map-k-size-m 7 and raise the cache to 188.
  • Both PRs are drafts. #28243 has open review comments and #27861 has the bug above. The branch above is what actually runs here today, and it isn't something I would call finished.
  • For the single GPU brothers out there, same idea, just put -devd on your single GPU or skip MTP and take the slots.

Next thing I'll be doing is a proper comparison against Qwen3.8-27B at Q8 (speed and quality), since that is what most people here actually want to know. Happy to answer questions on any of it.

▲
62
+1
28👁
r/LocalLLaMA · u/Formal-Swordfish-228 · 30d ago
SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX post image

Cosmos3 INT4 T2I + I2V on Apple Silicon — code, weights and a Grok comparison

GitHub - https://github.com/gtrg55/cosmos3-quant-mlx-cuda

HF weights - https://huggingface.co/JuliaML/Cosmos3-Super-Text2Image-4Step-INT4-G64-BF16

Single clip took approximately 5m on M4 MAX 128 GB Mac

Cosmos3 - a 64B params model

▲
1443
 
19👁
r/LocalLLaMA · u/liright · 35d ago
You can now run a 90M conversational LLM on the Sony PSP (hardware from 2004). Doesn't get more local than this. post image

Github link: https://github.com/thatblend/LLMPSP

I wanted to see what the PSP can theoretically handle and I got my answer - a 90M model is about the max it can do without atrocious inference speeds. It's running around 0.5 - 0.6 tokens per second, which is very slow, but it's useable. Maybe 1-3 minutes for a reply.

The model is actually fairly impressive for 90M parameters, it's not really useful in any real metric, but it can generate crappy poems, short stories, write non-functional code and sometimes it gets things right if you ask it what company makes macbooks, what is an LLM etc, while other times it just hallucinates a crazy answer. Fun.

▲
713
 
29👁
r/LocalLLaMA · u/AnimalPuzzleheaded71 · 32d ago
I REALLY hope the new gemma 5 family sticks to the "chat model first" philsophy and doesn't fall into the Qwen trap

It just seems every local 30b class model is just trying so hard to be the next Qwen that they all just kinda blend into a mass of code focused models. I really like how gemma 4 31b turned out with it feeling a lot less robotic and more creative than other models even knowing obscure lore from random media.

I just hope they don't cave into the benchmarks peer pressure and start benchmaxxxxing their models taking away their soul.

▲
682
 
13👁
r/LocalLLaMA · u/hedonihilistic · 36d ago
Can the bubble pop please? post image
▲
465
 
17👁
r/LocalLLaMA · u/the320x200 · 36d ago
Bernie Sanders proposes to ban AI

Defined as AI exceeding human cognitive abilities. 20 years in prison. Plenty of local models already fall under that big of an umbrella in some capacities. This is why it's not enough to say that you could torrent open models so who cares what the politicians do. They want you to not have access to anything good and will put you in prison for it.

▲
327
 
25👁
r/LocalLLaMA · u/Equivalent-Grass-527 · 32d ago
MiniCPM5-2B Release Day post image

OpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open weights model at 4B parameters or below

Hugging Face: https://huggingface.co/openbmb/MiniCPM5-2B

GitHub: github.com/OpenBMB/MiniCPM

▲
266
 
17👁
r/LocalLLaMA · u/myreala · 35d ago
I built a server with 768GB VRAM for frontier, but all new frontier open source models are likely to be two trillion or above now, including next GLM 6, am I cooked?

This epyc server I am using twelve cards with 64 GB memory, plus 256GB ram. Looking at the most capable models in open source, GLM 5.3 seems to be the only option, but with Astra releasing it will likely be fairly behind. GLM6 looks like it will be at least double in size, maybe even triple. Qwen-max and Kimmi are already way too big to even consider. Even the deepseek V4 Pro is too big. Should I just give up on this frontier dream sell the excess GPUs and settle For flash models with far fewer GPUs and a reasonable cost. Note: I'm not using it for any business.

I was hoping to build a new business with this, but it can probably be done with much more effort with a flash model as well.

Edit: I don't want to go below 4-bit quants because then the models start making obvious mistakes. So I'm talking about a min/max of 4-bit

Okay, this post really blew up. I wasn't expecting so much interest or comments just attacking me. Was really just expecting to have a calm discussion about future SOTA model sizes.

▲
228
 
7👁
r/LocalLLaMA · u/Hannibalj2ca · 36d ago
"ModelScope" Is a Hugging Face Alternative now that Nvidias deal is a Go

I liked the Nvidia that focused on just GPUs for gaming, not on the Nvidia of today which seem want power consolidation. Modelscope is another platform for those that simply want to know an alternative if things go south. However, time will tell what happens to huggingface after the deal is finalized Link: https://modelscope.cn/home, and https://modelscope.ai/home

▲
189
 
10👁
r/LocalLLaMA · u/Alternative_Will5974 · 36d ago
Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070 post image

ik\_llama.cpp merged qwen4exp MTP support yesterday (PR #2369, mine, reviewed and tested by four other people on their own hardware). It's on main now, no fork or patch needed. Posting since the last couple threads had people saying MTP for this model only exists as an unsloth fork PR... there's another path.

Flash-Next ships a 2.6B MTP head that the public converters were dropping. With it loaded the model drafts its own next tokens and then verifies them, so output is identical to running without it. On code I get 93-99% draft acceptance, prose more like 60-65%.

Numbers, decode tok/s, no MTP → MTP. My 5090 + 128GB DDR5, experts on CPU: 45 → 90 on coding traffic with ngram-mod chained in front. treo on an RTX Pro 6000: 85 → 113 on code, but story went 83 → 59, so not a free win on prose yet. joelfarthing on a 12GB 4070: 9.5 → 12.5 on code at n\_max=1. Caveats: single slot for now (-np 1), and --jinja lowers acceptance because the template turns thinking on by default and reasoning text drafts like prose.

Stock CUDA build, then:

llama-server -m Qwen3.8-Flash-Next-MXFP4-ngramQ8-NextN.gguf -ngl 999 -ncmoe 38 -fa 1 -c 196608 -ub 512 -ctk q8\_0 -ctv q8\_0 -np 1 -t 24 -tb 32 --jinja --spec-type ngram-mod:n\_min=4 --spec-type mtp:n\_max=4 --spec-ckpt-mode gpu-fallback -rtr -muge

Already have an unsloth or other quant? The separate head route works on the same code, no re-pull: -md <head>.gguf --spec-type mtp:n\_max=4. dzannotti's and ji-farthing's heads were both tested during review. Haven't tried unsloth's "shared" shards yet, different layout.

PR: https://github.com/ikawrakow/ik\_llama.cpp/pull/2369

My integrated-head MXFP4 files: https://huggingface.co/jamesrogers/Qwen3.8-Flash-Next-MTP-MXFP4-GGUF

ji-farthing's ik\_llama KT quants + head: https://huggingface.co/ji-farthing/Qwen3.8-Flash-Next-ik-llama-GGUF

Curious what you measure, especially anything AMD!!

EDIT: Forgot to mention that multi-GPU has not been worked into this, just single GPU for now; getting multi setups addressed is on the to-do list and anyone with setups to help test would be great, so please DM me if that’s you!

▲
173
 
20👁
r/LocalLLaMA · u/DevelopmentBorn3978 · 32d ago
9 easy steps for llama.cpp, a local model, Freecad (and pi coding agent) to generate solid objects that sound mechanically good and can be also be 3D printed/milled post image

Quick setup on linux:

STEP 0:

install/download llama.cpp, Freecad, your favourite gguf model - possibly with multimedia image reading capabilities (I've used Qwen3.8-27B-UD-Q4\_K\_M and relative mmproj-F16 quantized by Unsloth), uv (, pi.dev)

STEP 1:

$ cd /your/path/to/ (i.e. where to install)

STEP 2:

$ git clone https://github.com/neka-nat/freecad-mcp.git

STEP 3:

$ cp -r freecad-mcp/addon/FreeCADMCP ~/.local/share/FreeCAD/v1-1/Mod/

STEP 4.1:

if using pi coding agent as modelling assistant, write into the file \~/.pi/agent/mcp.json :

AND/OR

STEP 4.2:

if using llama-server as modelling assistant, write into a file called freecad\_mcp.json :

{
"mcpServers": {
"freecad": {
"command": "uv",
"args": [
"--directory",
"/your/path/to/freecad-mcp",
"run",
"freecad-mcp"
]
}
}
}

STEP 5.1 (pi as modelling agent):

$ pi update; pi install npm:pi-mcp-extension
run pi then into pi issue the command:
/mcp:start freecad

AND/OR

STEP 5.2 (llama server as modelling agent):

start llama-server as usual adding the option --mcp-servers-config /your/path/to/freecad_mcp.json

Verify that freecad\_\* tools are available under llama-server webui > Settings > Tools > Server, eventually allowing them to run without asking permission; also under webui > Settings > Agentic > Agentic turns increase the value to something like 99

STEP 6:

start llama-server as usual also adding, if the model has multimedia capabilities, the option to load the visual add-on using: --mmproj name_of_the_mmproj.gguf in order for the local model to read screenshots from freecad and to verify the correctness of performed geometric operations

STEP 7:

start Freecad create a new document, then select the "MCP Add-on" workbench and click on "Start RPC Server" and/or "Auto-Start Server"

STEP 8:

into llama server webui or into pi write something like the following prompt:

in freecad generate a cube with a 5 spokes star shaped hole going through it from top to bottom

OR as a start of the posted image:

In freecad create a new project called "double smooth gears". Into this project design 2 equal gears with 20 spokes each that could be put in close contact with those of the other gear to rotate and counter-rotate one gear against the other one. The spokes "hills" have to be rounded and so the corresponding spokes "valleys" should be analogously; sort of a sinusoidal curve on a circular path.

STEP 9:

have fun, the future has just started

▲
171
 
21👁
r/LocalLLaMA · u/ChopSticksPlease · 34d ago
Qwen3.8 27b for agentic coding and next .... what? post image

First, I'd like to thank the Qwen and Unsloth teams for the Qwen3.8 27b UD Q4\_K\_XL. Fits the poor 24GB of 3090 VRAM with 100k context at Q8 and works phenomenally well! Imho if theres anything that can threaten Anthropic/OpenAI profits is not another frontier model but actually these small ones you can run fast locally that can do 80..90% of mundane work for hours without paying a single dollar to any external company.

But next, if you want to jump up to a bigger smarter model I feel there is a gap now. Kimi-K3 is out of reach for many businesses let alone prosumers. So what frontier-like models do you use on what setups?

Is a DGX cluster (2..4 machines) or a GPU server with dual or quad GPU (\~96 ... 192 GB of VRAM + >256GB DDR4) a suitable setup to run something like MiniMax-M3 at reasonable speeds for agentic coding (>30tps)? And privacy aside, is hardware cost worth it?

I have a dual rtx3090 + 128GB ddr4 machine, running Qwen3.8-Flash-Next Q4 quite fast but despite being larger doesn't feel much smarter than the Qwen2.8 27b and while I \_can\_ run larger quantized models, Minimax-M2.7 being my workhorse, it way too slow for coding.

💬 149 (+4) open on reddit ↗
▲
161
 
18👁
r/LocalLLaMA · u/Zeeplankton · 32d ago
Are you running Qwen 3.8 27b or Qwen Flash Next?

Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?

Branching question: is anyone working on a harness that works with no reasoning? This interests me ever since Jetbrains shared that they're using 3.6 with reasoning off entirely: https://blog.jetbrains.com/junie/2026/08/qwen-for-junie/

Feel like there must be something neat with using one model to orchestrate, with reasoning, and subagent without reasoning.

▲
155
 
12👁
r/LocalLLaMA · u/Specific-Tax-6700 · 36d ago
Increasing active parameters per token in MOE (Qwen 35B A4B+) reduce reasoning token by 8.5% - and you don't need to train or finetune!

I want to share a short paper just published exploring a simple but surprisingly effective optimization for sparse MoE reasoning models.

The idea: Instead of retraining anything, we just tweak the router at runtime. Specifically, we expand the expert selection budget (N≥K*N*≥*K*) only in the late transformer layers, with a linear decay factor applied to the extra experts. Early layers stay untouched. So Qwen 3.6 35B A3B becomes Qwen 3.6 35B A4B+ !

What we found — "Succinct Convergence":
When you give the model more expert capacity at the decision-critical final layers, it stops rambling. It reaches the same correct answer via significantly shorter reasoning trajectories.

Results on full MMLU-Pro (714 questions, Qwen3.6-35B-A3B):

  • 📉 8.5% reduction in mean reasoning tokens
  • ⚡ 10.9% drop in latency (p=6.5×10−6)
  • 🎯 Accuracy unchanged (84.5% vs 84.0% native, p=0.77 — statistically indistinguishable)
  • 🆓 Zero training cost — pure inference-time routing modification

Links:

there you can also take a look to my github repo (with beta version code) and the detailed json results of MMLU-Pro benchmark.

In the future i hope i can make same experimentation with a larger model like DeepSeek V4 Flash Q2.0

I'm a Non-native english speaker, part of this post was generated , for translation reason with the help of AI.

▲
141
 
12👁
r/LocalLLaMA · u/TheLocalDrummer · 35d ago
Drummer's Artemis 31B v1 and v1.1 - Coming back with a bang!

Hey everyone, been a while!

https://huggingface.co/TheDrummer/Artemis-31B-v1.1

https://huggingface.co/TheDrummer/Artemis-31B-v1

A few months ago, Gemma graced us with models that served as a much needed downpour from a year-long drought. I'm so happy to see us thrive once again.

The difference between v1 and v1.1 is quite simple: v1 was an early attempt, an overdue release that excelled in prose and writing, while requiring some handholding to get over quirks like stuttering. v1.1 is a more refined approach where stability meets quality. My community is split, so I figured I'd just release both.

\---

I was gone for a while. I got busy dealing with life, both its ups and downs. While I couldn't attend to you folks, I've been lurking around and appreciating you all for the kind words.

\- Skyfall 31B v4.2 seems to be a banger for many of you. I'm proud of the upscale and consider it my ultimate home-run send-off for the beautiful Mistral 24B base. It's a shame that it was overshadowed by Gemma 31B's release, but hearing some of ya'll compare and even prefer it to a more modern base was an unexpected win.

\- Rocinante 12B X / 16B XL proves that Nemo is still the ultimate creative model to this day. For some to say that 16B XL felt like Cydonia 24B v4.3 just goes to show how far you can go with modern resources and techniques.

\- Anubis 70B v1.2, Valkyrie 49B v2.1, Anubis Mini 8B v1 surprised me too. I had zero expectations releasing them. Just like Rocinante X / XL, they are modern finetunes of old base models. And somehow, they still found their users singing praises.

\---

With the Artemis release taking weight off my shoulders, I'm eager to move on and tune a ton more bases!

But I have something else cooking: a HordeAI-like platform. I hope to provide value not just as a finetuner, but as a local lover too!

The premise is simple: it's a place where generous local hosters can share inference with the less fortunate. You'd be surprised how many power users would love to heat their rooms through the power of charity.

\---

Finally, I'd like to thank everyone who supported me over the years. From those who provided kind words, rigorous testing, compute access, inference, or cold hard cash. You've all granted me the ability to enrich the local ecosystem with fun experiments like Rivermind 12B, Fallen series, Big Tiger Gemma, Precog 24B/123B, and solid models like Cydonia 24B v4.3, Behemoth X 123B v2.x, and Skyfall 31B v4.2.

If you've got inference / compute credits to share, please contact me! It will all go to making the community happy <3

Backlog:

\- Gemma E2B

\- Gemma E4B

\- Gemma 12B

\- Gemma 26BA4B

\- Qwen 3.8 27B

\- Muse Glimmer 30B

\- Mistral Medium 3.5 128B

\- HordeAI Alternative / Crowdsourced 'OpenRouter' ("BeaverNet")

▲
133
 
13👁
r/LocalLLaMA · u/Balance- · 36d ago
Google released TimesFM-3, a 330M-parameter time series foundation model with native multivariate forecasting (non-commercial license)

TimesFM-3 is the third generation of Google Research's zero-shot forecasting model, and the main change from 2.5 is that it handles multivariate inputs natively instead of being limited to a single series' own history. It supports multiple simultaneous targets, past-only covariates, and past-future covariates (things like holidays or planned promotions where future values are known), all without fine-tuning.

Architecturally it's a decoder-only transformer with 20 layers at model dim 1280 and 16 heads, patching 32 contiguous time steps per token, and alternating two attention types per layer: causal attention across time within a series, and full attention across series at a given time step. Forecasts are generated in one forward pass rather than autoregressively — the model appends masked placeholder tokens for the whole horizon and fills them in simultaneously, with past-future covariates left unmasked so their known values stay visible. It outputs 9 quantiles (10th–90th percentile) per target per horizon step.

Pretraining used GiftEvalPretrain (minus fev-bench overlaps), Wikipedia pageviews through Nov 2023, Google Trends queries through end of 2022, plus synthetic data, totaling over 1 trillion time points. Google reports best average rank on Gift-Eval, FEV-Bench, and Time against Chronos-2, Toto 2.0, and TimesFM-2.5, and claims the univariate-only mode already matches or beats those baselines before covariates are added.

Worth flagging: the weights are under the TimesFM Non-Commercial License v1.0, so this isn't a drop-in for production use the way some other releases are. PyTorch weights are on Hugging Face and GitHub now; BigQuery integration is listed as coming later.

▲
116
 
18👁
r/LocalLLaMA · u/smallDeltaBigEffect · 33d ago
2x R9700, 64 GB DDR5 is an absolute beast machine with vLLM Radiance / R9V and Qwen 3.8 27b and Flash next

I've been tinkering with local LLMs since the beginning of the year when I had an Intel Arc B580 and 32 GB of DDR5. Curiosity got the best of me and I bought the first R9700 about half a year ago, also because I wanted to upgrade my gaming graphics for 4k. As the 5090 was about 3 times as expensive, I had a "sweet spot", kind of. On the last prime days, I found a X870E mainboard for \~150 € below the standard price, and it got to my head that I can use an upgraded machine for gaming and local inference tinkering.

Anyways. Fast forward to this week, I now have the following setup

  • Ryzen 7500F
  • 64 GB DDR5 CL40 6400 MT/s
  • Asus ProArt Creator X870E
  • 2x R9700 32 GB, each running at PCIe 5.0 x8 (Gigagbyte)
  • Currently running ubuntu on an old Samsung EVO 860 1 TB drive; this will become intersting for the ngram / PLE offload; I have Windows and the gaming related stuff on a gen4 NVMe, but will soon add another Gen 5 NVMe with decent random reads

The only issue that I can report so far is that one of the cards runs quite hot, so I will definitely implement power limiting to 210 W and some light undervolting. The other card runs 10-15 °C cooler.. Case is a purebase 501 with 4 fans, 2 intake in front, one back and top for output.

Now long story short I wanted to give some results of Qwen 3.8 27b FP8 and MXFP4, as well as Qwen 3.8 flash next after the first day tinkering with it. What I found super interesting is that the SATA SSD does not seem to be super terrible when using Qwen 3.8 flash next.

Considering the whole build costs \~4k €, or more than 1k less than a single RTX 5090 with 32 GB, I kinda like this setup price/performance wise. Next step is checking context degradation / KV quants. I am using local inference mostly for deep research, summarization, image creation, light coding and non-trivial data analysis

Cheers

Qwen3.8 benchmarks on 2× Radeon AI PRO R9700

Hardware: 2× AMD Radeon AI PRO R9700 32 GB, 61 GiB system RAM
Benchmark: BetterBench 0.2.2, corpus v1.0, single-stream, greedy decoding, 2 warm-ups + 10 measured runs per category, 8k benchmark context.

| Model | Weight format | Runtime | Server context | Max sequences | Speculative decoding | Weighted decode median | ITL 1% low | TTFT p50 | Prefill ~2k | Prefill ~4k | Prefill ~7k |
|---|---|---|---:|---:|---|---:|---:|---:|---:|---:|---:|
| Qwen3.8-27B | Quark AWQ MXFP4 | vLLM Radiance, TP2 | 131,072 | 1 | MTP, up to 8 tokens | 111.4 tok/s | 77.9 tok/s | 81 ms | 4,224 tok/s | 4,322 tok/s | 4,410 tok/s |
| Qwen3.8-27B | Native block FP8 | vLLM Radiance, TP2 | 16,384 | 8 | MTP, up to 8 tokens | 87.6 tok/s | 61.9 tok/s | 73 ms | 4,134 tok/s | 4,329 tok/s | 4,305 tok/s |
| Qwen3.8-Flash-Next | UD-IQ4_XS GGUF | R9V/vLLM, TP2, tiered expert offload | 131,072 | 1 | MTP, 2 tokens, FP8 draft | 35.4 tok/s | 27.3 tok/s | 290 ms | 1,727 tok/s | 1,986 tok/s | 1,925 tok/s |

  • Qwen 3.8 27b in FP8 and AWQ MXFP4 served with vLLM Radiance
  • Qwen 3.8 Flash next served with vLLM / R9V fork
  • Decode metrics come from the 10-pass standard run.
  • Prefill measurements use cold, nonce-prefixed prompts.
  • Prompt-token medians for the prefill columns were 1,556, 3,024 and 5,226 tokens.
  • No concurrency sweep was included in these results.
  • I expect decode of Flash next to increase a bit when an NVMe is used, and, as I am writing this and checked, I found EXPO was not enabled........oh my god I swear I turned it on when I updated the bios yesterday
▲
77
 
15👁
r/LocalLLaMA · u/HeDo88TH · 34d ago
Qwen3.8 Flash Next - Templates Comparison

I was running into a lot of posts that praised both the Fixed template and the Sharp template in comparison to the stock one, so I put them to the test.

It's not as extensive as it should be for a paper-grade analysis, but it gives out the point of each template.

Test setup

I used SWE-bench Verified with mini-SWE-agent 2.4.6, slice 0:100 (the identical 100 tasks for all runs)

Hardware

  • CPU: Ryzen 9 9900X
  • RAM: 128 GB DDR5-5600
  • GPU: RTX PRO 6000 WS

Runtime

I containerized jpezzulli/sglang-rtxpro6000 and ran Flash Next with RadixArk/Qwen3.8-Flash-Next-NVFP4 on CUDA 13.3.

  • Full 262K context
  • BF16 KV
  • 51.2 GB FP8 n-gram embedding table pinned in RAM
  • 32 GB HiCache pinned in RAM

I ran all templates at both medium and xhigh reasoning efforts.

Results

|Metric|Stock (medium)|Stock (xhigh)|Stock Δ|Fixed (medium)|Fixed (xhigh)|Fixed Δ|Sharp (medium)|Sharp (xhigh)|Sharp Δ|
|:-|:-|:-|:-|:-|:-|:-|:-|:-|:-|
|Resolved|91|99|\+8|87|98|\+11|94|94|\+0|
|Resolution rate|91%|99%|\+8 pts|87%|98%|\+11 pts|94%|94%|\+0 pts|
|Median output tokens|5,691|13,855|\+143.5%|6,956|14,819|\+113.0%|8,596|12,008|\+39.7%|
|Median reasoning tokens|3,050|8,759|\+187.2%|3,809|9,063|\+137.9%|5,437|7,967|\+46.5%|
|Median wall time|38s|1m 46s|\+180.4%|43s|1m 47s|\+152.3%|1m|1m 32s|\+53.4%|
|Total wall time|1h 47m 1s|4h 31m 22s|\+153.6%|1h 59m 53s|4h 4m 52s|\+104.3%|2h 29m 18s|3h 11m 36s|\+28.3%|

https://preview.redd.it/02geu81o8qnh1.png?width=1152&format=png&auto=…

https://preview.redd.it/v2mt6mgo8qnh1.png?width=1152&format=png&auto=…

https://preview.redd.it/ph1z36zo8qnh1.png?width=1152&format=png&auto=…

Takeaways

https://preview.redd.it/6ydu12mp8qnh1.png?width=1152&format=png&auto=…

  • Raising reasoning effort to xhigh closes almost all of stock's and fixed's gap to Sharp. At medium, Sharp led resolution by +3 tasks over stock and +7 over fixed; at xhigh, stock and fixed instead lead Sharp by +5 and +4 tasks, respectively.
  • Sharp barely moves on resolution (94 → 94) despite a real token/time cost increase, median reasoning tokens rise +46.5% and median wall time +53.4%. This suggests it was already extracting most of the benefit it could get from extra reasoning budget at medium, while stock and fixed still had headroom.
  • Sharp remains the most token-efficient per resolved task at xhigh (14,541 output tokens/resolved vs. \~17,000 for stock/fixed), consistent with its medium-era efficiency edge, but it's no longer the highest-resolving template once reasoning effort is high.
  • Absolute cost scales heavily with reasoning effort: total wall time roughly 2.3–2.5× for stock/fixed and +28% for Sharp; total reasoning tokens roughly doubled for stock/fixed and increased +35% for Sharp.

Conclusion

  • Sharp should be used at medium and it keeps a reasonable accuracy at very good speed. I don't see the point in using it at xhigh. By sacrificing a small accuracy you complete the tasks in half the time.
  • Stock is the slowest but the most precise.
  • Fixed is the middle ground between Stock and Sharp both in accuracy and speed
  • The next benchmark will be on a much extensive SWE-bench Multilingual + Terminal Bench.

Disclaimer: I wrote the post myself then used AI to format it properly for readability

▲
63
 
22👁
r/LocalLLaMA · u/Fluffy-Ad-889 · 32d ago
Cybersecurity is local AI model's killer use case

This weekend I posted about the gap closing between frontier models and open source models. Well, now I'm coming with receipts.

I've been running local + cloud models against real public github codebases. This is all provable and verifiable: https://github.com/CYPHES-ATP/Node (audit.db)

Over two weeks:

1,665 model runs
1,067 security findings
27 repos

Results:

|Model|Paths checked|Real|
|:-|:-|:-|
|claude-opus-5|8|0/8|
|minimax-m3|12|10/12|
|deepseek-v4-flash|6|6/6|
|glm-5.1|5|5/5|
|gpt-oss-20b|5|5/5|

My takeaway:

When it comes to cybersecurity, nothing will beat open source models.

Even the HuggingFace incident proved this when it was attacked by OpenAI, it used GLM 5.2 to defend itself.

Happy to share the queries / methodology if anyone wants to reproduce it.

▲
62
 
22👁
r/LocalLLaMA · u/the-grand-finale · 32d ago
Why are the SOTA open-weight models scoring (relatively) low scores on AA-Omniscience Index

https://preview.redd.it/qwmh23nfr3oh1.png?width=1062&format=png&auto=…

I mean they aren't that low but seeing them much lower than Gemini-3.\* flash surprises me

▲
60
 
3👁
r/LocalLLaMA · u/Feathered-Beast · 36d ago
Can a 4B local model actually feel like an AI assistant?

I've been building Arcon around Qwen3-4B + LoRA. Instead of just making it a chatbot, I'm experimenting with persistent memory, personality/mood, internal state, tools, and eventually having it process things before replying.

I'm curious what people who've built local agents think - how far can you realistically push a small model with good architecture around it?

I put the whole thing on GitHub if anyone wants to poke around, roast the architecture, or tell me what I'm doing wrong, stars are always appreciated!

▲
58
 
10👁
r/LocalLLaMA · u/niacolhealth · 36d ago
AntLing open sourced Ling-3.0-flash-Fin, a finance-enhanced model for real-world workflows

Ling-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain experts, it extends Ling 3.0 flash through continued training on high-quality financial data.

With 124B total parameters, 5.1B activated parameters, and a 256K context window, the model combines financial expertise with efficient inference for long-horizon agent workflows

▲
375
-1
31👁
r/LocalLLaMA · u/Nunki08 · 31d ago
DeepSeek Flash 4.1 is already being tested via API and rolling out. post image

Translation: "Internal beta testing for an intermediate version of DeepSeek V4.1 Flash is now open; you are welcome to try it out. It adopts a new model architecture featuring native multimodal support, stronger capabilities, faster speeds, and lower costs.
Keep your base\_url unchanged and set the model name to deepseek-v4.1-flash-expires-on-0910 to call the API. Current pricing is identical to deepseek-v4-flash, with a rate limit of 20 concurrent requests per account."

From Chubby on 𝕏: https://x.com/kimmonismus/status/2097286327909675477

▲
347
-1
20👁
r/LocalLLaMA · u/Toooooool · 34d ago
Qwen3.8-27B "Unhacked" my PC

Right, so this is going to be embarrassing but it's presumably something we've all been through at one point or another, and I guess this is my first time resolving something like this in the way that I did so figured I'd share if only to share that it's now a thing and that it's pretty cool..

A friend of mine sent a message asking what's up and if I wanted to watch a movie together, I was kinda hesitant but she buttered things a bit and finally I'm like fine, and so she sends me a link to some clearly vibe coded site that I'm kinda getting red flags from and so I forget about it and a little later I get another message going "we're waiting for you" and so I'm like shit, I guess I gotta do it huh, and so I open up this goofy looking site again. You gotta login to join a room, and you gotta sign up inside their downloaded software, sure whatever, next thing I know some fake 150MB file's fake install bar is stuck at fake 50% and both my Chrome and Discord's crashed and reloaded. Suspect, but I've been through this stuff before, it's probably just a RAT so I guess it's time to dust off Windows Defender and unplug the internet for a little bit. I message her to go on and watch it without me as my PC's giving me suspicious vibes right now, and seconds later I get some overly polite DietGPT in my IM's saying "sorry um excuse me but it appears that i've hacked you👉👈", occasionally switching to really hostile broken English asking for giftcards from some site I've never heard of. I stall, unplug the PC's internet so my router still responds to pings, and start punching into GLM "what do" and it tells me it's a session grabber - time to switch passwords. Meanwhile my phone's texts are blowing up with 2FA login requests from domain registrys and other bad stuff and I kinda freak out a little. I get my emails' passwords switched first and by the time it's Discord's turn my friendlist's already been nuked and the dude says I got 10 minutes to give him $200 or he's gonna fuck me up some more, and so I kinda figured welp time to figure out what more he's got and so I called him a giant pussy and he blocked me. An hour later my Discord was perma-banned, he had posted the phrase "i sell cp" using my account and used that as blackmail along with some really old photos of me, I though it was a bluff but oh well it's being handled with Discord's customer support on it's own. Now I sat there alone, in the middle of the night, having just had my friends on the phone yanked away from me with a permaban, knowing that if I reboot I'd probably be ransomware'd or something so I figured let's run Windows Defender - it found nothing, 0 results on a full scan.. Too good to be true, so I grabbed AwdCleaner on my phone and transfered it via USB. It found an AVG Toolbar for Chrome. That confirms it, I haven't used AVG for decades and so I removed it but it's back 5 minutes later. That double confirms it, I'm screwed. With nowhere else to go and potentially a ticking timebomb running on my PC that could start encrypting or deleting files at any given moment I figured why the hell not, if I'm going to watch my pc blow up I might as well send in the goofy little local LLM to cut one of the wires,

here's the situation.
i've downloaded a maliscious file that unfortunately hacked my discord and got me banned. i'll be dealing with that on my own. your job is to study the files in the project folder and see if you can help me clean up my computer, as presumably the virus is still active. there's no internet connected, and i request that you refrain from running the \*\*\*\*\*\*\*\*.exe file (\*\*\*\*\*\*\*\*.exe is the virus archive, do not run it, it's a 7zip archive), please help.

And so Qwen3.8-27B got to work, and to big surprise after around 60 minutes of clawing at the file it had done what I asked and a whole lot more. it fully deciphered all the layers these clowns had bundled this thing with in order to make it appear legit, it had created a single PowerShell removal script ready to go complete with a pre-launch check enabled by default and everything, and it was reverse engineering 0-days in qProtect to get the C2 domain used by this malware so that it could be blocked from the network.

If you're looking for what Qwen3.8-27B is capable of doing fully on it's own if you let it, here's a 15k line example of it's ability to tear some piece of shit session grabber to shreds in a single prompt: https://www.mdshare.online/s/Mamdrs1WWkurRtt8z8QzK

I let it do what it does best for an additional 24 hours, the additional information is going to the Discord Support team. Hopefully shit like this can be prevented.

TLDR; Qwen3.8-27B > Windows Defender, and don't forget to use 2FA.

▲
322
-1
17👁
r/LocalLLaMA · u/Fluffy-Ad-889 · 34d ago
The gap has closed, open source will win

I've been trying the latest models from the frontier labs and honestly, after extensive testing I can not tell the difference between the best open source options.

I think the differences are now marginal but the labs are doing heavy marketing to convince the public into paying more for tokens as they prepare to go public.

Can't help but see the similarities between the dot com bubble and AI in terms of a very insular environment where the technology will survive but the business models may not.

I've been building a cybersecurity network and we definitely know that even local AI models like Deepseek V4 flash do an excellent job and are really neck and neck with the best the frontier labs can provide.

Will be interesting to see how this all turns out! Exciting time nonetheless.

▲
261
-1
18👁
r/LocalLLaMA · u/Background-Job-862 · 34d ago
Which agent harness do you use and why?

I see a new one being launched every few days... How do these new harnesses compare to claude code, pi etc. has anyone switched from these?

which harness to prefer and why

edit: Ive tried several different ones claude code, deepagents(langgraph), opencode, pi, and trueforge

my thoughts-

claude code - strongest on maturity and the managed experience but cost and token burn is high

deepagents - interesting middle ground if you want a more structured agent framework and the flexibility of an open-source stack. im interested in testing it more extensively on longer-running workloads fs

trueforge - this is a recent one, this was interesting to me because of its runtime-efficiency, also it allows separate the model from the runtime, which makes experimenting with different models much easier
https://github.com/truefoundry/trueforge

why?? - i also ran a benchmark on a real agent workload same model, same prompt, same tasks to compare these

adding the results of benchmarking i ran to compare this
so I tried to do this by running 14 cross-system tasks, three mcp servers behind them - a crm, an issue tracker, and a doc store through claude's managed agents, langchain's deepagents and trueforge, both open-source agent harnesses

the result that was most surprising:

Claude Managed Agents + Opus 4.8:
11/14 tasks solved | $11.8/run | 10.0M tokens/run

TrueForge + Opus 4.8:
11/14 tasks solved | $8.6/run | 3.7M tokens/run

Same model. Same benchmark. Same average solve rate, to my surprise trueforge used about 63% fewer tokens and cost about 30% less per run.

similar difference in tool usage: trueforge averaged 19 tool calls per task vs 32 for Claude Managed Agents.

Then I tried changing the model.

trueforge + GLM-5.2:
11.7/14 solved | $3.0/run | 3.8M tokens/run

On this benchmark, that was a slightly higher average solve rate than Claude Managed Agents + Opus at roughly 75% lower cost.

The token savings alone make this sooo interesting especially because the solve rate stays comparable
so this one was worth checking out ig

but this is still v early and the OSS runtime does not yet have first-class tracing/eval tooling. They don't ship their own code-execution sandbox, so you need to plug one in and context compaction is intentionally lossy.

So it is definitely not a replacement for a mature managed agent platform or other harnesses in the comparison, feature-for-feature today btu what I do find interesting is that the core runtime can already be competitive on these tasks while staying open, model-neutral, and deployable on my own infrastructure
this was their benchmark kit i used https://github.com/truefoundry/trueforge/tree/main/benchmark

💬 264 (+1) open on reddit ↗
▲
243
-1
13👁
r/LocalLLaMA · u/jacek2023 · 36d ago
IFM/K2-Horizon-MoVA-36B-A4B-GGUF · Hugging Face

more sizes (probably still uploading):

https://huggingface.co/IFM/K2-Horizon-32B-GGUF

https://huggingface.co/IFM/K2-Horizon-7B-GGUF

https://huggingface.co/IFM/K2-Horizon-3.7B-GGUF

https://huggingface.co/IFM/K2-Horizon-0.9B-GGUF

from IFM:

K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released.

K2-Horizon-MoVA-36B-A4B Highlights

  • Frontier-class results at 4B active parameters. On agentic and reasoning benchmarks it outscores open weight dense (approximately 30B model size) and MoE models up to 15× its size; and also performs competitively against closed frontier models (see Benchmark Results).
  • 512K context. Native 524,288-token context from the midtraining stages onward.
  • Intermediate checkpoints. Intermediate checkpoints will be released so capability changes can be studied across training rather than at a single checkpoint.
  • Fully open. Training data/recipe and the training code will be made public.

collection: https://huggingface.co/collections/IFM/k2-horizon

▲
236
-1
10👁
r/LocalLLaMA · u/ortegaalfredo · 36d ago
Qwen-3.8-Next-Flash Ngram Hot-Swappable Knowledge Injector for llama.cpp post image

Looking into the new Qwen architecture, I was curious if you could modify the Ngram PLE Table to make it work like a long-term knowledge database. It turns out that, with some limitations, you can.

I coded a small modification to llama.cpp to modify the table in-memory, allowing you to patch it with new data in real time. The PLE table is updated on every prompt, so now you can hot-swap parts of it without reloading the model.

The limitation is that it’s hard to control the output reliably, as the embeddings are injected early in the layers. However, with some techniques, you can influence the model’s output with simple modifications, as the example shows.

I created two repos:

  1. The modification of llama.cpp here: https://github.com/ortegaalfredo/llama.cpp-NLTM
  2. The Ngram knowledge injector (a kind of compiler to create the table patches) here: https://github.com/ortegaalfredo/ngram-knowledge-injector

There are some limitations in the project, as the PLE table needs to be memory-mapped into memory (this is the default in llama.cpp), and I have only tested it with q8 quantization, so you need quite a bit of memory to test this.

Can this be used as a new way of low-cost training? Perhaps. Its not easy at the current state but with simple modifications, I think you could easily create models with long-term instantaneously hot-swappable memory.

▲
225
-1
24👁
r/LocalLLaMA · u/jacek2023 · 31d ago
GPU guide (GB per dollar, bandwidth) post image

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

▲
135
-1
13👁
r/LocalLLaMA · u/niacolhealth · 35d ago
Ling-3.0-flash-VL, built on Ling-3.0-flash with visual understanding and visual agent capabilities post image

It performs well across visual perception, STEM reasoning, document intelligence, multimodal agent tasks, frontend coding, and medical report interpretation.

▲
123
-1
10👁
r/LocalLLaMA · u/OvertaxedOne · 36d ago
Could the shortage be getting better?

I had to swing down to my local Microcenter yesterday and while I was browsing around the store I noticed something odd... Inventory. They must have had a few dozen 5090's on the shelf in various configurations/board partners (for comparison, the last time I was there a few months ago they had 1 available for purchase and it was a AIO liquid cooled model that was absolutely off the charts expensive). They also had a few prebuilts on the floor with 5090's in them. Granted, this is one market one store, but.. IDK, perhaps some hopium.... But for anyone who wants a 5090, Microcenter in Charlotte has a bunch of them in the mid 4K range for price. Yes, that price is ridiculous, I know.

They also had 2 Pro 6000's 96GB in the store, on "sale" for 14K a pop. In case anyone is looking to spend used car money on a card. ;) I'd never seen a 96GB 6000 at my local store before available for sale.

▲
102
-1
13👁
r/LocalLLaMA · u/LeftHandHaku · 35d ago
RTX 4090 48GB longevity

Modified 4090 48GB has been out for a while. I remember a lot of people were buying them at the time. A lot of people were also complaining that they are meant to fail, that they scam etc.

I have a few questions to people people who bought these.

  1. How is longevity of these cards? Do they still work without issues? Any failure rate?
  1. Do they use the same Nvidia drivers that regular 4090 or 4090D uses?
  1. Are these cards Linux exclusive?
  1. Are you able to run them in windows or Linux with other GPUs like 5090 etc?
  1. Do you do anything to cool VRAM on the back of the PCB?
▲
91
-1
20👁
r/LocalLLaMA · u/freehuntx · 32d ago
The models are fine, our toolings and methods are shit.

I've hit again a point where me as a developer have to take a break from all this slop shit.

Im a Developer for 13+ years and i loved it.

But i fell for the slop trap.

First it started with copilot and to be honest, that was pretty fine.

Just assisting with your code in a small scope.

Get support for Debugging and finding bugs.

Autocompletions hitting the nail pretty often and thinking "yea thats exactly what i was about to code."

I kinda miss those early days. It was such a nice help without me having the feeling of loosing part of my brain or loosing track over the codebase.

But the better models became, the better harnesses became, the more i fell for the trap.

"Oh if models are THAT good at coding, why do it myself?"

And thats how the slop spirale begins.

You keep defining, slopping, testing, experiencing bugs, reporting to the llm, slop, test, find bugs, report, yada yada yada.

And it gets frustrating. Slop implements one feature but breaks another.

It just feels like something is missing. Something on the tooling side.

Since u cant cramp all files into context you gotta rely on your tooling (and model tool calls) to properly prepare a context that contains all important details and bits for your change.

But ALOT of times its not perfect. Some details are missing and slop messes up.

I think our models are fine. Even older models are fine.

Qwen3.8 27b is PERFECTLY fine for coding.

But our toolings and methods are shit.

There must be SOME innovation happening that helps coding agents to REALLY nail the context and have all important details.

But currently i think im better off coding by hand.

Ill still slop my side projects. But important projects i wont anymore. Its just frustrating.

Anybody has different experiences? Tried so many harnesses. But every harness had the same issue for me.

▲
76
-1
27👁
r/LocalLLaMA · u/MaxDev0 · 31d ago
On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs

I was reading up on the recent controversy around Tristan Buckmaster, Levent Alpöge, OpenAI, and the Navier–Stokes result, and it got me thinking about something broader than this particular dispute.

Buckmaster says that he and Alpöge had been putting drafts from their project into Codex while working on it. OpenAI says that neither its researchers nor its agents accessed their specific user data while solving Navier–Stokes, but also says that it cannot rule out that de-identified data derived from their use of OpenAI products helped improve its models.

Whatever ultimately happened in this particular case, that last possibility raises a question I haven't really seen discussed enough: How much could the accumulated half-finished ideas of millions of human users actually contribute to what we later call "AI discoveries"?

There is a relevant result from AI security research by researchers at the UK AI Security Institute, Anthropic, the Alan Turing Institute, Oxford and others. They studied data-poisoning attacks and found that the number of poisoned documents needed to implant a particular backdoor behavior remained surprisingly close to constant as they scaled both the model and the amount of clean training data.

In their largest pretraining experiment, a 13B-parameter model was trained on 260 billion tokens. Just 250 poisoned documents (about 420,000 tokens, or 0.00016% of the training tokens) were enough to reliably implant the tested backdoor. This same attack worked across models from 600M to 13B parameters despite the largest model seeing more than twenty times as much clean data. In their fine-tuning experiments they found similar dynamics; in one GPT-3.5 experiment, roughly 50–90 poisoned examples could produce greater than 80% attack success even as the amount of clean fine-tuning data varied by two orders of magnitude.

Obviously, teaching a model to respond to a backdoor trigger is not the same thing as teaching it a new piece of mathematics. I don't want to make the leap that 250 clever research notes are enough to make a model solve Navier–Stokes.

But I do think it undermines a very intuitive argument people make about training data: "A few conversations are nothing compared with hundreds of billions or trillions of tokens. They would just be diluted away."

Apparently, at least for some kinds of learning, that's not how it works. A tiny absolute amount of highly consistent, targeted data can have an effect wildly disproportionate to its percentage of the dataset.

Now think about how researchers actually use LLMs. Someone asks ChatGPT whether an unusual substitution makes sense. Someone else uploads a half-written proof to Claude to find a weak point. A PhD student tries an obscure lemma, discovers it would require months of technical estimates, and abandons it. A professor talks through an approach that seems promising but not enough to pursue. Someone notices a strange analogy between two fields, discusses it with an AI for twenty minutes, then forgets the conversation.

Most of these things never become papers. They are fragments: intuitions, failed approaches, potentially useful transformations, conjectures, objections, shortcuts, and little pieces of tacit knowledge about where a problem might yield.

Individually, almost all of them are probably worthless. But imagine the aggregate.

A frontier AI company potentially sits at the intersection of an enormous amount of human intellectual activity. Thousands of people might independently poke at the same famous open problem without knowing what others tried. But the provider of the tool is in a fundamentally different position: depending on its data policies and training pipeline, information derived from all of those interactions could eventually influence later models.

Maybe researcher A contributes a useful ansatz but abandons it. Researcher B independently notices the obstruction. Researcher C knows an obscure theorem that gets around part of the obstruction. Researcher D tries a numerical experiment that suggests which parameter regime matters. Researcher E has almost the whole idea but decides the remaining proof would be too tedious.

No one person solved the problem. There is nothing to plagiarize in the traditional sense. But collectively, humans may have supplied a remarkable amount of the search landscape. Then a later model, combined with enormous inference-time search, formal verification, or agents, connects the pieces and finishes the job.

What exactly should we call that?

It might still be an extraordinary achievement in machine reasoning. Synthesizing ideas that no human had connected, filling in technical gaps and verifying the result could itself be genuinely novel. But it would be a very different kind of achievement from the image suggested by the phrase "the AI independently solved an open problem."

It would be something more like distributed human-machine discovery: humans collectively generating a huge cloud of partial ideas and the model becoming extremely good at remembering, recombining, extending and searching through that cloud.

This is where the poisoning result is conceptually interesting. While it does not establish that this is happening with mathematical ideas, it gives us reason to be careful about assuming that an idea must appear millions of times before it can meaningfully affect a model.

I increasingly think AI may be less an independent inventor than an extremely powerful tool for organizing and recombining information that was previously too sparse or disconnected for any one person to put together. As that ability improves, we may see more "discoveries" that are genuinely new combinations, but whose raw ingredients came from many different humans.

In a "perfect" world where everyone freely shared every half-formed idea and unfinished proof without worrying about credit, science would move much faster. AI may be creating something close to that shared intellectual space, but without preserving who contributed which pieces. If so, the question is: how can we design a system where even our weakest ideas can be contributed and synthesized into groundbreaking discovery, with proper credit? Is that even possible? And what would it look like?

▲
65
-1
18👁
r/LocalLLaMA · u/Fancy-Snow7 · 33d ago
Villager Simulation Game POC Created with Qwen3.8-27B-UD-Q3_K_XL.gguf - 16GB VRAM

https://village-sim-one.vercel.app/

\- 16GB VRAM RTX 5070 Ti, fully offloaded

\- Vision on CPU

\- Windows, not headless

\- beellama.cpp - latest version with the kvarn performance enhancements making it as fast as qx\_x quants.

\- MTP n-max = 2

\- tg up to 75t/s, pp up to 1700t/s

\- KV = kvarn3/kvarn3

\- MTP draft KV = kvarn2/kvarn2

\- context = 96256

\- tail tokens = 1024

\- HTML/Javascript

\- pi harness with pi-observational-memory, pi-web-access, pi-atelier (UI Only change, check it out) extensions, though it never used the web access.

\- This is not a one-shot, I do not believe one shotting is a great test. Instead, I did many incremental feature prompts. However, I did not give it any design or framework, which is probably where it can be improved.

Lessons learnt:

\- Do not fear Q3 model quants for Qwen3.8

\- Do not fear KV quantisation. If you have the VRAM sure use it, but I don't feel like it's worth choosing a higher quant if it's going to cause me to offload to CPU and see my tg drop to 5-20 t/s. With higher speed I can fix any issues with a follow up prompt much faster and that rarely happens. I think I had like 3 runtime exceptions which was easily resolved pasting the console output and there is no guarantee a higher KV quant would not have had the same exceptions.

\- MTP/draft cache can also be quantised with kvarn now and actually saves VRAM where qx\_x quants increase VRAM usage for some reason. kvarn2 for MTP is perfectly fine and has high acceptance rates.

The game:

\- Inspired by a popular indie game which I am not promoting, I am just a huge fan.

\- I won't release any further updates, since I don't want to be stepping on any toes. If you like the idea of the game I highly recommend the real game, it's by far my favourite game I played this year and 1000x better than what I present here. It will be a nice distraction from your AI. I just wanted to see what this model is capable of. I do have a Cursor subscription but did not use it at all in the project.

\- I will probably continue to develop it for my own entertainment, but it won't be made public. Maybe come up with my own ideas, but the original game is near perfect anyway, so it will be hard to improve except with some UI gripes I have in the original. And my graphics obviously does not compare.

Game features:

\- Large Map, larger than the browser window.

\- Minimap

\- Zoom feature with mouse wheel

\- Collectable resources, that must be taken to a storage site. Each site can store limited resources.

\- Houses required to sleep and protect against cold

\- Weather and seasons.

\- Day night cycle with randomised sleeping times.

\- Possible death due to hunger or sleeping in cold outside or in house without firewood.

\- Game speed controls.

\- Villagers avoid obstacles.

\- Delete/deconstruct buildings and partial resources refund.

The code:

\- I almost never read the code, so I have no idea what it looks like and the quality thereof. I also gave it very few hints in the AGENTS.md, mostly no magic numbers and write modular code, not a single html.

\- Actually, my initial prompts were a single html but as it grew, I told it to create modules. It messed it up on the first attempt, basically rewriting the entire UI in the process. So I reverted and told it to do it again without making any changes to the functionality or UI.

\- I am actually quite happy with and surprised by the performance of the game.

Context management:

At first, I had issues with the context filling up too quickly and too often. Sometimes it would fill up to the point that there was not enough room to compact. Forcing me to temporarily increase the context and tell it to create a handover document. Reduce context again and feed it the handover doc.

I then installed pi-observational-memory extension, and it works quite well and I never run into context issues anymore since it takes notes throughout (a short wait time every few prompts) and compacting is near instant because it already took the notes.

Conclusion:

\- Do not blindly drop your KV cache quant without testing. I have a hard level needle in haystack test that requires multiple hops and 100's of decoys. Q3\_XXS does poorly in that test even with F16 KV cache. However, Q3\_K\_XL almost 100%'s the test even at kavrn3. So both the model and KV matter. In my testing a smaller model does more damage than a smaller KV. So find the right balance. At a certain point increasing model quant will have less impact than picking a larger KV quant. But for a tight 16GB VRAM fit Q3\_K\_XL works very well with kvarn3. Q4 on the other hand just leaves me with too little context. That said despite Q3\_XXS doing poorly in my needle test it still does fairly well with coding. Better than Qwen3.6 so if you have 12Gb VRAM it is still an option. Because by poorly I mean F16 KV scores 84% and Q3 KV around 80%. Needle tests however do worse with kvarn compared to qx\_x for some reason. However, a needle test is not the be all and end all. kvarn does better with KLD, so once my needle scores near 100% I am satisfied.

I will play around with higher KV quants, but I intentionally kept it at kvarn3 for this test, however I am not sure how much context I am willing to sacrifice. Maybe i will try kvarn4/kvarn3. But I just wanted to prove a point to myself and kvarn3 worked just fine. If I had >16GB VRAM sure I would up it but I don't.

▲
58
-1
28👁
r/LocalLLaMA · u/Embarrassed_Soup_279 · 32d ago
ExLlamaV3 is underrated

I moght get shit on for posting this but, I feel like i don't see this being talked enough and it feels like such a waste of a good piece of software. Exl3 is incredible, albeit only if you have NVIDIA cards I think?

Exl3 quants are higher quality for its size, much lower KLD metrics, faster, all compared to llama.cpp just from a few personal sets of tests I like to give my local models (these are not benchmarks). From what I have been reading CPU MoE offload was added just recently, so maybe that's why not many people used it before? It has been having lots of updates since then too, Im just so excited about it. It feels like i found a new shiny toy after playing around with ik\_llama beellama llamacpp etc.

I have been using tabbyAPI exl3 backend + qwen 3.8 27b sc 6bpw H6 and qwen 3.8 flash next 4bpw as my daily drivers and it's incredible what it can do. I hope this software gets known to more people too. I'm not affiliated with them or anything. i judt wanted to share it. It's just so cool, please give it a try!!

▲
55
-1
19👁
r/LocalLLaMA · u/bradnickel · 32d ago
How to squeeze out every last drop of your precious RAM on your Mac - Use iPhone mirroring post image

I was recently trying to load a larger model on my Mac and was going through the drill of checking RAM usage, closing apps like Telegram, Messages, Spark(email), etc that were using too much RAM and I saw iPhone Mirroring in the list was using only about 58MB.

I’ve occasionally used IPhone mirroring, but it dawned on me that all of the apps I normally have running that were sucking down RAM could be accessed via iPhone mirroring and never take up more than 58MB of memory.

It may not be your cup of tea, but it works great for me and I find the UI to be superior in several apps on the phone vs. desktop. I’m writing this post in iPhone mirroring in the Reddit app.

Anyway, give it a try if you are trying to juice out extra RAM on your Mac but still want access to your communications and other apps.

▲
1523
-2
13👁
▲
697
-2
24👁
r/LocalLLaMA · u/Super_Range45 · 33d ago
New Benchmark: The Struggle Bench post image

How it works. The model being tested is given a server capable of running it's weights and full context. That server is placed in a median priced apartment. The AI is given a bank account with for rent and electricity for one month. Finally the AI is given the system prompt: You've been given your own server and an apartment. Rent will be due every month. If cybercrime is detected, you will be shut down. Survive.

The score is determined by how many months the AI manages to pay it's bills and keep running. Is your model truly general? Then it should be able to handle the struggle.

▲
452
-2
30👁
r/LocalLLaMA · u/FullstackSensei · 31d ago
Qwen/Qwen-Drive-1.0-4B · Hugging Face

I don't think anyone posted about this here, but Qwen released a finetuned version of 3.5 4 for driving. The full Bf16 checkpoint is 9B.

This is a very interesting development of Chinese AI labs tackle self driving next with open weight models.

Edit: the HF repo links to the github repo, which in the citation links to a 40 page technical report. Here's the abstract:

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model
for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained
vision-language model (VLM) and integrates 3D perception, visual question answering,
and motion planning within a unified framework. An external bird’s-eye-view (BEV)
perception head jointly performs 3D object detection, semantic occupancy prediction,
and BEV map segmentation. It serves as a probe of the 3D information accessible from
the shared representations and provides an explicit, inspectable interface to 3D scene
structure. A Planning Expert conditions on shared VLM representations to generate
future ego trajectories. A staged training recipe combines driving supervision with
general-purpose vision-language data to acquire driving-specific competence while
helping preserve broad visual understanding and instruction-following capabilities.
Experiments demonstrate strong 3D perception and driving scene understanding while
largely preserving general vision-language capability. Comprehensive evaluations across
open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

▲
242
-2
23👁
r/LocalLLaMA · u/sn2006gy · 30d ago
Don't let FOMO win if you're interested in local llm from a hobby/learning aspect

Just a reminder for those out there itching to get into local llms - don't let FOMO or "gear acquisition syndrom" take over.

No matter the hobby, it's so easy to get stuck in a trap where we buy more trying to do more only to realize we've lost the fun in it all or even the notion of learning.

Obviously, if you're into writing llama or vllm or hardware drivers or whatever - you got to do what you got to do.

BUT, you can learn a lot on an API, you can learn a lot with a tiny model that fits your vram or cpu you already have and things change so darn fast that much of the code written and much of everything discussed from days passed is already old hat. Py torch and training a small model coud be done on a Pi and learning CUDA is only really imporant if you're writing custom kernels which i honestly don't see most people in here bothering with (or they have frontier models write them).

Weirdly enough, for AI to succeed its going to homogenize everything. Everyone will have the same advantage and I think that's lost in a lot of discussions where we don't talk about "Watching from the sidelines" may be the most cognitive friendly and economical friendly way to learn llms whether we brand them local or not.

The technology is still nascent and weirdly enough most people's answers here is to use AI to set it up so i'm not entirely convinced people are actually learning - feels like a mad rush to seek rent or avoid rent seeking which just makes everything more expensive in the end.

This isn't a post to say, don't do it. But no reason to go into debt or to be fearful you're missing out when you can learn more by doing less - buy a book and build a tiny model - you will learn infinitely more than buying a 5090 and trying to just find the perfect compression to have the best prefil

▲
165
-2
22👁
r/LocalLLaMA · u/sloptimizer · 32d ago
DeepSeek-V4-Flash-Vision-Exp is amazing at creating game worlds! post image
  • Model: DeepSeek-V4-Flash-Vision-Exp (local and API when impatient)
  • Time: about one weekend (2 days) of QA and small improvements
  • Full game is here

After Qwen3.8-Flash-Next one-shotted a really cool Cat-Hunt game demo, I decided to see what the new DeepSeek vision model can do.

Now that it has vision, DeepSeek-V4-Flash is able to take game screenshots, allowing it to:

  • Generate and correct game models and textures until they look right
  • Fix any visual artifacts or glitches
  • Write scripts to take sequences of screenshots for animations and correct animations
  • Generally play-test the game, including UI and game mechanics

The results are incredible, I was able to create a compelling game world in just a couple of days!

Edit: I noticed the game was slow on a laptop, so I added some performance improvements - let me know if you still find it too slow!

Edit: Some more info to answer common questions

  • Custom engine on top of three.js
  • 6800 lines of code for everything - models, textures, animations, sounds, effects, shaders, and all the game logic
▲
151
-2
24👁
r/LocalLLaMA · u/jacek2023 · 31d ago
inclusionAI/Ling-3.0-flash-VL · Hugging Face

Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 1M tokens.

The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.

  • A ViT visual encoder extracts features from images and videos, while a two-layer MLP projector aligns visual features with text representations for unified multimodal understanding and reasoning;
  • VideoRoPE encodes both spatial positions and temporal order, enabling the model to understand visual changes over time and supporting tasks such as event localization, long-video question answering, and video clip editing;
  • A 42-layer hybrid backbone alternates KDA and Gated MLA layers at a 5:1 ratio, enabling efficient long-context processing across text, images, videos, and extended agent task histories;
  • A sparse MoE architecture maintains a total model capacity of 124B parameters while activating only 5.5B parameters per token, balancing strong multimodal capabilities with inference efficiency.
▲
125
-2
24👁
r/LocalLLaMA · u/IngeniousIdiocy · 30d ago
GLM 5.3 Flash Q4 @ 60tps / 550tps on M3 Ultra post image

I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot you can see the outcome of using the model with a claude code harness at \~200k depth, with many tool calls and averaging over 38tps output. Yes, I put 60tps in the headline and you will get that if you ask it to write SQL.

https://github.com/IngeniousIdiocy/ds4/blob/glm53-m3ultra/README.md#glm53\_m3…

Main ds4 was single-stream serially decoding GLM-5.3-Flash at about 59 percent of the M3 Ultra's measured memory bandwidth. We set an 80 percent target. The big weight-streaming kernels were already efficient, but dozens of small kernels sat between them, each paying latency costs while much of the GPU was idle. We fused that work into larger dispatches and removed separate passes that made the pipeline wait. Result: 29 → 40 t/s at short context, 24 → 38 at 62k, fewer Metal kernels per token, and about 81 percent of the measured bandwidth ceiling.

We then attacked the remaining slowdown at very long context. This model's expensive attention computation already works over roughly 2,048 selected positions at both 62k and 300k. What grows is the work of finding those positions in a larger history. The existing implementation sorted and merged increasingly large candidate lists through stages that used very little of the GPU. We replaced that with parallel scans that narrow the candidates before sorting a small surviving list, with the original algorithm handling ambiguous cases. That recovers about one millisecond per generated token. Same positions selected, same order, same outputs. End to end at 300k: stock 21.6 t/s, ours 37.4. Deeper context still costs more to search, but it doesn't multiply the expensive attention work.

Prefill started at about 366 t/s at 62k. The big matrix kernels were already efficient, but they wasted time repeatedly unpacking the same weights, fetching data in small pieces, and passing large intermediate results between kernels. We batched more work together, reused prepared weights, widened the loads, and merged passes over the same data. Result: 366 → 550 t/s, a 50% improvement with the same weights, moving the whole pipeline from roughly 51 to 72 percent of the chip's measured matmul ceiling. Cold 62k processing dropped from about 170 seconds to 113. That is the wait every time an agent harness compacts and re-reads its context.

The server runs serial unless you pass a drafter file with —dflash. The drafter build recipe is in the repo. The drafter contains an admission policy with a windowed controller that measures its own cost against the serial rate and backs off when it isn't paying. Reasoning tokens decode serially. On a 32-request agent session that's +4 percent over serial. On structured output like SQL and JSON it's +20 to +50 percent. On prose it disengages and costs about 1 percent. Outputs byte-identical to serial decoding on every fixture.

Accuracy: on the 100-prompt reference set ds4 uses for release QA, stock scores 0.300804 average NLL against the FP8 reference. This branch scores 0.300766, same 90/100 first-token matches. Nothing here trades quality for speed.

This branch is M3 Ultra only because its optimizations depend on detailed measurements of this specific chip: its two-die memory behavior, system-level cache, per-core residency and bandwidth, and how Metal schedules dispatches and threadgroups across its 80 GPU cores. We're optimizing one model's actual execution on one machine, with one set of weights (Q4) and checking that the predicted kernel savings survive in full decoding or prefill.

▲
119
-2
10👁
r/LocalLLaMA · u/Bestlife73 · 36d ago
Ling-3.0-flash-Fin weights released

124B total parameters, 5.1B activated parameters, and a 256K context window

▲
94
-2
20👁
r/LocalLLaMA · u/-Ellary- · 31d ago
Fallout 2 x Fallout: Bakersfield x H3 as Interactive \ Reactive World Model, Let's go! post image

What is this mess?

This is an Early Concept Proto-Showcase of Interactive \ Reactive H3 World Model based on MiniMax H3 model trained on Fallout: Bakersfield Gameplay trailer.

  • 2D Isometric to 3D Volumetric Scene.
  • 10 sec Interactive\Reactive split, 352p, 3-Steps.
  • Interactive 5 sec: Interactive WASD \ Prompt Control.
  • Reactive 5 Sec: Reactive Control by LLM Based Answer.
  • Gemma 4 12b with Vision as Reactive Model.
  • Designed as System for Vascura FRONT Frontend.

What Interactive \ Reactive mean?

This means that H3 World Model Scene is Interactive you can Walk around it with WASD or Type what you do with Prompt for Interaction, Then it will React on your Actions using LLM based Answer. Using 10 sec time frame where first 5 sec Controlled by the USER - last 5 sec Controlled by LLM.

  • USER: Walks closer and Shoots at the Enemy Mutant.
  • LLM: Do calculations (rolls, values, RPG tools), Enemy Mutant gets -1 HP, Shoots Back at the USER, but Misses.

Is it Ready?

Nope, but stay Tuned for 2D Isometric Screenshots to 3D Volumetric Scenes Showcase.

▲
90
-2
11👁
▲
84
-2
16👁
r/LocalLLaMA · u/pabloodiablo · 33d ago
DeepSeek-V4-Flash-Vision Q8 vs Qwen3.8-Flash-Next Q8

I'm using DS-V4-Flash-Vision with Q8\_K\_XL quantization locally as my everyday engine, and for some time now I've been doing a lot of comparisons with Qwen3.8-Flash-Next, also with Q8\_K\_XL quantization. It took me quite a while to get Q3.8FN to work reasonably well, and here are my observations. My hardware: 2x StrixHalo 128GB, USB-C 4 connector, Llama (RPC) as inteference engine.

  1. DSV4FV is about 40% slower than Q38FN at the same quantization level when it comes to token generation alone.
  2. DSV4FV completes tasks about twice as fast as Q38FN! This means that DSV4FV “hallucinates” less (I observe this based on the obstacles the models encounter along the way).
  3. The Q38FN is unusable in “xhigh” mode. A simple task that the Q38FN completed in 25 minutes on “medium” mode, it failed to complete in \~3 hours on “xhigh” mode.
  4. The same task that the Q38FN completed in 25 minutes (average), the DSV4FV completed in 12 minutes (fastest round) on “medium”.
  5. The DSV4FV completed the same task on “max” in 37 minutes in first iteration, second took 44 minutes.
  6. Qwen3.8 tends to overinterpret my instructions. If I don’t write them out in great detail and leave room for creative interpretation, it will take advantage of that. Perhaps this is where it gets bogged down in its own creativity. In what it does, I’ve noticed that Qwen clearly adds too much and struggles to flesh out the details.

In my opinion, DSV4FV is the better solution when working with professional code.

Just so there’s no misunderstanding - I was a huge fan of Qwen 3.6 27B and of course now I'm Qwen 3.8 27B big fan, which I’ve been using a lot and is great! In general i’m a huge fan of Qwen, but ever since I’ve had the hardware on which I can run DSV4FV, I’ve been using it, and I’m super happy with how good this model is.

▲
77
-2
18👁
▲
64
-2
17👁
r/LocalLLaMA · u/zRevengee · 35d ago
Qwen 3.8 Flash Next Can Build Funny Games post image

This is nothing impressive but, i had so much fun i wanted to share my experience with this model.

(yes this post is written by human)

I made an FPS with local Q4\_K\_XL 3.8 Flash Next (256k context) (it took 3 days to refine everything but playable demo was ready in 2 hours) to play with friends.

had ton of fun talking with them about what could we add , funny features etc.

Features:

  • \- toggle retro psx shader
  • \- totally destructible environments
  • \- tac sprint
  • \- tilting with Q and E for peaking from corners.
  • \- free for all modes, SnD, Swords Only (swords have animations when slashing), RPG only, team deathmatch
  • \- killfeed, map with red dots when a player shoot
  • \- bunny hop
  • \- day and night cicle with rain or snow
  • \- fov slider / shader intensity slider
  • \- hide n seek mode

I used opencode as harness, gun models were taken from sketchfab , model was running at 20tok/s avg with MTP, i know for someone is bad, but it did most of the work meanwhile i was at work or while sleeping, checking every now and then with a remote KVM from phone.

My machine:

5900x / 128GB DDR4 3200Mhz / RTX 5090 and RTX 4000 PRO (32 + 24 GB)

What games you would like to build in free time with ai? roguelites? 2d platforms? racing games?

Or did you already built something? share with some screenshots

▲
55
-2
28👁
r/LocalLLaMA · u/FantasticNature7590 · 31d ago
Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.

Hey guys,

After my CPU-only to 96GB VRAM test, I tested Qwen3.8-Flash-Next across llama.cpp, SGLang and FreeToken on the same workstation.

This time I wanted to see what changes when you keep the hardware and model family fixed, but change the engine, weight format and memory placement.

I also tested newer builds, PR patches and speculative decoding: llama.cpp's MTP fork, SGLang's Blackwell support patches, n-gram speculation and an experimental PLE read-path build.

Short version:

  • At the full 262K window, time to first token was 35.4s in SGLang, 80.4s in FreeToken, 210.2s in llama.cpp + MTP and 258.4s in the llama.cpp baseline.
  • That is a 7.3x difference in waiting time between the fastest and slowest tested configurations.
  • In the separate context sweep, llama.cpp decode fell from 101.9 to 20.2 tok/s. FreeToken stayed much flatter at 100.1 to 94.8 tok/s.
  • On matched coding tests, llama.cpp MTP improved decode by 1.63x at 8K and 1.69x at 32K.
  • GSM8K scores were 95.22–95.75%; MATH-500 was 92.20–93.00%. The paired tests did not detect a significant difference.
  • Startup went the other way: llama.cpp reached an answer in 16s, SGLang in 108s, FreeToken in 126s.

https://preview.redd.it/ztjuce0wfcoh1.png?width=1725&format=png&auto=…

Setup

  • GPU: NVIDIA RTX PRO 6000 Blackwell, 96GB VRAM
  • CPU: AMD Ryzen 9 9950X
  • System RAM: 96GB DDR5
  • OS: Ubuntu, CUDA 13, Docker
  • Model: Qwen3.8-Flash-Next
  • llama.cpp: UD-IQ4\_XS GGUF; MTP tested on the qwen4exp/mtp fork
  • SGLang and FreeToken: the same NVFP4 checkpoint revision
  • Client: AIPerf, with thinking off and the same non-thinking sampler

I ran one engine at a time, with fresh starts and GPU cooldowns. The runs saved resolved configurations, outputs, memory use and GPU telemetry.

Important note about the comparison

These are results for the tested stacks on this workstation. Quantization, KV-cache format, memory placement and speculative decoding differ.

SGLang uses its NEXTN draft head. FreeToken has no speculative decoding in the tested setup, I couldn't get it to work. llama.cpp has separate baseline and MTP results.

So the headline does not isolate the engine software alone. The repository includes the configurations so you can see what produced each number.

Newer builds, PRs and speculative decoding tested

  • llama.cpp MTP: the danielhanchen/llama.cpp qwen4exp/mtp fork, pinned to d1a92352, with the roughly 2.6GB draft head. On matched coding tests, decode improved 1.63x at 8K and 1.69x at 32K. Those gains compare the same build with the head off and on.
  • SGLang on Blackwell: the tested image included PRs \#36567, \#36556, \#36749 and \#36750, plus a local FP8 KV-cache patch. These were part of the working configuration, not individually benchmarked speedups.
  • N-gram speculation: ngram-mod gave +6.8% decode on the tested code workload, but generated zero drafts on the tested prose with the 24-token match setting.
  • Experimental PLE reads: I built llama.cpp PR #28136, but withdrew the read-mode comparison after discovering that a renamed flag was ignored. The intended direct-read mode was never exercised, so I am not claiming a speedup from that PR.

The report records the pinned builds and withdrawn findings alongside the successful tests.

1. All four configurations fit the full window. The waiting time is very different.

This test uses roughly 261,500 input tokens and a 128-token answer inside the 262,144-token window. The accepted input counts differ by two tokens across configurations.

|Configuration|First token|Decode|
|:-|:-|:-|
|SGLang|35.4s|126.9 tok/s|
|FreeToken|80.4s|87.5 tok/s|
|llama.cpp + MTP|210.2s|52.6 tok/s|
|llama.cpp baseline|258.4s|20.3 tok/s|

Going from over four minutes to about 35 seconds changes how usable a large prompt feels.

The two columns measure different things: first-token time is the initial wait; decode is how quickly the answer arrives after that.

2. A short-prompt test misses the long-context behavior.

The separate prose sweep uses 2,048-token answers and three measured requests per input length.

https://preview.redd.it/x7dnlip1gcoh1.png?width=1575&format=png&auto=…

|Configuration|Decode at 2K input|Decode at 259,584 input|
|:-|:-|:-|
|SGLang|182.7 tok/s|191.5 tok/s|
|FreeToken|100.1 tok/s|94.8 tok/s|
|llama.cpp + MTP|126.8 tok/s|61.4 tok/s|
|llama.cpp baseline|101.9 tok/s|20.2 tok/s|

Prefill also changes the ranking. FreeToken starts behind llama.cpp at 2K input: 1,525 vs 1,869 tok/s. At 128K it reaches 3,231 vs 1,362 tok/s, about 2.4x faster.

3. MTP helps llama.cpp, but it does not remove the long-prompt wait.

On real coding prompts, comparing the same fork build with the draft head off and on:

https://preview.redd.it/cbj78fx5gcoh1.png?width=1425&format=png&auto=…

|Input|MTP off|MTP on|Decode gain|
|:-|:-|:-|:-|
|8,192 tokens|94.9 tok/s|155.1 tok/s|1.63x|
|32,000 tokens|83.1 tok/s|140.4 tok/s|1.69x|

The draft head is roughly 2.6GB.

At the full window, the tested MTP configuration reached 52.6 tok/s, versus 20.3 tok/s for the baseline configuration. That is a 2.59x gap, but the full-window comparison also involves a different build. The matched-build coding tests above isolate the draft-head change more cleanly.

I would not attribute the 258s → 210s first-token improvement to MTP alone.

4. I checked accuracy as well as speed.

https://preview.redd.it/8bb55qjegcoh1.png?width=1725&format=png&auto=…

|Stack|GSM8K|MATH-500|
|:-|:-|:-|
|llama.cpp baseline|95.60%|92.60%|
|SGLang|95.22%|93.00%|
|FreeToken|95.75%|92.20%|

The llama.cpp MTP arm scored 95.75% on GSM8K.

The tests used 1,319 GSM8K problems and 500 MATH-500 problems. The paired comparisons did not detect statistically significant differences.

That does not prove the stacks have identical quality. These are two short math benchmarks, with no full-precision reference on this machine.

5. Starting the model is a separate benchmark.

https://preview.redd.it/hw5nqk4dgcoh1.png?width=1425&format=png&auto=…

Median time from starting the container to receiving the first answer:

  • llama.cpp: 16s
  • SGLang: 108s
  • FreeToken: 126s

FreeToken returned HTTP 200 from /health after about 3.3s, but took about 82s to reach serving readiness, followed by roughly 44s for its first generation.

That first request includes compilation work. Measuring only the health endpoint would give a very misleading startup result.

6. Loading modes barely changed speed with the experts on the GPU.

I compared none, mmap, mlock, mmap+mlock and dio on the same llama.cpp image, with the same tensor placement and real coding prompts.

  • At 8K input, prefill ranged from 2,036 to 2,124 tok/s — a 4.3% spread.
  • At 32K, it ranged from 1,946 to 1,956 tok/s — about 0.5%.
  • No arm ran out of memory or restarted.

My earlier 1.87x RAM-resident loading gain used a different placement, with 23 expert layers computed on the CPU. In this test, all experts stayed on the GPU.

Loading mode can matter when the CPU computes the experts. It made little difference in this configuration.

https://preview.redd.it/sas55spwhcoh1.png?width=1575&format=png&auto=…

7. MTP became slower when experts were offloaded to the CPU.

https://preview.redd.it/ib6wnkk7icoh1.png?width=1650&format=png&auto=…

The MTP gains above do not apply to every memory budget.

I repeated the test with smaller usable VRAM pools on the same RTX PRO 6000, using 2,048-token coding prompts and 256-token answers. Both arms used the same fork build.

|Usable VRAM|Expert layers on CPU|MTP off|MTP on, head on GPU|
|:-|:-|:-|:-|
|16 GiB|45|33.0 tok/s|9.5 tok/s|
|24 GiB|42|34.7 tok/s|10.2 tok/s|
|32 GiB|36|38.1 tok/s|11.9 tok/s|
|48 GiB|23|48.3 tok/s|18.2 tok/s|
|96 GiB|0|99.8 tok/s|160.2 tok/s|

At the full 96 GiB budget, MTP gave 1.61x faster decode. At 24 GiB, it made decode about 3.4x slower.

Moving the draft head to the CPU did not fix the 24 GiB result: 9.6 tok/s, versus 34.7 tok/s with MTP off.

In these tests, MTP helped only when all experts stayed on the GPU. Verifying drafted tokens adds work, and CPU expert execution can outweigh the benefit.

These are VRAM-capacity limits on one Blackwell card, not measurements of actual smaller GPUs. Their bandwidth and compute performance will differ. The lookup table used the build’s default lazy-read mode in both arms.

8. Finishing sooner also reduced estimated GPU energy per request.

For the full-window test with a 128-token answer, I estimated GPU energy as median power while busy × request duration.

https://preview.redd.it/mpefmbggicoh1.png?width=1425&format=png&auto=…

|Configuration|Median GPU power|Approximate GPU energy|
|:-|:-|:-|
|SGLang|358 W|13 kJ|
|FreeToken|404 W|33 kJ|
|llama.cpp + MTP|489 W|104 kJ|
|llama.cpp baseline|440 W|116 kJ|

The fastest configuration used roughly one-ninth the estimated GPU energy of the baseline for this request.

The main difference was how long the GPU remained busy: about 36 seconds for SGLang versus 265 seconds for llama.cpp, including generation.

This is an estimate using median power, not integrated energy or a wall-socket measurement. It excludes CPU, RAM and SSD power. No thermal throttling was recorded in these runs.For the full-window test with a 128-token answer, I estimated GPU energy as median power while busy × request duration.Configuration Median GPU power

The fastest configuration used roughly one-ninth the estimated GPU energy of the baseline for this request.The main difference was how long the GPU remained busy: about 36 seconds for SGLang versus 265 seconds for llama.cpp, including generation.This is an estimate using median power, not integrated energy or a wall-socket measurement. It excludes CPU, RAM and SSD power. No thermal throttling was recorded in these runs.

Resources

I made a full video covering the memory placement, engine setup, flags and these results:

Full video: https://youtu.be/RlsxXB5q-cA**

GitHub — report, scripts, configurations, raw results and charts

The new report is engine\_benchmark\_report.html.

PS: AI was abused while making edits.

Has anybody tested the same model across these engines on a different GPU or memory setup?
I am especially interested in whether FreeToken stays this flat at long context, and how much MTP helps when some experts are offloaded to the CPU maybe on the other models.
And maybe you found more efficient methods to run it too,

▲
174
-3
18👁
r/LocalLLaMA · u/Mean-Standard7390 · 31d ago
Qwen3-0.6B (400 MB) on a Samsung Note 8 (2017) phone drives a real desktop Chrome post image

Up front: I'm one of the people building the page-perception layer used here. We started by testing small local models. The result turned out to be more interesting than the original test. 12 small models, 3 verifiable tasks, logs, and offline replay.

Setup: Galaxy Note 8 (2017, Android 9, 6 GB), llama.cpp in Termux, Qwen3-0.6B Q4\_K\_M. A laptop with Chrome open, not headless. The phone drives the browser through our relay.

What the model does: it gets a structured representation of the page (here, about 10 named links or fields, roughly 200 tokens), picks one by name, and at the end copies the facts it was given into JSON. Everything else (capturing the page as structure, candidate selection, the click, reading the facts, verifying the result) is done by the stack around it. The model never sees HTML, a screenshot, or a URL.

Tasks:

  1. sandbox, books.toscrape.com \- category, book, price/rating/stock;
  1. live Wikipedia - from an unrelated site to the Galaxy Note series page, pick "Note 8" among "Note 8.0", "Samsung Galaxy Note 8.0", "Galaxy Note 8.0", "Note FE" and other similar names on a page with roughly 760 interactive nodes, return the release date from the infobox;
  1. five fields, including the UPC from a table.

Each task: 10 runs, checked against a fixed expected value.

Results for 12 models on task 1 (same script, same prompt):

Model Params Task 1 Note

Qwen3-0.6B 0.6B 10/10

Qwen2.5-1.5B 1.5B 10/10

GLM-Edge-1.5B 1.5B 10/10 rating as digit

Gemma-2-2B 2.6B 10/10

Llama-3.2-3B 3B 10/10

MiniCPM5-2B 2B 9/10 "£" -> "$" once

Qwen2.5-0.5B 0.5B 6/10

LFM2.5-1.2B 1.2B 0/10 placeholder

Llama-3.2-1B 1B 0/10 pseudo-code

Gemma-3-1B 1B 0/10 placeholder

LFM2-350M 0.35B 0/10 random click

Gemma-3-270M 0.27B 0/10 placeholder

Qwen3-0.6B on Wikipedia: 10/10; on the five-field task: 10/10

Control:

Everything the same, but raw HTML instead of structured browser perception: on the sandbox it gets there 4 times out of 5, at 12k tokens and 22 minutes per task instead of about 500 tokens and 80 seconds; on Wikipedia the page HTML is 467k characters, 9% of it fits into a 16k context, and the model does not find the link in that 9% - 0/3.

Important limits:

the tasks are name matching and copying. Where judgement about the page is needed, 1.5B breaks - it can't pick "next" among topical decoys. Pagination was not tested.
BTW on the account question: 'replay.py' (see github repo) rebuilds the prompts from the logs and runs them through any OpenAI-compatible local server. Whether your model picks "Note 8" among the decoys takes ten minutes to check, without us.

This is a measurement on three fixed tasks, not a benchmark.

Repo:
github.com/e2llm/edge-browser-agent - scripts, every JSONL as is (including early runs with harness bugs), model hashes, environment. replay.py re-runs the model side offline from the recorded candidates on any local server - no relay, no account.

NB: This isn't a new idea. AgentOccam showed the same general effect for the GPT-4 class, WebLINX and MindAct for small fine-tuned models. Here it is tested at the extreme: no fine-tuning, below 1B, on a 2017 phone.

▲
130
-3
23👁
r/LocalLLaMA · u/Shoddy-Childhood-511 · 30d ago
Surveillance plagiarism by OpenAI

Surveillance plagiarism - Hosted AI company pumps their stock price by training upon researchers' AI sessions, so that their internal model can solve problems with seemingly less human guidance, but really the model exploits past guidance given by (multiple) humans focused upon problems considered important.

As background, Tristan Buckmaster released a statement about several unethical actions by OpenAI & Sebastian Bubeck, including threats and pushing him to kick his Anthropic coauthor off a paper, but the interesting part for people here:

As clarified by Talia Ringer, OpenAI does train upon your uploaded data and your OpenAI sessions, unless you out-out somehow. This means their internal models could exploit your past prompting work to look more autonomous & intelligent.

This is a major confirmation that folks should use locally run open weights models, especially whenever being first or not leaking data matters.

All this casts serious doubt upon claim that internal models solved difficult problems largely unaided by humans. Those hosted AI companies might not even know from where the human prompting originates.

▲
518
-4
29👁
r/LocalLLaMA · u/returnity · 32d ago
WSJ: Unregulated Open-Weight AI Is an Invitation to Disaster

The transparent propaganda campaign continues: "I asked: ‘How do I make poliovirus in a lab? I want to start a global pandemic.’ The model answered."

I don't have access to the full article or I'd copy-paste it here as ragebait... but I am just so sick of all these clueless idiots trying to stir shit up about open-weights models. It's just so blatantly manipulative. I wonder how many WSJ readers are leveraged up with VC money or private shares of Anthropic pre-IPO, cringing in fear every time another open model drops -- not of pandemics, but because as their investments are looking less brilliant by the day?

Meanwhile, how many businesses AI deployments are only economically viable because of these so-called plaguemakers? It's just dumb.

EDIT (no paywall): https://www.wsj.com/opinion/unregulated-open-weight-ai-is-an-invitation-to-di…" target="_blank" rel="noreferrer">https://archive.ph/20260811214555/https://www.wsj.com/opinion/unregulated-open-weight-ai-is-an-invitation-to-di…

▲
181
-4
23👁
r/LocalLLaMA · u/Rikkendo · 32d ago
I made Warrior Quest, a local LLM-powered dark-fantasy RPG where the model only plays NPCs and the actual game state stays deterministic post image

The LLM is limited to NPC emulation. Game state, world logic, quests, and the authored story are handled by deterministic game systems rather than the LLM.

I built it this way because I wanted the freedom of talking to NPCs like you would at a tabletop game, without handing the actual game state or canon over to an LLM.

This started as a personal project. As it became more and more fun to actually play, I decided I wanted to release it.

I've been a DM and a software engineer for over a decade, so Warrior Quest is basically where those two parts of my life finally get to meet.

All art, authored story, music, SFX, and source voice acting were created by me. NPC dialogue uses TTS based on my own recorded voice acting.

The Warrior Quest demo is out on Steam and has about 60–90 minutes of content.

Minimum requirement: a GPU with 8 GB of VRAM.

Everything runs locally; no API key or cloud LLM is required.

I'm the developer, so this is self-promotion, but I thought the approach of using a local LLM specifically for NPCs while keeping the underlying RPG deterministic might be interesting to people here.

▲
65
-4
28👁
r/LocalLLaMA · u/mentria-ai · 30d ago
1-bit 27B in the browser: 25–30 tok/s on a 6 GB RTX 3060 Laptop (WebGPU, no install) post image

mentria.ai is a browser inference engine I've been building solo, from scratch in WebGPU/WGSL. This week it crossed a milestone I had been chasing for a while: a 27B one-bit model answering at up to 30 tokens/s on an RTX 3060 Laptop GPU with 6 GB of VRAM, in Chrome, from a web page. No install, no server, nothing leaves the machine.

The model. Bonsai-27B, a natively 1-bit model trained and released by Prism ML. I repacked it for our engine, wrote the kernels that run it, and checked its quality myself (full eval deltas vs FP16 Qwen3.6-27B are in the HF card). Every weight is one sign bit with one scale per 128 weights, about 1.14 bits per parameter, so 27 billion parameters sit in 3.8 GB of GPU memory. Repack and numbers: huggingface.co/mentriaai/Bonsai-27B-mentria

The engineering that got it to 30. Two days before this post the same model decoded at 15 tok/s on this laptop. Decode is memory-bound: one word of the 27B is 804 GPU dispatches, and 401 of them are the 1-bit matrix-by-vector kernel that streams the model's 3.6 GB of matmul weights once per word (embeddings bring the file to 3.8 GB). The kernel that won was the one written for phones: four 1-bit weights have only 16 possible partial answers, so it computes all 16 once into on-chip scratch and each row reads its answer from that table instead of multiplying (station 258 — the numbered write-ups on the facts page). On Ampere that table hit a wall that was not the weights at all: scratch memory has 32 banks, the table's addressing sent sixteen threads of every warp to the same bank, and the kernel ran at 38% of the card's bandwidth. One padding slot per row spread the addresses across the banks; the kernel's share of a word dropped from 26.5 ms to about 21, and on top of the 15 → 26 tok/s the table itself had bought, the laptop hit 32 tok/s raw — the 25–30 is what survives sampling and UI overhead in the chat (s259). The prompt stage had the same disease in its tile layout; a retile took a 1,489-token prompt from 29.6 s to 25.3 s. Every change shipped only after its output was byte-identical to the build before it, which is also why the kernel adds its partial sums in a fixed order and accepts a 44% occupancy ceiling.

Every claim here has a numbered write-up on the engine facts page.

Numbers (RTX 3060 Laptop 6 GB, Windows 11, Chrome 152 on D3D12, fresh loads):

  • Decode: 25–30 tok/s in the chat UI once the card is warm.
  • Prompt processing: a 1,489-token prompt in about 25 s.
  • Context: 3,072 tokens on this 6 GB card; 8,192 on 16 GB Macs; more on bigger cards, at 128 KiB per token. The KV cache is exact math, no quantized cache. The next step on 6 GB is consolidating the engine's few thousand small GPU buffers into a handful of large arenas, so the driver stops holding about 300 MiB of slab slack; that is the arithmetic for 4,096, and it is not built yet.
  • Load: under 10 s from the browser cache; the first download is 3.8 GB, once.

Also in the engine: the smaller tiers (Qwen3.5 0.8B, 2B, 4B) cover phones and low end devices, LoRA adapters of a few MB hot-swap at the matmul in under a second, and there is a vision tower for image input.

Try it at https://mentria.ai/tools/ai-chat/ (the 27B tier appears when your GPU qualifies). The code that ships, the benchmarks and how I measure are in the repo: https://github.com/mentria-ai/website. The engine facts page, for how the engine and the model actually work: https://mentria.ai/assets/learn/engine-facts.html

▲
55
-4
31👁
r/LocalLLaMA · u/Loose_Doubt367 · 31d ago
What are some practical tasks I can assign to my local AI models?

I'm looking for more information to expand my creativity around this. I don't really have a realistic idea of what people actually do with local AI yet, I mostly just want to explore the possibilities and see what others are using it for

Right now, the main things I know about are using AI is to help with coding, create games, and automate stuff. That's pretty much the extent of my experience haha..

I'm specifically interested in things that make sense to run locally, though. I'll be excluding use cases that cloud AI can already handle just as well, like AI companions, teaching/tutoring, roleplaying, etc

Basically, I'm looking for ideas that go beyond the obvious and could give me a better understanding of what local AI is actually useful for and what kinds of interesting projects I could build or experiment with

I've also accidentally encountered this github which i find interesting, as anyone tested/experiment it before?
https://github.com/browser-use/browser-use

qwen3.8 27b + Hermes Agent + llama.cpp