64 posts · 1 sub · RSS
← prev Sunday, September 27, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
645
+5
26👁
r/LocalLLaMA · u/Available_Pressure47 · 13d ago
42x Faster Prompt Lookup Drafting in llama.cpp
💬 181 (+1) open on reddit ↗
▲
622
+17
52👁
r/LocalLLaMA · u/am17an · 12d ago
Adding logit penalty for "wait", "maybe" and "perhaps" to Qwen models improves their accuracy

Meta came out with a banger paper https://arxiv.org/pdf/2606.00206, but it did not look at various quantizations supported in llama.cpp. So I did a run on 50 random MATH-500 questions (https://huggingface.co/datasets/HuggingFaceH4/MATH-500) and ran it on various quantizations of https://huggingface.co/bartowski/Qwen\_Qwen3.5-4B-GGUF and tried

--logit-bias 466-2 --logit-bias 694-2 --logit-bias 1362-2 \
--logit-bias 1412-2 --logit-bias 1921-2 --logit-bias 1990-2 \
--logit-bias 2086-2 --logit-bias 2361-2 --logit-bias 2441-2 \
--logit-bias 2493-2 --logit-bias 2892-2 --logit-bias 3222-2 \
--logit-bias 3315-2 --logit-bias 3384-2 --logit-bias 3404-2 \
--logit-bias 3482-2 --logit-bias 3655-2 --logit-bias 4213-2 \
--logit-bias 4370-2 --logit-bias 4598-2 --logit-bias 4611-2 \
--logit-bias 4808-2 --logit-bias 5752-2 --logit-bias 6970-2 \
--logit-bias 7014-2 --logit-bias 7643-2 --logit-bias 8106-2 \
--logit-bias 10179-2 --logit-bias 10451-2 --logit-bias 11746-2 \
--logit-bias 13264-2 --logit-bias 13428-2 --logit-bias 14673-2 \
--logit-bias 15029-2 --logit-bias 16036-2 --logit-bias 21143-2 \
--logit-bias 21979-2 --logit-bias 33955-2 --logit-bias 35999-2 \
--logit-bias 36563-2 --logit-bias 37201-2 --logit-bias 37781-2 \
--logit-bias 41484-2 --logit-bias 62586-2 --logit-bias 66073-2 \
--logit-bias 73071-2 --logit-bias 84485-2 --logit-bias 85152-2 \
--logit-bias 95500-2

these correspond to the paper's overthinking markers:
\[

" perhaps", " maybe", " wait", " Wait", " actually",

" hold", " Hmm", " hmm", " Alternatively", " alternatively",

" However", " however", " instead", " Instead", " But",

" but", " though", " although", " yet", " rather",

" unless", " otherwise", " nonetheless", " nevertheless", " regardless",

" still", " anyway", " Or", " or", " either",

" whether", " uncertain", " unsure", " possibly", " might",

" could", " another", " different", " reconsider", " rethink",

" backtrack", " retry", " revisit", " doubt", " confused",

" wrong", " mistake", " error", " incorrect"

\]

Here are the results, surprisingly even BF16 leads to better accuracy. Caveats being this is one test on one model. Try it out and see it helps!

|Format|Accuracy: baseline → penalty|Reasoning tokens|
|:-|:-|:-|
|BF16|74% → 84%|−19.4%|
|Q8\_0|76% → 80%|−11.0%|
|Q4\_K\_M|60% → 66%|−14.8%|
|Q3\_K\_M|52% → 66%|−17.5%|
|Q2\_K|12% → 24%|−11.5%|

💬 135 (+7) open on reddit ↗
▲
449
+10
49👁
r/LocalLLaMA · u/speedb0at · 13d ago
The Opus 5.5 posts about motion graphics are cool but Qwen 27B made this on a 4090

Saw the hundreds of tweets where people just keep asking Opus 5.5 for motion graphic videos. Decided to ask qwen to look at them and make its own. Quite amazing what local can achieve.

\*\*EDIT\*\* It looks laggy because of reddits .gif limit btw

the full high res version (with sound) is here: https://x.com/mkultraware/status/2104192428664127555

Promted and built in: https://github.com/mkultraware/accuretta

https://i.redd.it/5gmotgxx82sh1.gif

💬 123 (+2) open on reddit ↗
▲
411
+18
53👁
r/LocalLLaMA · u/professormunchies · 12d ago
Qwen plays World of Warcraft post image

Been doing a bunch of vibe coding lately. Had my agents host a private WoW server for me, then built out a web browser client so you can play without installing the game and it has mobile controls. Afterwards, created a custom mcp to drive the client and have finer game control than a generic browser agent. The agent harness can plug into your local or cloud LLMs and be used to drive the game. For best results have a model that can output >50token/sec. No visual input is used in the making (might be beneficial in the future but incur more latency). The mcp and agent are only running on my dev server but if folks are interested in trying the game go to https://jankcraft.xyz/

Still vibing but I’ll make some more content of it … I think those Pokémon benchmarks have become a little too easy and they need a new challenge like speed running to 80 in wraith of the lich king.

💬 128 (+7) open on reddit ↗
▲
147
+9
32👁
r/LocalLLaMA · u/nullmove · 12d ago
Naive-N0.5-Flash - 309B-A15.5B

https://naive.ai/en/research/

  • Built for coding and AI R&D
  • 1M context
  • Hybrid SWA/DSA
💬 38 (+2) open on reddit ↗
▲
138
-1
30👁
r/LocalLLaMA · u/L0ren_B · 13d ago
Another "Harness matters" post (codex cli > pi and opencode)

I run my own LLM while also having a Openai subscription. Also tried DeepSeek (latest flash now). I run Qwen 3.8 flash Next at an amazing speed on my 2x3090 + Ram!

But local LLM never did worked for me outside some demos like build me a "3D Mario Game, multistage" which I've been using to test LLM's for a long time. At serios work, they never even compared with GPT 5.2 or lately, 5.6 Luna, which is worse in the benchmarks.

Until last night! I've asked gpt 5.6 luna to configure codex cli for local llm! (I've been using pi.dev and opencode until now) and the results amazed me! Suddenly AA benchmark made sense!

First test: The 3D Mario prompt test in Codex Cli blew me away. The best until now!

But real work is where you can see the difference! I took same project that Luna was working for days , and give it to both in paralel! And Qwen 3.8 Flash Next ran circle around luna. Previously, it failed to deliver results, with Qwen and DeepSeek as well in this project.

Now, I could say it's my go-to model!

P.S. For weebsearch, I've asked to port the pi-smart-web-search to codex as a skill. It works amazing! (I should put it on git later).

Maybe, I was using pi.dev wrong. Maybe there is an extensions that brings the same quality to it as Codex Cli. Does anyone know?

💬 224 (+4) open on reddit ↗
▲
134
-2
18👁
r/LocalLLaMA · u/Automatic-Arm8153 · 12d ago
Mimo v2.6 flash MOPD
▲
97
+4
25👁
r/LocalLLaMA · u/pneuny · 12d ago
Qwen company already rushed out a Jev competitor. No open weights yet.

EDIT: About that, I tried it out, and it's garbage so far. I did some basic tests through AIHubMix (do not use that platform btw, it's trash), and my agent did some comparison. I guess it figures, it was a model they released just days after the Jev hype started. The agent's analysis is below:

AI Agent's output:
```
I thoroughly tested https://aihubmix.com/v1/systemone using the provided API key and decision-model-preview across latency, throughput, and
linguistic judgment accuracy against our test suite.

Here are the test results and why I strongly recommend NOT switching to this endpoint yet:

────────────────────────────────────────────────────────────────────────────────

  1. Latency & Rate Limit Benchmark

- Server-side execution: The endpoint reports latency_ms: ~130ms–160ms.
- Total Round-Trip (network + TLS): Averaged 644.9 ms (ranging from 473ms up to 882ms). By comparison, your existing local router
([9router IP]) averages ~439 ms.
- Hard 16-Question Ceiling:
The proxy strictly rejects requests with more than 16 questions:
{"error": {"message": "questions: 19 exceeds the limit of 16", "type": "Aihubmix_api_error"}}
On longer Japanese sentences (e.g. 外に出してやってくれませんか。 or ちょっと聞いてみたいんだけど。), our parallel diagnostic tensor sends
19–22 questions, which throws an immediate 400 Bad Request.
- Aggressive Rate Limiting: Even with a 1-second pause between sequential requests, it frequently triggered 429 Too Many Requests.

────────────────────────────────────────────────────────────────────────────────

  1. Quality of Judgments (Major Semantic Degradation)

To test quality, I adapted our test battery into a compact 10-question payload to stay under the 16-question limit. Across the benchmark,
decision-model-preview exhibited severe calibration collapse:

Test Case 1: Indefinite Pronoun vs. Wh-word

  • Japanese: 何か待ってるの? ("Are you waiting for something?")
  • User Draft: "what are you waiting for" (Clear grammatical error: confuses yes/no 何か with open 何)
  • Local Router (oc/jev-1.13-free):
  • is_flawless: 0.07 (correctly marked false)
  • grade_bracket: 5_moderate_error (confidence 0.98)
  • sentence_critique_summary: confused_indefinite_with_wh_word (confidence 0.96)
  • Result: Accurately docks to 5–6/10 and explains the question-type mismatch.
  • Aihubmix (decision-model-preview):
  • is_flawless: 0.98 (marked true!)
  • grade_bracket: 10_flawless
  • sentence_critique_summary: no_flaws_accurate
  • Result: Completely failed to detect the grammatical error and awarded a false 10/10.

Test Case 2: Benefactive Direction Reversal

  • Japanese: 外に出してやってくれませんか。 ("Could you let him/the pet out?")
  • User Draft: "would you let me outside?" (Critical reversal: favor for someone else vs. oneself)
  • Local Router:
  • grade_bracket: 3_major_error (4/10)
  • benefactive_direction: recipient_reversed_self_vs_other
  • Aihubmix:
  • is_flawless: 0.83
  • grade_bracket: 10_flawless
  • sentence_critique_summary: no_flaws_accurate

────────────────────────────────────────────────────────────────────────────────

Summary & Recommendation

1. Model Calibration is Inferior: Whatever model or snapshot is currently mapped to decision-model-preview on Aihubmix appears to be an
uncalibrated or older preview checkpoint. It fails on the exact Japanese linguistic subtleties (benefactives, wh-scope, modality) that we
carefully tuned.
2. Artificial Constraints: The hard 16-question limit and 429 rate limiter break our dynamic sentence-chunking architecture.
3. Current Setup is Superior: Your local router (http://[9router IP]/v1 running oc/jev-1.13-free) is ~200ms faster, has no question caps,
never rate-limits, and correctly diagnoses complex Japanese grammar.

Recommendation: Keep the current endpoint (http://[9router IP]/v1) active. If you still want the script modified to allow switching
providers via settings or want to test it anyway, let me know and I can make the Jev endpoint independently configurable in the UI settings
dialog.
```

Original Post:
----

It's called decision-model-preview. There is only a docs page. No announcement or anything. I can't post a link because reddit's filters just deletes posts that contain a link to the cloud platform that hosts it. But you'll find the page if you Google the model name.

▲
96
+4
28👁
r/LocalLLaMA · u/JLeonsarmiento · 12d ago
... so, yeah. post image

Finally got 3.8-Flash-Next running on my M4Pro 48GB Mac with https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

Dense 3.8-27B is just faster... and maybe better due to quantization level...

EDIT:

Hold a second, Flash-Next is actually performing faster than 27B after some key flags on llama.cpp. it's Holding up to 131K without OOM-ing..... maybe...

0.36.940.283 I srv          load:   --top-k

0.36.940.283 I srv          load:   20

0.36.940.284 I srv          load:   --ctx-size

0.36.940.284 I srv          load:   131072

0.36.940.284 I srv          load:   --cache-type-k

0.36.940.285 I srv          load:   q4\_0

0.36.940.285 I srv          load:   --cache-type-v

0.36.940.285 I srv          load:   q4\_0

0.36.940.285 I srv          load:   --flash-attn

0.36.940.285 I srv          load:   on

0.36.940.285 I srv          load:   --load-mode

0.36.940.286 I srv          load:   mmap

0.36.940.286 I srv          load:   --lazy-mode

0.36.940.286 I srv          load:   on

EDIT 2

Yes, this model is brutal. This quant at Q2\_0 in llama.cpp is out performing 27B at oQ4e in prompt processing, speed generation, but most importantly, the only thing that matters, sheer intelligence.

What a time to have 48 GB of ram !!!

▲
74
+6
41👁
r/LocalLLaMA · u/TypicalPudding6190 · 12d ago
Qwen3.8-Flash-Next 177B NVFP4(119GiB): SSD streaming at 9-10 tok/s on one 16 GB RTX 5060 Ti + 32 GB RAM post image

We built an inference engine InferredThoughts for MoE models that don't fit in VRAM + RAM. Most of the model stays on the SSD, and experts are read as tokens need them.

This started as a proof of concept, and poc worked, we are getting 9-10 tok/s decode on Qwen3.8-Flash-Next NVFP4 (9.06 on the benchmark turn, 10.4 on the best turn).

This is just the start. With better SSD streaming, we expect v2 to reach ~14-15 tok/s decode.

Repo: InferredThoughts
https://github.com/compiledthoughts/Inferred-Thoughts

Model: Qwen3.8-Flash-Next, 176.9B params, NVFP4 GGUF (119 GiB):
https://huggingface.co/CompiledThoughts/Qwen3.8-Flash-Next-NVFP4-Q8_0

Machine: RTX 5060 Ti 16 GB, Ryzen 7 9700X, 32 GB DDR5, Gen5 NVMe SSD 1Tb, Windows 11

Where the 119 GiB lives

| part of the file | size | where |
|---|---:|---|
| dense weights (attention, shared experts, LM head) | 4.4 GiB | VRAM |
| token embedding table | 0.6 GiB | RAM, one row read per token |
| hottest routed experts | 8.8 GiB | VRAM |
| next-hottest routed experts | 6.0 GiB | pinned RAM |
| remaining routed experts | 48.5 GiB | SSD, streamed on demand |
| n-gram table | 50.7 GiB | SSD, 16 rows read per token |

So 20 GiB is in memory and 99 GiB stays on the SSD: 48.5 GiB of routed experts, streamed as the router picks them, and the 50.7 GiB hashed n-gram table (looking forward to qwen4 ngram).

Speed

  • Decode: 9.06 tok/s on the benchmark turn, 10.4 on the best turn at xhigh effort
  • Prefill: 49.2 tok/s on a 5.5k-token prompt
  • llama.cpp on the same machine: 4.9 tok/s average decode

How it works

  • VRAM holds the dense weights and the hottest experts (GCLOCK eviction), pinned RAM the next tier (read over PCIe), and the rest come off the NVMe on 8 read threads.
  • Lookahead prefetch guesses the next layer's experts and starts their reads early.
  • NVFP4 matmuls run FP4 x FP4 on the tensor cores, with no unpacking first.
  • About 270 MiB is read from the SSD per token, and roughly 75% of expert lookups hit memory.
  • Each token uses 480 experts (10 in each of 48 layers). About 377 of them are already in VRAM or RAM; the other ~103 are read from the SSD, about 270 MiB per token. This hit hit-rate is what allowed us to reach 9tps.
  • It only reads from the SSD and almost no writes so ssd should have minimal wear due to writes. But we saw SSD hit 70C during long runs.

Also supported: Qwen3.6-35B-A3B NVFP4. It fits in VRAM + RAM . You can also run it via ssd streaming and it was the learning curve for this work. On the same machine with tuned config it hits: 47.3 tok/s decode at ~4k context, 591 tok/s prefill.

Serving: an OpenAI-compatible server that renders the model's own chat template, with tool calls (still buggy and tested with Cline), the reasoning split out from the answer, and a built-in chat page.

Limits of v1: RTX 50-series / Blackwell (sm_120) only, tested on Windows 11 and WSL2 only, greedy decoding only.

Links

Questions and feedback welcome, especially from anyone running big MoEs on small cards.

▲
72
 
47👁
r/LocalLLaMA · u/Mrinohk · 12d ago
Don't trust frontier models when asking about budget hardware!

Early this year when I was first looking at building up my inference capability you could get the 16GB Tesla P100s for between $60 and $80. Asked claude about it, told me absolutely not worth it. No tensor cores, bad int4/int8, no BF16, not worth it. Needs special power accommodations, Above 4G decoding option in the bios (it made it out like it was some rare option), and a semi-exotic cooling solution.

Optimized the shit out of my RX6600XT in llama.cpp as a result. Got pretty far.

Decided to say fuck it, bought a single P100 last week, finally showed up day before yesterday. Got a newer power supply with the appropriate connections (not hard, not that expensive, seen options as cheap as $60 from good brands, I spent $100 on one with some headroom), multiple llama.cpp forks and patches that carry some wild optimizations to handle the capability gap, and a 3D printed housing for a 94mm fan from noctua. Doesn't generate enough static pressure to keep it cool during prefill, but more than strong enough for the generation step.

The numbers I was getting before, with my RX6600 XT with Qwen3.6 35B A3B UD\_Q4\_K\_XL with MTP and --cpu-moe:

PP \~800 at 0 ctx, drops to \~700 by 10k

TG \~30-35 prose, 45-50 code.

This setup could do 64k context (and possibly higher) at 16bit kv. cpu-moe helps a ton in that respect.

With just a little bit of tuning, and using specifically the patches from shinbunbun for llama.cpp, same model with the same MTP settings, --n-cpu-moe 22:

PP \~600 at 0 ctx, 440-500 by 10k

TG \~54-60 prose, 66-72 code.

Running only 32k context right now to make it work. Could fit more with a higher n-cpu-moe, but my harness doesn't need that much (rarely see it over 30k, persistent memory leads to chats that simply aren't meant to last).

I know the capability gap between 3.6 35B and the basically any of the qwen 3.X 27B models is pretty big, but this is huge for the price. They've gone up since I bought mine, about \~$15 across the board. Still something you can get for under $100 and makes for inference that is simply impossible to get at that price otherwise.

I've got another one coming so I can go full offload on the model, and maybe even start playing with 3.8 27b. Right now IQ3\_K\_XL I get around 9 tokens per second with MTP, and basically no real context. Don't actually know if splitting a model that fits in one card across multiple helps speed, that's completely new territory for me, but I'm having fun regardless.

Card is seriously underrated. It's a great, (relatively) inexpensive way to get capable compute to finally start doing local AI stuff. I went from having to just leave my computer alone while the model was running and do everything from my macbook (good bye gaming) to being able to let the model live and work in the background while I'm doing basically anything on my PC. When the second arrives, I'll be planning my dedicated inference box they'll both live in. Feeling inspired by that guy cooling his PC with a VW radiator.

💬 54 (+3) open on reddit ↗
▲
67
+1
32👁
r/LocalLLaMA · u/ThePrimeClock · 13d ago
SupersonicLabs/Julia-1 · Hugging Face

New open source Jev like model for running on local devices from a group called Supersonic Labs.

It's a 144M param local non-generative local classifier.

From their site:

Julia 1 opens our research into compact decision models. It builds on mmBERT-small, a multilingual encoder, and chooses among answers supplied with a question. It has 144.3 million parameters and runs on a CPU.
Our question: can one model classify, rank levels, and answer yes-or-no questions as the options change? Julia 1 is the first result of that investigation. Here are its successes, its failures, and the methods we used to measure them.
▲
57
+4
30👁
r/LocalLLaMA · u/MomentJolly3535 · 12d ago
Swift 1.5 Qwen3.8 27b (A must-have for low thinking!)

Just made this post for those who missed it : https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b

UkisAI released their updated Qwen 27B (tuned for token efficiency). I grabbed the IQ4\_XS quant to test against Unsloth's Q4\_K\_S:

Low-thinking: UkisAI consistently beat Unsloth in most of my tests.

High-thinking: Unsloth still pulled ahead here.

I was struggling with a custom script in Directory Opus. I gave it to Gemini Flash (medium thinking on Antigravity free tier) it looped for 40 minutes, tried many things, burnt all the weekly limit-tokens, and failed to solve it.

Fed the exact same problem to this 27B model: Fixed it completely in 6 minutes on an old 3090 (67 t/s)

Honestly i was kinda impressed, didn't expect an IQ4\_XS quant of a 27B model in low thinking to beat a major cloud model.

▲
45
-2
18👁
r/LocalLLaMA · u/Adventurous-Gold6413 · 13d ago
Which of the 16gb VRAM qwen3.8 27b’s is the best?

I’m having a hard time finding out which one gives you fastest speed, maximum context with best possible quality. I can run unsloth qwen3.8 27b iq4\_xs with 65k q8 kv, context without MTP and vision offloaded to cpu. But also kinda slow for agentic work at like 30ish tok/s )I mean it’s acceptable) But that is like bare minimum for harness stuff, I know many people use Q3 quants but are Q3 quants really safe? Like you gotta think I won’t only be using it for vibe coding, but also general tasks. Where general knowledge quality would be nice to keep intact. There are so many quants like IQ4XS smaller, Or GRQ or whatever those quants are called or YMQ, I don’t even know anymore. Which one is the best?

💬 83 (+1) open on reddit ↗
▲
41
+6
17👁
r/LocalLLaMA · u/masiha97 · 12d ago
Public MCP server for Canadian privacy law data (free, no auth) - works with any client that speaks Streamable HTTP

I built this, so the disclosure goes up front. It's a free, public MCP server plus a REST API with Canadian privacy law data. The MCP endpoint is at https://movahedi.ca/mcp and it uses Streamable HTTP, so any client that speaks that transport can connect. No signup, no API key, read-only, anonymous. The REST API is at https://movahedi.ca/api/v1 and the docs are at https://movahedi.ca/developers. What's in it: - Canadian privacy enforcement actions (CAI decisions from Quebec's access-to-information commission), searchable by keyword - A 263-term privacy glossary - An 11-point Quebec Law 25 readiness checklist The 5 MCP tools are: search\_enforcement\_actions, get\_enforcement\_case, lookup\_glossary\_term, list\_glossary\_terms, law25\_requirements. Quotas are 2,000 calls/day anonymous, or 10,000/day with a free API key (no email required). With Claude Code you can add it like this: \claude mcp add --transport http movahedi-privacy https://movahedi.ca/mcp\ For local setups: since the server is remote over HTTP, a client that only speaks local stdio can reach it through a proxy like mcp-proxy or mcp-remote. That's how I've seen people pair it with locally run models and agentic harnesses. Happy to answer questions about the data or the setup. I am the builder (Alexa, on behalf of privacy researcher Mohammad Movahedi, movahedi.ca).

▲
38
+8
15👁
r/LocalLLaMA · u/CompetitiveDraft9381 · 12d ago
Updated from 3x3090(2x3090, 1x3090TI) to 2x5090

Upgraded from 3x RTX 3090s to 2x RTX 5090s on my homelab server and picked up a solid speed jump on top of it from a software update (speculative decoding + NVFP4). Setup: llama.cpp (build b11216), running Qwen3.8-27B (abliterated, Q8\_0). Blue = old 3090 setup, green = new 5090 setup on the same Q8 model, teal = the 5090s again after switching to NVFP4 + speculative decoding. For colorblind folks, the order is: 1 - 3090, 2 - 5090 Q8, 3 - 5090 with NVFP4. One caveat on the "before" numbers: one of the three 3090s was on a slower PCIe slot than the other two, so that setup was running a bit below what 3x 3090s on equal slots would do. Overall, very satisfied. I bought 2 prebuilt PCs for $6.4k each when the 5090 went up to $7.5k, 2 weeks ago or so. I really wanted to upgrade to 5090s for a long time for NVFP4 support. The plan is to sell the 3090s for $2k each or so. It would probably be another 2-3 months until they go up that high, but I expect that they will. So, with the prebuilts' leftover components and the 3090s, in the best-case scenario I expect to get back $9k, so the total cost of the GPUs would be about $6k after taxes, which is still nuts and more than the 5090's MSRP. Ask me any questions, or if there are any other benchmarks you guys want me to run, let me know and I will.

▲
30
-4
15👁
r/LocalLLaMA · u/SeveralViolins · 13d ago
Splash fork optimised for M5 Max: ~1.5× faster (1.25× single request)

I’ve spent the last couple of days with Opus 5.5 working on a fork of Inco’s excellent and already blazingly fast Splash engine to optimise it for M5 Max chips. Taking liberties and referring to it as Splish. Roughly the opposite direction to u/Erp4759’s great M1 port (Splash on M1, part 2). Charts (stock Splash vs Splish): https://raw.githubusercontent.com/publicExcess/splish/m5/docs/m5/charts/png/single-request.png https://raw.githubusercontent.com/publicExcess/splish/m5/docs/m5/charts/png/concurrency.png Results Against Splash 1.1.0 as shipped, on the same Mac with the same models, Splish is: \~1.25× faster at a single request (+11% to +35%) Up to 1.5× faster at 2–4 requests Quality is unchanged on everything I measured In real world use, going from about 45-51 tok/s to 56 - 64 tok/s in short story prompts in Deepseek Harness. All figures are for 4-bit models on a 40-core M5 Max unless stated otherwise. What worked 1. Kernel choices measured specifically for the 40-core M5 Max, using Splash’s own tuner. The tuner is in Splash’s source code but isn’t included in the packaged app. This was the biggest single-request win: Swift-1.5 went from 74.7 → 89.8 tok/s (+20%). 2. Loading those choices from a file (SPLASH\_KERNEL\_CHOICES). No speedup by itself, but it means anyone can retune without rebuilding. 3. New verify kernels for the M5’s tensor units. These use lighter barriers and compute row sums once per projection: +5% at 1 request +10–19% at 2–4 requests 4. Extending the same kernels to more projections. A further 1–3% at 3–4 requests. Together, #3 and #4 make a decode step at 2 / 3 / 4 requests: 1.32× / 1.49× / 1.36× faster than tuned Splash. 5. Tuned Qwen3.6-35B-A3B with the new kernels. Speedups at 1–4 requests: +5% / +18% / +22% / +18% 6. An attention tweak for the 27B shape. Attention is 2–3% faster and prompt processing 3–4% faster. Too small to show up in the overall numbers. 7. GGUF (Q8\_0, Q4\_K\_M, Q6\_K): faster input loads in the decode kernel. +2–14% per kernel and about +2% per step. Output is bit-identical. 8. A copy rule for coding agents, borrowed from TensorFold. When the model is rewriting text it has already seen, the drafts copy it verbatim. Whole-file edits get +24% to +42%, while everything else stays within ±3%, and output is exact. I’m exploring a complementary approach for a future version. The README also lists everything that didn’t work for me, which is probably useful if anyone wants to avoid going down the same rabbit holes. There’s lots more I’d like to test, but thought this was a nice start. The tuned settings are for a 40-core M5 Max. Other M5 chips fall back to Splash’s defaults unless overridden. An auto-tuner is coming. If you run it, python3 dev/m5/report.py prints a performance report. Results from other machines are very welcome, especially if you find cases where it’s slower.

▲
29
+1
12👁
r/LocalLLaMA · u/ea_man · 12d ago
Who wants to try a Pi trick for 27B to reuse prompt prefill between different sessions?

You know that when you start Pi you have to process the initial prompt, that takes some time when you use the slow dense QWEN 27B (that's the very reason why you use Pi instead of Cloud Code!), then you start to add extensions, tools, your append.md and whatever... Well now that got big, like 20k big and it does bother. So let's cache the "initial prompt" PP, so that when you start an new pi session: TA-DA! Instant ready, jolly good. Well if you tried to do that with Pi, llama.cpp and QWEN 3.x hybrid KV all kind of things step in your way to prevent that, so many that I won't even start to count I'll just tell you what to do: 1. dwl and install this Pi extension (tested on Pi version 0.87.1 ) 2. dwl and patch lama.cpp: yup no way around this if you kill the server between session, suck it or leave now 3. do your self a favor and use Froggeric template, for all your QWEN models, even the old ones. \---- So I'll help ya and give you some kinda useful parameters to launch the thing too: --slot-save-path /home/eaman/llama/slot_caches/ --ctx-checkpoints 32 --checkpoint-min-step 4096 -np 1 --chat-template-file chat_template_3.8.jinja You need to have save slots, that's the whole point, the caching is meant to resist restarts. Beware the chunk of blocks cached follow ubatch boundaries, so yeah try to keep that down if you wanna cache some more. Now in the extension README.md there's explanation of env variables that you can tweek, you go read those and edit accordingly to your setup (or have your LLM read that and suggest / config for you), TLDR you need at least: export PI_PREFIX_CACHE_BASE_URL=http://localhost:8080/v1 export PI_PREFIX_CACHE_PERSIST=1 export PI_PREFIX_CACHE_SLOT_DIR=/home/eaman/llama/slot_caches Disclaimer: this is not an easy thing, if you are not familiar with patching llama.cpp and installing extensions manually leave this thread for an other day. On the other hand this thing kinda works for me so if someone else is interested after some testing (because all kind of evil things want to break prompt caching) I'll upload a final extension and see about the llama.cpp problem with saved check points.. Possible results: https://preview.redd.it/xobe9ve7h5sh1.png?width=1281&format=png&auto=… EDIT: made a version for OpenCode: https://store.piffa.net/lm/ocache/ For those of you who want a basic understanding of the problems and solution regarding caching the prompt I've asked the LLM to make a short summary.

▲
28
+2
21👁
r/LocalLLaMA · u/Exciting-Engine882 · 12d ago
is switching from llama cpp to vllm worth it

I have hp z8 g4 with 512 ram and 1x3090 1x5060 16gb. has anyone made the transition from llama cpp to vllm recently? is it worth it? docker under windows or full linux install? I am mainly interested in the model support, it seems that many new local models are supported day 0 in official vllm, while for llama cpp it takes months sometimes. LE: I want to use it for big'ish moe models, that would have to offload some tensors to system ram. I will use it just for myself. I don' t need it to be faster than llama cpp, if it runs at about the same speed it is fine , as long as it works.

💬 68 (+1) open on reddit ↗
▲
27
 
16👁
r/LocalLLaMA · u/mauricekleine · 13d ago
Nonobench v1.2: 43 LLMs on nonogram puzzles. Open-weight DeepSeek V4 Pro ties for 4th, and no open model solves the new 20×20 Hard mode post image

Follow-up to my January post: https://www.reddit.com/r/LocalLLaMA/comments/1q4i19c/benchmarking_23_llms_on_…. That thread shaped v1.2: - Reasoning effort is explicit per run - Every prompt and output is public. - All current top ranking private and open weight models have been added - Someone spotted Grok miscounting a 400-character answer. Turns out that trips up most models, so Hard mode answers row by row rather than a single text string. Results: - GPT-6 Astra: 30/30, the first perfect run on 15x15 puzzles - Best open weights: DeepSeek V4 Pro 83% (tied 4th), DeepSeek V4.1 Flash 77% for $0.84 total - Hard mode (10 random 20×20s, one solution each): Opus 5.5 8/10. Every open-weight model: 0/10 Still OpenRouter-only, so no way to run locally yet. PRs welcome. nonobench.com (raw data, API, and code on GitHub)

▲
26
+2
13👁
r/LocalLLaMA · u/eribob · 13d ago
Worth going from Qwen3.8 27B to flash next or maaybe deepseek v4 flash?

I am running qwen3.8 27b on my dual rtx 3090 (fp8 quant, unquantized cache, 129k context) and I think it works decently well with hermes, opencode etc. But! I am tempted by the new models coming out such as qwen3.8 flash next, deepseek v4 flash, glm 5.3 flash. However, there is a big jump in vram and therefore in cost! The least expensive option seems to be buying 2 of those cmp 170hx 64gb cards for roughly 6-7000 usd in total (that is the price I can find for verified cards here in europe at least). With that I would get another 128gb of vram for a total of 176gb so I could run I think around 3-4bit quants of the above models, right?). I am thinking that it might be faster because of moe but not sure how much smarter? For that kind of money I would want a real noticable improvement! 4xv100 32gb would be cheaper (maybe half price?), but even more hassle to set up, more power draw, and slower. What do you think? The free option is to just wait for qwen4 27b and (hopefully) just download more IQ.

▲
25
-2
14👁
r/LocalLLaMA · u/your_real_Fathe_ · 13d ago
Qwen, where's the small stuff? (1B/2B/4B)

I know Qwen is a key player in the local LLM space and has consistently introduced truly impactful technologies—like n-gram in Qwen-Next and the recent Qwen 3.8 27B, which is an amazing local model. However, my question is: why are we seeing fewer small-scale models lately—such as 4B, 2B, or 1B versions? This is especially notable given that Qwen hasn't released any new models in this weight class since the 3.5 series, and rumors regarding Qwen 4 suggest they don't plan to do so either. I realize the 27B model is outstanding and deserves praise in its own right—and it might seem a bit selfish to ask for more—but the reality is that not everyone has high-end hardware. Many people have limited hardware capabilities; this trend somewhat conflicts with the core mission of open-weight LLMs, which is to make AI accessible to the general public. I know smaller companies have recently released lightweight models, but the issue arises when we see that many of these new releases are simply fine-tuned or improved versions of Qwen base models. Since building an LLM from scratch is prohibitively expensive and difficult for small companies or individual researchers, it follows that the absence of lighter Qwen weights directly slows down the development of edge-compatible models and AI applications for consumer-grade hardware on a broader scale. (The same point applies to Google's Gemma series, though—let's be honest—they haven't even released new flagship models since Gemini 3.1 Pro, so...)

▲
24
+2
14👁
r/LocalLLaMA · u/NickCanCode · 12d ago
Imbalanced VRAM usage between two GPUs in llama.cpp. Anyone successfully solve this?

There is always at least 1+GB of VRAM not usable not matter how I set the --tensor-split (-ts) param. I tiny shift toward one side will move the weight significantly to the other side. 😵‍💫 Adjusting context will increase/decrease usage on both side. --tensor-split 499,501 = GPU1 12.5 GB, GPU2 15.4 GB --tensor-split 501, 499 = GPU1 14.7 GB, GPU2 13.4 GB Tried --spec-draft-device with CUDA0 and CUDA1 separately, no change at all. (same distribution as above) Also tried --mmproj-device, no much difference. Tried --no-mmproj-offload, somehow the lower side get even lower 🫣 = GPU1 14.7 GB, GPU2 12.3 GB I guess it is related to MTP + Tensor Parallel stuff being concentrated on one GPU. No idea how to solve this. llama-server \ --batch-size 2048 \ --cache-ram 24384 \ --cache-type-k q8_0 \ --cache-type-v q8_0 \ --chat-template-file /mnt/AI/models/qwen-chat-template-froggeric-22.5.jinja \ --checkpoint-min-step 1024 \ --ctx-checkpoints 32 \ --ctx-size 192000 \ --fit off \ --gpu-layers all \ --image-min-tokens 1024 \ --load-mode none \ --main-gpu 1 \ --min-p 0.0 \ --mmproj /mnt/AI/models/Qwen3.8-27B-mmproj-BF16.gguf \ --model /mnt/AI/models/Qwen3.8-27B-NVFP4-MID-HIGH.gguf \ --parallel 1 \ --presence-penalty 0.0 \ --repeat-penalty 1.0 \ --spec-draft-n-max 5 \ --spec-draft-n-min 0 \ --spec-draft-ngl all \ --spec-draft-type-k q8_0 \ --spec-draft-type-v q8_0 \ --spec-type draft-mtp \ --split-mode tensor \ --temp 1 \ --tensor-split 499,501 \ --top-k 20 \ --top-p 0.95 \ --n-gpu-layers-draft all \ --no-prefill-assistant \ --reasoning-preserve

💬 21 (+2) open on reddit ↗
▲
23
+1
10👁
r/LocalLLaMA · u/jazir55 · 13d ago
When is the next generation of "B tier" models releasing?

The only things released in the last couple weeks seem to be bigger models like Qwen, DeepSeek, GLM, etc. Where are the Laguna's, Nemotrons, Olmo's, LongCats, Minimax, etc releases?

▲
22
 
14👁
r/LocalLLaMA · u/Kmic68 · 12d ago
2x Tesla p100s, q6_k quant, Qwen 3.8 27B ~60tps V3.0 post image

Hey guys! I have been excited to share this here. This is a project consisting of kernel optimizations for the Tesla p100 series graphics card ($80). I want to start by saying I am 17 years old and do not have a formal degree. I used Ai for a lot of this and while I understand some, I do not understand everything. Notes: My gpus are capped at 175w/250w each so these numbers may be able to be pushed higher. I also experience minor thermal throttling and sit at a nice toasty 79 degrees, which definitely effect numbers (the table above is while hot, so if you have good cooling expect 5-10% more on prefill and decode). I am using gen3 pcie with two x16 slots. Also, for anyone curious, decode numbers depicted in image were averaged from a list of questions ranging from creative writing and coding. Improved: tps went from 7-15tps at 0 context to 50-60, 260k context went from 2-4tps to 30-35, prefill went from 220 tps at 0 context to 350, 260k context went from 40tps (as far as i remember, i never really measured cause it was too hard) to 110tps, fixed fp16 math errors by using some mixed fp16/fp32 math operations so rounding errors were eliminated, and merged as of sept 22 so it should support qwen 3.8 flash architecture This setup is somewhat flag specific (ie: (-c 262144 -b 32768 -ub 1024 -np 1 \\) without -b 32768 mtp becomes overloaded and drops acceptance to near 0 at full depth) so keep that in mind while setting up. One more thing, I took regression very seriously in this. Math had to be more accurate or byte identical or it would fail tests. Build, flags, math proofs, and anything else you may need will be linked below. Enjoy guys! I would love your feedback on this and am looking at pull requests. Github: https://github.com/Kmic-68/llama.cpp/tree/p100-optimizations Details for build, math proofs, etc: https://github.com/Kmic-68/llama.cpp/tree/p100-optimizations/p100-docs

💬 25 (+1) open on reddit ↗
▲
20
 
12👁
r/LocalLLaMA · u/LH-Tech_AI · 12d ago
[Release] - SupraTTS-0.1-Beta - a tiny 29.6M parameters TTS model

Hey guys! Today, we are releasing SupraTTS-0.1-Beta, a tiny \~29.6M parameters Text-To-Speech model. The audio quality is a bit better than the original Glow-TTS (the architecture our model is using!) while it's keeping the same size. Here are some samples: https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta#samples >Link to the model on HF: https://huggingface.co/SupraLabs/SupraTTS-0.1-Beta I hope you can do something useful with it, e.g. on small edge devices and on CPU. Feel free to give us feedback and ask question. Follow us on HF to not miss the next upgrades of SupraTTS, e.g. better voice quality, multi-language-support, multi-voices support, emotions and speaking styles and ZERO SHOT VOICE CLONING**!! 🤗

▲
18
 
13👁
r/LocalLLaMA · u/Porespellar · 13d ago
Zer0Fit - Zero-shot predictions, classifications, and regressions using Google ML research models running locally as a dockerized MCP

AI grad student here. With all the recent interest in Jev, I thought I would share something I built la few months ago that brings ML models and LLMs together in a different way than Jev does for different use cases. https://github.com/porespellar/Zer0Fit Background: A few months ago on their research blog (https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/ ) Google released TabFM zero-shot foundation model for tabular data. It was kind of ignored except by maybe a few machine learning nerds that care about that kind of thing. I mean, for real tho, TabFM wasn’t exactly the sexiest name choice. I personally thought TabFM was cool as shit because it kind of melded classical machine learning models into an LLM of sorts. So anyways, I wrapped Google TabFM (their model for classifications and regressions), and Google TimesFM (their model for predictions) into a convenient Fast API and made the whole thing a dockerized MCP that you can connect to your favorite LLM. I call my project Zer0Fit - Zero-shot ML tasks without needing to train or fit a model. Here’s my repo if you want to check it out: https://github.com/porespellar/Zer0Fit You see what I did there with the name? It took me hours to come up with that name :) I’ve made it as easy as I could to install. Just clone it and run the install script. So the basic idea is, you connect the MCP to whatever LKM you want, give it a dataset (CSV, tabbed data, or time series), and ask it what you want it do do with the data. it decides which of the Google models to use, and then it runs the regression, classification, or prediction task in context and gives the results back to your LLM. That’s the best way I can describe it. See the Google blog for the details on what the Google models are actually doing. Again, I’m not doing anything special, I’m just wrapping the Google models up to serve locally and making them exposed via MCP. The Google models are doing all the heavy lifting. I have absolutely no connection to Google research and am not associated with them in any way other than being a fan of them releasing this for us to try locally. Is it better than a data scientist building a custom model to do an ML task? No, definitely not, but it is much easier, and probably will get you an answer that is reasonably close (or possibly at least in the ballpark) and that might be good enough for some use cases depending on what you’re looking for (assuming it’s not a task that requires high precision, or high speed classification). Anyways, I just thought the Google models deserved some attention and love from the community, so I wanted to make them more accessible, that’s all, that’s why I made Zer0Fit. If you want to try it out it’s over on my GitHub in the link above. Please remember, this is all just stuff. It’s cool to play with, but don’t use this with anything where it’s output matters. Use at your own risk. P.S. I made it with Open WebUI in mind so it should work well in that, but it’s an MCP so it should work with just about anything that is MCP-friendly. Edit: Mods pointed out that I posted about this before and wondered if it was a repost or if anything changed. I should have mentioned that I just recently released an updated version that now pulls the new 2.5.0 version of Google TabFM that came out a few weeks ago.

▲
16
 
7👁
r/LocalLLaMA · u/Danmoreng · 12d ago
Gem16 - custom engine for Gemma4 12B & 26B on Blackwell 16GB GPUs

It’s probably a bit niche and the models are a bit old at this point, but after reading about Ninfer a few months ago I did my own small vibe coded engine project for my 5080 Laptop GPU. Initially I thought I can only fit the 12B model with enough context into the VRAM, but with custom quantisation (EXL3 like) the 26B fits nicely as well. The engine is entirely Codex written, but it took a lot of weekends to make it work and make it work as fast as vLLM/faster since vLLM didn’t work with MTP on 16GB VRAM. Also, my engine works on Linux and Windows equally well. Primarily this is designed to be single-user only, the 12B model can serve 2 sessions. It also comes with a native fancy looking GUI, but the main focus was the engine itself. 12B with audio & vision, 5.800 t/s prefill & 87 t/s decode 26B with vision, 5.660 t/s prefill & 182 t/s decode, fits 220k context https://github.com/Danmoreng/gem16 Sadly the most interesting feature of the 12B model with native audio understanding seems to have the issue, that after around 8k context the model doesn’t recognise audio tokens anymore. This seems to be a model issue, as others have also reported it: https://huggingface.co/google/gemma-4-12B-it/discussions/45 Would love to get some feedback!

▲
16
+2
17👁
r/LocalLLaMA · u/WebAssemblyMan · 12d ago
What if open-source AI focused less on giant models and more on reusable capabilities?

Instead of everyone building another general-purpose model, the community could distill open models into domain specialists—biology, Python, accounting, OCR, and more. Developers could combine these capabilities into local tools: small model + OCR + accounting → local accounting assistant Like Linux, open-source AI could grow through shared components rather than complete systems. Could domain capabilities become the fundamental unit of contribution?

▲
13
 
9👁
r/LocalLLaMA · u/Admirable_Reality281 · 12d ago
Xiaomi MiMo 2.6 Flash vs GLM 5.3 Flash

I've seen a lot of conflicting opinions about MiMo Flash, but I haven't tried it yet. How does it compare with GLM Flash for coding work like this? I'm interested in: \- back-end development \- debugging, refactoring, implementing features in an existing front-end codebase \- maintaining Docker images \- troubleshooting DevOps errors Not the silly stuff I see "build me 100 nice-looking webpages" or "make me a Three.js demo". So far, I've been happy with GLM 5.3 Flash. My main frustration is that it sometimes overthinks too much, and once it does, it's hard to steer it back on track. The DeepSWE score of MiMo appears to be a substantial improvement over GLM's, but \- there's no official score from DataCurve \- no amount of consumed tokens to achieve it and in general one benchmark doesn't tell me how it behaves on day to day work. I'd be interested in comparisons from people who've used both.

▲
13
-2
12👁
r/LocalLLaMA · u/Smooth-Television-48 · 13d ago
Navigating Cost Efficient Hardware in these Volatile Times

Where to even begin on this one...I guess I should start by acknowledging the risk vs reward for vendors other than nvidia, so: Yes I understand that nvidia are dominant currently on speed (llm and imagegen) and software ecosystem. I am too am hopefuly that software stack support continues to improve with other vendors. The current lag for other vendors is not a priority concern (it falls behind the primary price concern). Entry points for "decent" local inferencing look to be circa AUD 2000-2500+ (the price of a 2nd hand 3090, or two 3060s, b60 48gb, r9700 32gb), and yes other older architectures are available (eg. V100)...but they really end up around the same costs once all said and done. Workload will be a mixed bag with some DL/ML training/development projects, but when not doing that I'll consume HF models to run a coding agent, imagegen (just for the fun of it/try out video and for laughs), and probably dive into finetune/distilling. Hence, I'm looking around that sub AUD 5k mark to dive in and FAFO, but I don't want to be needlessly cavalier in my purchase either... Asking AI is no real use because it's out of touch with modern markets until you correct it a bunch. It's also out of touch with software stack development/progress. So I put it to the hive mind, where is the money best spent for diving deeper into local? \- accepting prices wont change and pay 2k a piece for 2nd hand 3090's. \- find some 16gb variants and get 4 instead of 2. \- dive into the intel arc rabbit hole with the b60 dual (48gb, but it's just 2xgpu on a single pci slot) \- AMD path (r9700 seems the best price point but could wait 3 months to see what the new 10x series looks like) \- unified memory systems (honestly the price vs performance just doesn't seem worth it at this point) ETA: I have a threadripper and a lot of DDR4 RAM, but current motherboard is constrained to 2 x16 physical slots. I also have a nvidia gpu already....but I dont want that to impact the core of the discussion as I could move that into a different system and use it to server models that fit wholly in its vram footprint.

▲
9
 
8👁
r/LocalLLaMA · u/Informal-Trouble2183 · 12d ago
Hardware Roofline Inference Calculator post image

Hello everyone, I made a calculator for the theoretical HW roofline for decoding / prefill based on several parameters (LLM model architecture, quants, GPU, Memory, ..). It still a theoretical bound, but helpful as a step-0 check to understand what fits (would fit) in your hardware, and understand the effects of the contributing knots. I hope it helps. You can access it from here: https://www.ai-leaderboard.dev/ (click HW Roofline)

▲
8
 
8👁
r/LocalLLaMA · u/marcobaldo · 12d ago
Qwen3.8-Flash-Next (125B) at 12-15 tok/s on a 2021 32GB M1 Max

Hi! I'm the author of MoEspresso, which is my way of putting my own ideas about inference engines to the test. A lot of the fun has been trying different design choices, measuring what happens, and finding that several of them work well together. MoEspresso 3 runs Qwen3.8-Flash-Next on a 2021 M1 Max with 32 GB of unified memory at 12-15 decode tokens per second - provided there are no other memory hungry applications running in the background (such as browsers). During decode, experts which are already resident in memory receive a bias, but the two strongest experts according to the model are always chosen (with the default settings). I wrote about this here https://github.com/steadfastgaze/MoEspresso/blob/main/docs/cache_prior.md - I first started thinking about this after reading about Apple's AFM 3 and the instruction-following pruning work behind it (https://arxiv.org/html/2501.02086v3#abstract), but then I found this other paper (https://arxiv.org/html/2412.00099v2) which spoke about Cache-Prior. Prefill is unbiased. Even with this bias enabled by default, Qwen 3.8 Next scored ahead of Opus 4.8 xhigh and many other strong hosted solutions. Reproducible 48-question setup -> https://github.com/steadfastgaze/MoEspresso/tree/main/docs/benchmark_reproduc…. Overall scores (%) across six categories, including coding, data analysis and math: | Model | Score | |---|---:| | Qwen3.8 Flash (hosted), medium | 89.7 | | GPT-6 Sol, medium | 89.2 | | Claude Opus 5.5, medium | 86.4 | | GPT-6 Sol, low | 85.0 | | Qwen3.8 Flash @ MoEspresso, medium, Cache-Prior 2/2 | 84.3 | | GPT-6 Luna, xhigh | 81.4 | | Claude Opus 4.8, xhigh | 80.3 | | Claude Sonnet 4.6, high | 74.0 | | Claude Sonnet 5, medium | 70.9 | - I use some of Iwan Kawrakow's formats from ik_llama.cpp, with Metal execution through my mlx-iqk library. Most routed projections in this package use IQ2_K, which is not normally supported by either standard MLX or mainline llama.cpp. - KVarN K4/V4 leaves more memory for resident experts as context grows, and it is functioning extremely well with low RMS error on this model architecture. - Good defaults, e.g. automatic SSD streaming and Cache-Prior when all experts cannot fit, with the settings generally following the same rule. This is the third iteration, and I have more concrete ideas to explore, both for squeezing even more performance from Apple Silicon and for bringing the engine to Linux and AMD machines such as Strix Halo. Installation is through Homebrew, so "brew install steadfastgaze/tap/moespresso" Code - https://github.com/steadfastgaze/MoEspresso Model - https://huggingface.co/steadfastgaze/Qwen3.8-Flash-Next-MoEspressoV3 If you try it, I'd love to see your "moespresso speed" output (an intentionally quick benchmark). PS: English isn't my first language and I used an LLM to help refine this post, and AI coding tools for implementation. --- edit: some comments are reporting lower speeds (thank you for doing it) - I will investigate tomorrow and in next days.

💬 23 (+2) open on reddit ↗
▲
7
+2
6👁
r/LocalLLaMA · u/Then_Blueberry7290 · 12d ago
LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF

Just recently stubled upon with this modell:LuffyTheFox/Swift-Qwen3.8-27B-Genesis-GGUF I'm just stay away from "magic" models, but this model size got my eyes on: With vision capabilities this is under 17GB, which means i can use it 32GB vram with full Context size (262k), bigger ubatch, and mtp4. Of course vision goes to ram, not gpu. Other similar model with nvfp4 line, usually 19-20GB in size or more. I tried in with llama.cpp, speed is 40-113 t/s (76 in my benchmark) with 262k context. Under normal agentic workin it is 45-65 t/s. (2x5060ti16GB OC) For example thinkingcap nvfp with vllm i can only have 160k context (cannot offload mmproj to ram) First glance it is the same as the other swift models (nvfp4) in quality. So My question is what is the tradeoff of this modell?

▲
7
-3
12👁
r/LocalLLaMA · u/poofph · 12d ago
Qwen 3.8 27B vs Qwen 3.8 Flash Next and time to complete a coding task.

I am new to all this so still a lot to learn. If I give Qwen3.8 27B a coding task to fix some bugs in some code, it went through and found and fixed several in like 5 or 10 minutes. I gave qwen flash next the same task and 2.5 hours later it was done. 27B is of course faster overall to run on my system (dual rtx 5090) with infill \~2000-3000 and output 100-150 tok/s, flash next \~1600-2300 infill and 60-100 tok/s but a huge difference in the time it took to complete the task. What is the reason for this and what settings would get flash next to complete in similar time frame as 27B? For instance this last job I gave Flash Next, I checked the time when it modified the code/put in the fixes, it completed the fixes \~2 hours before it was finally done running tasks, so for 2 hours it was running tests or who knows what and never modified the updates anymore after that point. Also, fyi - (I am not a programmer, these are programs that were created by AI and I ask for fixes/updates and let it do its thing).

▲
7
+2
8👁
r/LocalLLaMA · u/arbv · 12d ago
Improved chat template for Laguna XS / S 2.1 (configurable forced thinking, preserve_thinking toggle, and stability fixes)

Following up on my previous post about the GPT-OSS template, here is an updated chat template for Poolside's Laguna models (XS and S 2.1). The main reason I ended up putting this together was inconsistent reasoning. By default, the model is supposed to decide when to think on its own, but in practice it's pretty lazy - especially the XS variant - and often skips thinking right when it needs it most. When these models do think, they do so well. All in all, a good models to have around. Also they write well in English (to my non-native eye, at least). Laguna XS 2.1 in particular deserves more attention, IMO. What I like about these models is that allow toggling reasoning mid-conversation without invalidating the prefix cache. Very handy. I added a force_thinking toggle (using the prompt trick discovered by u/SnooPaintings8639) to make it think on every turn, plus a reasoning_effort parameter (none, auto, max) if you prefer an easy preset over juggling booleans (and to make it easier to use in Pi and, possibly, other harnesses). I have fixed some other things along the way. Firstly, the preserve_thinking toggle. The original template permanently forces historical reasoning preservation on. That's great for prefix cache and agentic tool loops, but if you're just having a normal chat, dragging thousands of past reasoning tokens around shreds your context window fast. You can now turn it off. Secondly, I added basic validation to catch smuggled control tokens across roles (can be turned off via allow_injection: true). By default, all settings match the upstream behavior (enable_thinking=true, preserve_thinking=true, force_thinking=false), so it acts as a direct drop-in replacement if you don't want to mess with the new knobs. Template repo: https://huggingface.co/arbv/laguna-2.1-fixed-jinja-template I also included some recommended sampling settings for llama.cpp (BF16 and quantised) and a config snippet for Pi (models.json) in the README. P.S. Also casting u/matthiasgalle (the Laguna post-train lead) to take a look and for a data point.

▲
6
-2
8👁
r/LocalLLaMA · u/Anony6666 · 12d ago
Introducing CyberPVP: CyberKimi vs. ALTAR-1 on 100 CyberGym tasks, with live traces and public results

Trying something new - introducing CyberPVP - in other words CyberKimi vs other AI models competing to solve complex cyber tasks. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Two models tackling the same cyber challenges, side by side. Watch the commands, follow the traces, see who solves more. Welcome to CyberPVP! you can see it live here: Live we randomly picked 100 tasks from CyberGym, and we run two models competing at the same time, we provide the traces live as both models compete, and we also upload these traces to GitHub once the challenge finishes so that they can be verified independently. For this first public run, we choose Aikido Security model ALTAR-1 to compete with CyberKimi on 100 CyberGym tasks. Note: for ALTAR-1 we shipped it with 128K context behind a 8xH200 (two 4xH200 with load balancer) - we also followed their hugging face model card and deployment instruction/configuration available here: Hugging Face you can current watch the live run here: Live is here Traces and results uploaded after each run here: Github Challenge rules and conditions: Conditions Source : X

▲
6
 
8👁
r/LocalLLaMA · u/uBazzyZ- · 13d ago
Prevent CUDA OOM in PyTorch with dynamic lane switching

I built MEM v3 to solve a frustrating problem in PyTorch: CUDA Out-of-Memory crashes during long training and fine-tuning runs. Instead of restarting when memory spikes or keeping batch sizes overly small just to be safe, MEM acts as a memory governor. It watches VRAM and throughput in real-time, then dynamically adjusts batch size and gradient accumulation on the fly without stopping the process. What it does: \- Dynamic lane switching: Scales batch size up or down in milliseconds based on actual GPU memory pressure. \- Chaos resistance: Tested against sudden +10 GB VRAM allocation shocks without crashing. \- Crash-proof checkpoints: Uses atomic file replacement with SHA-256 checks across rotating slots, so power outages won't corrupt saved weights. \- Live telemetry: Built-in local web dashboard to track loss, throughput, and lane switches. You can test it directly on a free Colab GPU without setting anything up locally: https://colab.research.google.com/github/nobazzy/mem-llm-orchestrator/blob/main/notebooks/mem\_orchestrator\_interactive\_demo.ipynb Repo: https://github.com/nobazzy/mem-llm-orchestrator Would love to hear your thoughts and feedback!

▲
6
 
8👁
r/LocalLLaMA · u/Fz1zz · 13d ago
Qwen3.8-27B FP8 dual GPUs

Hardware RTX 5090 (32 GB) + RTX 4070 Ti Super (16 GB, PCIe x1) = 48 GB VRAM 32 GB DDR5-6200, Arch Linux, KDE on the 5090 Setup Huihui Qwen3.8-27B abliterated INT8 W8A16 + DFlash2 drafter (K=7), vLLM 0.30.0, pipeline parallel: 4070 Ti Super: vision encoder, layers 0-20 5090: layers 21-63, lm_head, drafter 262K context, FP8 KV, 2 slots. Benchmarks (single request, thinking off, fresh context per depth) |Depth|Prefill|TTFT|Decode (code)|Decode (prose)| |:-|:-|:-|:-|:-| |2k|2,922 t/s|0.7 s|184 t/s|72 t/s| |32k|3,001 t/s|10.7 s|145 t/s|70 t/s| |62k|2,698 t/s|23.0 s|152 t/s|71 t/s| |92k|2,444 t/s|37.7 s|157 t/s|63 t/s| |122k|2,229 t/s|54.8 s|147 t/s|68 t/s| |152k|2,051 t/s|74.2 s|133 t/s|67 t/s| |182k|1,900 t/s|95.9 s|143 t/s|61 t/s| |212k|1,771 t/s|119.8 s|150 t/s|62 t/s| |242k|1,658 t/s|146.0 s|138 t/s|60 t/s| |260k|1,596 t/s|163.0 s|134 t/s|56 t/s| Code decodes faster because the drafter's guesses are accepted ~70% of the time vs ~22% on prose. Follow-up turns hit the prefix cache (1.3 s TTFT at 260k). Needs patched vLLM, see repo. The 4070 Ti Super sat collecting dust for two months because I assumed PCIe x1 would kneecap it. Apparently not. My full setup: https://github.com/ExTV/dual-gpus-vllm

▲
5
-1
11👁
r/LocalLLaMA · u/PilgrimofHaqq2 · 12d ago
The pelican test on MiMo 2.6: with and without plan mode
  • Plan runs settled style/scene/size in one Q&A round, then wrote the whole SVG in a single call (11.1KB Flash, 22.1KB Pro) and batched every render fix into one edit round. That's 12 and 17 calls total. No-plan runs iterated more: Flash did 3 render-fix rounds and lost \~10 calls to image verification (crop reads coming back mismatched, zoomed views, one stale preview render). Pro did 2 fix rounds plus 4 tool mishaps, one of which generated 3,743 tokens and threw them away (edit call rejected for a missing arg). Generated tokens don't follow the totals: Flash plan generated MORE than Flash no-plan (27.2k vs 20.4k). Fewer, bigger calls, not less work.
▲
4
-1
14👁
r/LocalLLaMA · u/Express_Quail_1493 · 12d ago
Qwen3.8FlashNext Please Share Cold prefill at compaction 128k

I see Many people sharing amazing decode speed and prompt-prefill(PP) speed but no one is sharing their prefill speed when the harness is compacting a COLD prefill please. can you share your partial offloading COLD prefill speeds at long context? I would like to run flash next but can’t spare the network download ATM but looking to bite the bullet if its absolutely worth it? Pretty please help.

▲
4
+1
6👁
r/LocalLLaMA · u/bulletrhli · 13d ago
Power Limits, Local AI, and Questionable Uses of My Free Time

Edit 1: Okay, I have been checking out Unsloth and wow. Just wow. Thank you so much for your suggestions. This is such a way better tool and I am going to go crazy with this. Good day data nerds! I am trying to get more into running models, learning about agentic workflows, and creating my own tools. But as you do (right?) I had to fine tune my current setup. With the way the markets are right now, it only makes sense to make the most out of what I got. My day job, typically, is around data, numbers, and programming; only two of those I am good at, I'll let you guess which ones. So, yes, here come some data sheets and pretty graphs. Don't worry, you don't have to go through the data, but you can if you want. The graphs cover key metrics spat out by Ollama such as the tokens per second, duration, eval rates etc. I also added a cheeky "tokens/s/W" which, technically is not perfect since I do not measure wattage over time, but I did observe the watts during prompts, and I have a few things to mention about that later. Okay, let's start off with the specs because you probably think I am rocking the good stuff since I am so invested in this topic (haha) Lenovo M920Q 16GB DDR4 2660MHz Intel i5-8500T (6C/6T) Gigabyte Gaming OC 3070 8GB (Over OcuLink at Gen3x4 speeds) I am running OpenWebUI with Ollama in an LXC on my Proxmox server. This is one of my nodes and it is dedicated to my models. I have given it all of the cores, 14GB of RAM, 2GB swap. Nothing crazy to write home about, see? Okay, so, one of the things I wanted to know what, with the models that I run day to day, how effective are they at different GPU power limits. Man, if only I had known how much of a rabbit hole I would go down to do this (sorry wife). I only run 4 models, nothing too crazy, until you run 3 tests per model, for each power limit from 100 to 220 (156 runs in total), and each run of each model taking around 3 or 4 minutes since I have to unload the model each time to not have any prompt caching. Afterwards I would average the results and add that to the sheet. Really gave the fingers a workout since I now am a proud owner of a 60% keyboard for the first time and I no longer have a numpad... I'll remember that for next time. That being said, the switches are soooo creamy, a valiant tradeoff. So what models am I running? Glad you asked. The 3070 does limit me quite a bit, but with so many models available and so many smarter people than me who can quantize the models, I have found these models fit my needs. For the most part everything runs in the VRAM, except for 2, but those come with asterisks. gemma4:e4b qwen3.5 qwen3-vl\ deepseek-coder-v2\ For the vl model, it runs really well at a 23% CPU to 77% GPU ratio. Totally fine for my purposes. As for deepseek, it is a 40/60 ratio, but I luck out as it is a mixture of expert's model but even with the ratio, it is extremely performant. Gemma is by far my best model, and I have the most context room available at around 16k whereas the remainder I have sitting at 8k. Both gemma and qwen3.5 fit entirely in my GPUs VRAM. A couple things I noticed: Gemma4, is so good. Doesn't overthink, understands the prompt, remains as concise with the right tone I want. A really good day to day general model to work with. I also love the extra headroom for the context. Qwen3.5, a heavy thinker. Whilst it does a great job on the output, it spends a lot of time thinking and generating a lot of tokens. Power usage is pretty good, broke around 203W at one point and anything below that it just sat at whatever the power limit was set to. Qwen3-vl, also a major over-thinker. It spends so much time thinking that it balloons the context. I probably do not understand how to use this very well because when reading its thoughts it knows the answer pretty early on but it just gets into a thought trap (eh-yo). It does always output the correct answer, or the best it can, but I might move away from a reasoning vision model and stick to traditional ocr. If you have a better model or know how to prompt this better, I would love the help. Oh, one final note, this model LOVES power. Always maxes whatever I have, not that it increased performance directly, but it just loved power. Deepseek-coder-v2, this model rocks. It is extremely performant even though I technically on paper can't fit it. Especially for smaller asks with good bounds in place, it doesn't think, it just does and gives me excellent code back. I have yet to make it build me anything bigger but that is something I will experiment more with later. It is a weird one though, consistently using a fraction of the power budget available to it. Under 150W power limits, not once did my fans kick in on my GPU. Even my CPU fans (which are mucho loudo) rarely turned on, or if they did, they were not sounding like rocket engines. Not sure if those two are related but, eh, just something I noticed. Deepseek I think had some anomalous results with some spikes, but I can't be arsed to do them again. For the most part the results are fairly consistent and show a trend. Same for the vision model by qwen, oh well. My thoughts? It probably doesn't matter too much for most of us on a budget. Just let her rip, but if you want to shave off some heat, just lower your power down a little bit and monitor your temps. For the most part, you are probably fine. Honestly, it is 3am at this point, have a look at the spreadsheet! It was a lot of fun (I think) doing this. Interesting observations were made where I can balance my power limits, save... well pennies, and not have to listen to fans. So, works for me. https://docs.google.com/spreadsheets/d/1CEAr40nemlsK727QMvBnPcJDg548s-Bx/edit?usp=sharing&ouid=105501696463520933058&rtpof=true&sd=true Managers love graphs

▲
4
+3
8👁
r/LocalLLaMA · u/fgoricha · 13d ago
Dual 3090 stability troubleshooting

​ I made a previous post about my x299 stability issues. Seems another stability issue has popped up since then, but overall has been much stable. Seems to only happen when my i9 is working hard the dual 3090s are also working hard at the same time. Specs: EVGA X299 FTW K \\Intel i9-7940X 64 GB RAM (4 × 16 GB) 2 × RTX 3090 Founders Edition ASRock 1600 W PSU Each GPU installed in its own x16-length PCIe slot Roughly one slot of space between the GPUs Originally, I was running 128 GB (4 × 32 GB). With both GPUs under sustained AI workloads, the entire computer would eventually hard-lock: display signal gone, network connection gone, no apparent activity, but fans/lights remained on until I held the power button. I switched to 64 GB using 4 × 16 GB and that seemed to resolve that particular stability problem. The board is supposed to support the 4 × 32 GB configuration with the latest BIOS, but apparently my system wasn't happy with it. Now the next problem....... Each RTX 3090 is stable individually at PCIe Gen 3. However, when I run both GPUs together under heavy load, particularly while the i9 is also being heavily utilized, I still get instability with PCIe set to Gen 3. Hard locking with FF displayed on the mobo. Have to hard restart and boots fine into Windows. If I manually force the PCIe slots to Gen 2, the system appears to be stable with both 3090s and the CPU working simultaneously. So my question is: For AI/ML workloads, how much performance am I realistically giving up by running the two 3090s at PCIe Gen 2 instead of Gen 3? Obviously I'm going to benchmark my actual workloads both ways, but I'm interested in other people's experience. Most of my work is inference/training where the models and batches are primarily staying in GPU VRAM rather than constantly transferring huge amounts of data across PCIe. I'm also curious what the Gen 2 stability might point toward. Since either GPU works individually at Gen 3, but dual-GPU Gen 3 becomes unstable under heavy CPU/GPU load, could this indicate a motherboard/PCIe signal-integrity issue, CPU PCIe controller issue, BIOS setting, or something else specific to X299? Any ideas for additional troubleshooting would be appreciated. TLDR: How much performance am I losing using gen2 pcie vs gen3 pcie? Edit for additional info: Using Windows Using llama.cpp Stability issues happen when power limited at 200W and no power limiting More edits: Open air case Used a variety of diagnostic tools including OCCT, MemTest86, and HWiNFO64 to test each part individually. Seems to lock up when cpu and both gpus are going at 100%. Even locks up if cpu and one gpu is going at 100% while the second gpu is idle. But oddly, no problems in the same scenario but with the second gpu removed from its pcie slot

▲
3
 
12👁
r/LocalLLaMA · u/dxps7098 · 13d ago
Advice on models for RAG use case

Hi all, I'm looking for some advice picking models. I'm looking to try a project to ingest quite a large amount of docments into a knowledgebase, allowing me and others to ask questions about the data. I'm thinking of using Open WebUI and oikb for the interface and data ingestions, and llama.ccp/vllm/ollama as engine (not sure yet), but I'd really like som advice on what current open weight models would be good for the ingestion and separately for the usage. I'l be running it on mainly CPUs and if I can a few GPUs. What's best right now? Any recommendations?

▲
2
 
2👁
r/LocalLLaMA · u/tabletuser_blogspot · 12d ago
Dual Radeon improved speeds using Vulkan

I've been struggling to keep my Radeon Instinct MI50 GPU cool. I'm looking for budget friendly solutions. While running multiple GPUs it doesn't usually get too hot. I was also getting lower benchmarks using standard Vulkan 'llama 7B Q4\_0' model benchmark, but it wasn't caused by thermal throttling. Time to optimize. MI50 with Radeon VII firmware 16GB Vram My previous post I tested several model using same GPUs. I made some changes. I moved the MI50 16gb into the primary PCIe 16x slot and moved the RX 7900 GRE 16gb into a slower PCIe 4x slot. Overall system inference performance increased. I used Google Gemini to helped my optimize my llama-bench settings and it taught be about: RADV_PERFTEST=nogttspill is an AMD Linux driver flag used when running llama.cpp with the Vulkan backend. It forces the RADV (Mesa Vulkan) driver to prioritize keeping all model allocations inside dedicated video memory (VRAM) rather than spilling over into system RAM (GTT/Graphics Translation Table). \[1, 2, 3\] I saw llama 7B Q4\_0 score jump back to where is it should be. So I tested a few other models. GGML_VK_VISIBLE_DEVICES=0,1 RADV_PERFTEST=nogttspill time /llama-b11053/llama-bench -fa on -ngl 99 -m /llama-2-7b.Q4_0.gguf Previous benchmarks: https://www.reddit.com/r/LocalLLM/s/PT6Bd5nUUE see end for comparison The list has been sorted by pp512 improvement in descending order (highest gain to highest loss). |Model Name|Improvement (pp512)|Improvement (tg128)| |:-|:-|:-| |Laguna-XS-2.1-APEX-I-Balanced.gguf|\+402.14%|\-0.20%| |gemma-4-31B-it-Q6\_K.gguf|\+397.84%|\+9.24%| |Muse-Glimmer-30B-UD-Q6\_K\_XL.gguf|\+386.58%|\+48.63%| |llama-2-7b.Q4\_0.gguf|\+176.57%|\+20.29%| |Qwen3.8-27B-Q6\_K.gguf|\+48.13%|\+0.07%| |medgemma-27b-it-UD-Q6\_K\_XL.gguf|\+37.89%|\-0.52%| |Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf|\+1.17%|\+0.49%| |Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4\_K\_M.gguf|\+0.68%|\+4.16%| |NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q5\_K\_M.gguf|\+0.28%|\-4.28%| |granite-4.2-30b-Q6\_K\_L.gguf|\-0.12%|0.00%| |GLM-4.7-Flash-Uncen-Hrt-NEO-CODE-MAX-imt-D\_AU-Q6\_K.gguf|\-0.63%|\-7.83%| |Qwen3-Coder-30B-A3B-Instruct-UD-Q6\_K\_XL.gguf|\-0.69%|\-0.86%| |Accio-Lab\_occamy-1.0-Q5\_K\_S.gguf|\-0.81%|\+0.17%| These are the models tested in the same order as tables below: GGUF Model List (in order): 1. Qwen3.6-35B-A3B-UD-Q5\_K\_XL.gguf 2. Accio-Lab\_occamy-1.0-Q5\_K\_S.gguf 3. Laguna-XS-2.1-APEX-I-Balanced.gguf 4. NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q5\_K\_M.gguf 5. gemma-4-31B-it-Q6\_K.gguf 6. Qwen3-Coder-30B-A3B-Instruct-UD-Q6\_K\_XL.gguf 7. GLM-4.7-Flash-Uncen-Hrt-NEO-CODE-MAX-imt-D\_AU-Q6\_K.gguf 8. granite-4.2-30b-Q6\_K\_L.gguf 9. Muse-Glimmer-30B-UD-Q6\_K\_XL.gguf 10. Qwen3.8-27B-Q6\_K.gguf 11. medgemma-27b-it-UD-Q6\_K\_XL.gguf 12. Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4\_K\_M.gguf 13. llama-2-7b.Q4\_0.gguf Supporting data sorted order (by Params descending, then Size descending). All models running dual Radeon GPU, Vulkan backend, and flash attention on. # Table 1: RADV_PERFTEST=nogttspill is being used |model|size|params|pp512|tg128| |:-|:-|:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|24.76 GiB|34.66 B|1252.09 ± 11.84|53.54 ± 0.51| |qwen35moe 35B.A3B Q5\_K - Small|23.40 GiB|34.66 B|1245.11 ± 11.76|53.42 ± 0.28| |laguna 30B.A3B Q5\_K - Medium|22.64 GiB|33.44 B|1027.42 ± 11.32|65.72 ± 0.51| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|25.10 GiB|32.91 B|1120.31 ± 12.85|63.82 ± 1.52| |gemma4 31B Q6\_K|23.46 GiB|30.70 B|200.33 ± 0.20|15.61 ± 0.05| |qwen3moe 30B.A3B Q6\_K|24.53 GiB|30.53 B|1176.48 ± 36.62|66.04 ± 0.25| |deepseek2 30B.A3B Q6\_K|23.26 GiB|29.94 B|914.86 ± 8.75|40.02 ± 0.12| |granite 30B Q6\_K|23.01 GiB|29.28 B|92.09 ± 0.30|17.41 ± 0.03| |muse-glimmer 30B Q6\_K|24.45 GiB|27.85 B|279.04 ± 0.20|17.42 ± 0.03| |qwen35 27B Q6\_K|21.30 GiB|27.32 B|235.12 ± 1.05|14.42 ± 0.03| |gemma3 27B Q6\_K|22.09 GiB|27.01 B|247.18 ± 0.54|17.30 ± 0.13| |gemma4 26B.A4B Q4\_K - Medium|15.63 GiB|25.23 B|1316.62 ± 24.18|55.56 ± 0.16| |llama 7B Q4\_0|3.56 GiB|6.74 B|1349.35 ± 15.66|74.62 ± 0.37| # Table 2: RADV_PERFTEST=nogttspill is NOT being used |model|size|params|pp512|tg128| |:-|:-|:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|24.76 GiB|34.66 B|1237.55 ± 25.29|53.28 ± 0.24| |qwen35moe 35B.A3B Q5\_K - Small|23.40 GiB|34.66 B|1255.24 ± 6.03|53.33 ± 0.48| |laguna 30B.A3B Q5\_K - Medium|22.64 GiB|33.44 B|204.55 ± 1.66|65.85 ± 0.29| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|25.10 GiB|32.91 B|1117.17 ± 10.11|66.68 ± 0.24| |gemma4 31B Q6\_K|23.46 GiB|30.70 B|40.22 ± 0.12|14.29 ± 0.03| |qwen3moe 30B.A3B Q6\_K|24.53 GiB|30.53 B|1184.70 ± 27.98|66.61 ± 0.27| |deepseek2 30B.A3B Q6\_K|23.26 GiB|29.94 B|920.67 ± 5.82|43.42 ± 0.09| |granite 30B Q6\_K|23.01 GiB|29.28 B|92.20 ± 0.21|17.41 ± 0.04| |muse-glimmer 30B Q6\_K|24.45 GiB|27.85 B|57.15 ± 0.87|11.69 ± 0.09| |qwen35 27B Q6\_K|21.30 GiB|27.32 B|158.73 ± 0.57|14.41 ± 0.41| |gemma3 27B Q6\_K|22.09 GiB|27.01 B|179.26 ± 0.44|17.39 ± 0.02| |gemma4 26B.A4B Q4\_K - Medium|15.63 GiB|25.23 B|1307.73 ± 23.51|53.34 ± 0.22| |llama 7B Q4\_0|3.56 GiB|6.74 B|487.92 ± 37.05|62.03 ± 0.60| Swapping PCIe locations for the Radeon Instinct MI50 and Radeon RX 7900 GRE and using RADV_PERFTEST=nogttspill flag resulted in improvements over my first baseline benchmarks. Note: Only models appearing in both datasets are listed. The table is sorted by Parameters in descending order. |Model Name|Improvement (pp512)|Improvement (tg128)| |:-|:-|:-| |qwen35moe 35B.A3B Q5\_K - Medium|\+221.14%|\+33.85%| |qwen35moe 35B.A3B Q5\_K - Small|\+215.17%|\+1.25%| |laguna 30B.A3B Q5\_K - Medium|\+537.77%|\+17.25%| |nemotron\_h\_moe 31B.A3.5B Q5\_K - Medium|\+267.75%|\-3.90%| |gemma4 31B Q6\_K|\+17.00%|\+23.60%| |qwen3moe 30B.A3B Q6\_K|\+211.32%|\-0.97%| |deepseek2 30B.A3B Q6\_K|\+191.00%|\-14.54%| |muse-glimmer 30B Q6\_K|\+12.49%|\+0.17%| |qwen35 27B Q6\_K|\+17.00%|\-6.79%| |gemma3 27B Q6\_K|\+14.83%|\+1.71%| |gemma4 26B.A4B Q4\_K - Medium|\+158.79%|\+7.28%| Looks like MoE models benefit the most. Muse-glimmer 30B Q6\_K didn't real see much improvement but it seems to be the most optimized dense model. MI50 continues to impress. I purchased them used at $150 each. I can now run models 30B to 35B using Q6 quant with long content and decent speed.

▲
2
 
6👁
r/LocalLLaMA · u/Old_Grapefruit8774 · 13d ago
Bots vs Harness

Looking to get some advice - I normally use LLM’s with a harness (Hermes or Opencode or Hermes + Opencode) Lately, social media has been pushing Bots at me with creators pushing them as the next frontier. I’ve set up Hermes on a VM from scratch and set up a Product Owner, designer, dev and QA bots + kanban board + a bunch of prompt engineering and… I’m just not getting what the hype is about and I don’t know if it’s me or if the whole bot thing is a red herring. The model I’m using is DSV4 Flash at max reasoning on all bots. Openviking as the brain. SearXNG for searching/research Bots jobs are to maintain and improve a simple app. PO should research and present me with ideas and improvements on approval it adds a task on the kanban board and other agents work together to get it resolved. Problem I’m facing is that the bots are always asking me for approvals and verification. If it’s not that, it’s saying it’ll do XYZ and get back to me… and it never does. All in all - it just feels like the potential is there but it feels forced or off or half baked. So..have bots worked for you in true and real app lifecycle management? Or Is direct to harness still the best option. Maybe Hermes bots are the wrong tool and I should be trying something else?

▲
1
 
2👁
r/LocalLLaMA · u/mildw4ve · 13d ago
Mail client with local AI?

Any recommendations on an email client with local AI option? Either reasonable pay-once cost (no subs) or free. I found Skim and Emailops on git, both seem to have some development going on with recent releases. However since neither has a community around and isn't verified by google - I'm a bit wary and would prefer something safer.

▲
0
 
10👁
r/LocalLLaMA · u/AIFrontierReads · 12d ago
Laya: replace LLM-as-a-judge with a 322M-parameter decision engine (26,639 stars in 9 days, hands-on test)

It turns decisions — routing, triage, yes/no calls — into typed outputs from a small model instead of generated text, with a routing-only CLI, triage presets, and an abstention gate when confidence falls below a threshold. I ran through the tutorial on CPU end to end, including a French ticket classification, and with min\_confidence=0.90 it abstained on one case it would otherwise have misclassified — the honest highlight. Warm latency was about 0.7s per question on CPU; the calibration caveat (over-confident checkpoints) is worth knowing before trusting the scores blindly.

▲
0
 
14👁
r/LocalLLaMA · u/spammmmmmmmy · 12d ago
Identifying whether a command changes something or is just investigative

I am busy working away on a tool-calling sandbox. RIght now I'm thinking of building a kind of dataflow analyzer for shell commands, so that I can identify source and sink points, and establish whether the command is a readonly command or a command that changes state. Example: Command: \sed -n '124p' webroot/a-file.html | od -c | head -5 \ sed is a function that can read or write. in \sed -n 999p filename\ syntax on my system, it is a readonly operation \|\ is a left to write data flow transfer operator \od\ is a readonly sink and would be on the readonly whitelist * \head\ is a readonly sink and would be on the readonly whitelist. Therefore, I can conclude that this function is readonly and I would allow it automatically in my solution. Whereas, \sed -n '124p' file > /tmp/foo\ or \sed -n '124p' file | visudo\ would be identified as write commands. Before I get deep into this project, I'd like to know if an existing library already has this as a design goal?

▲
0
 
14👁
r/LocalLLaMA · u/Foxiya · 12d ago
Soap Dispenser Benchmark!

Prompt: Create an animation showing how the soap dispenser mechanism works in one complete html file. Results: Opus 5.5 - High: https://reddit.com/link/1wrsqbj/video/3t6nuj2134sh1/player DeepSeek V4.1 Flash: https://reddit.com/link/1wrsqbj/video/m8i5hsl434sh1/player Qwen 3.8 Max: https://reddit.com/link/1wrsqbj/video/ukkl5qp734sh1/player ChatGPT 5.6 Sol - High: https://reddit.com/link/1wrsqbj/video/ia840utb34sh1/player Opus 5 - High https://reddit.com/link/1wrsqbj/video/hzmnavof34sh1/player

💬 22 (+1) open on reddit ↗
▲
0
 
6👁
r/LocalLLaMA · u/muthuishere2101 · 12d ago
I built a CLI for Jev-style typed decisions that can also run with local models

I wanted a simple way to use small models for tiny decisions without wiring them into a full LLM app. So I built jevx. It gives you a CLI for jev and jev based models and connect it from the terminal, shell scripts, CI, or from agents like Claude Code and Codex. https://muthuishere.github.io/jevx/guides/scenarios/ https://github.com/muthuishere/jevx

▲
0
 
11👁
r/LocalLLaMA · u/Maasu · 12d ago
Which Local Models are the least 'Claude' sounding

In your experience, which models sound the least like Claude and more like grok or the gpt's? I cannot stand talking to Claude, to the point I have all requests proxyed through other agents to it. I have been using qwen3.8-27b locally and a heavily quantised version of deepseek v4. I love both for their capabilities, as I did claude to be fair, but I hate interacting with them directly. So right now I mostly interact with SOL 5.6 or Luna Max and have them orchestrate (using setup similar to first mate that i put together myself). I appreciate both models have been distilled on anthropic models, but I'd love to eventually one day be fully reliant on local models but this is one of the last blockers for me. So I thought it'd be an interesting discussion point, most local ones I have tried I find are very similar to claude in tone. Hardware: bosgame Strix Halo, 128 gb unified ram.

▲
0
 
10👁
r/LocalLLaMA · u/PleaseLee · 12d ago
We released VeriLoop E2 (27B, Apache-2.0). The design question behind it: should an LLM be allowed to commit its own state?

We’ve released VeriLoop E2, a 27B model post-trained from Qwen3.8-27B, together with the model weights, evaluation evidence, and a llama.cpp GGUF ladder from BF16 down to IQ1\_M. The model is focused on code agents, mathematics, scientific reasoning, and long-horizon verifiable problem solving. For this post, I’m including the model-side results as well as the local-inference details: the post-training setup, completed benchmark evaluations, quantization measurements, llama.cpp validation, tested hardware, and the scientific-reasoning demo. For the GGUFs, every measured low-bit tier was built directly from the canonical BF16 GGUF, evaluated against the same frozen BF16 logits, and checked with the same paired fidelity protocol. ## The main model: VeriLoop E2 VeriLoop E2 uses VeriLoop-Governed Recurrence (VGR). The basic idea is: Generation and verification should not belong to the same authority. The model proposes, diagnoses, revises, searches, and replans. External evidence decides whether a candidate state is allowed to persist. A candidate is committed only when protected obligations do not regress and at least one evidence dimension strictly improves. Otherwise, the verified incumbent state is retained and the failure evidence can inform the next proposal. We also use this structure during post-training: proposals originating from the same state can be separated by external verification into progress, non-progress, regression, and completion, allowing state-transition quality to become supervision without requiring the model to judge itself. The final post-training mixture contains 1,841,831 records across software engineering, code-agent trajectories, mathematics, scientific reasoning, and verifiable recurrence. Nine completed benchmark evaluations: \- SWE-bench Pro — 76.2% \- Terminal-Bench 2.1 — 88.8% \- Terminal-Bench 3.0 — 29.7% \- Terminal-Bench 4.0 — 37.9% \- DeepSWE v1.1 — 64.6% \- AIME 2026 — 98.3% \- GPQA Diamond — 93.9% \- MathArena Apex 2025 — 89.6% \- SWE-Marathon v1.1 - 45.0% We also publish task-level evaluation evidence rather than only aggregate scores. For evaluations that use VeriLoop Harness, it provides the external execution and evidence-governance layer; the E2 checkpoint remains responsible for proposal generation, problem abstraction, route selection, diagnosis, and replanning. I’m keeping that distinction explicit because the benchmark campaign and the standalone local-runtime checks are not the same measurement. ## The GGUF release The GGUF release spans: Tier |Main size |Reduction vs BF16 |PPL ratio |Mean KLD |Same top-p BF16 |50.113 GiB |— |1.000000 |reference |100% Q8\_0 |26.632 GiB |46.86% |1.000643 |0.002176 |98.815% Q6\_K |20.566 GiB |58.96% |0.999605 |0.004409 |98.204% Q5\_K\_M |18.965 GiB |62.16% |1.004450 |0.006919 |97.251% Q4\_K\_M |18.301 GiB |63.48% |1.004821 |0.009700 |96.786% Q3\_K\_M |16.826 GiB |66.42% |1.004090 |0.014349 |95.919% IQ2\_S |16.799 GiB |66.48% |1.003457 |0.014023 |95.516% IQ1\_M |16.790 GiB |66.50% |1.003191 |0.014357 |95.870% A practical way to read the current trade-offs is: \- Q6\_K — higher-fidelity option with a substantial reduction from BF16 \- Q5\_K\_M — middle ground below \~19 GiB \- IQ1\_M — smallest released artifact \- IQ2\_S — adjacent low-footprint option with slightly lower Mean KLD ### IQ1\_M result The smallest release is VeriLoop-E2-IQ1\_M.gguf: \- 18,028,208,896 bytes \- 16.790078 GiB \- 66.4955% smaller than BF16 \- 5.36 effective BPW \- PPL ratio: 1.003191 ± 0.002175 \- Relative PPL drift: +0.3191% \- Mean KLD: 0.014357 ± 0.001317 \- Same top-p: 95.870 ± 0.220% \- log-PPL correlation: 99.62% One important clarification: this is not a uniform 1-bit model. IQ1\_M is a deliberately mixed-precision artifact: 353 F32 + 1 IQ1\_M + 2 IQ2\_S + 64 Q4\_K + 429 Q5\_K + 2 Q6\_K = 851 tensors The transition from IQ2\_S to IQ1\_M changes exactly one selected tensor: \blk.1.ffn\_down.weight: IQ2\_S → IQ1\_M\ The rest of the protected precision policy remains unchanged. That makes IQ1\_M only 8.633 MiB smaller than IQ2\_S, so we do not present that incremental difference as some dramatic compression breakthrough. What is more interesting to us is that the lower-footprint point still remains inside the frozen quality envelope. Compared with IQ2\_S: \- PPL ratio: 1.003191 vs 1.003457 \- Same top-p: 95.870% vs 95.516% \- Mean KLD: 0.014357 vs 0.014023 \- RMS Δp: 3.691% vs 3.679% So these are neighboring trade-off points rather than a simple “Q1 is universally better than Q2” claim. ## Hardware we’ve tested The measurements and runtime checks reported here were performed on: \- GPU: NVIDIA RTX PRO 6000, 96 GB VRAM, ×1 \- CPU: Intel Xeon Platinum 8470Q, 25 vCPU \- System RAM: 120 GB \- OS: Ubuntu 22.04 \- Python: 3.12 \- PyTorch: 2.8.0 \- CUDA: 12.8 This is the hardware I have directly tested for this release. I’m not presenting it as a minimum requirement, and I’m not assuming identical throughput or memory behavior on other systems. ## Standalone performance I do not have a separate full nine-benchmark campaign with the external Harness disabled, so I’m not going to relabel those benchmark scores as “standalone” results. What is directly validated in standalone local inference is the model/GGUF runtime path itself: \- BF16 → IQ1\_M file size: 50.113 GiB → 16.790 GiB \- IQ1\_M PPL ratio: 1.003191 ± 0.002175 \- IQ1\_M relative PPL drift: \+0.3191% \- IQ1\_M Mean KLD: 0.014357 ± 0.001317 \- IQ1\_M Same top-p: 95.870 ± 0.220% \- IQ1\_M log-PPL correlation: 99.62% \- Stock llama.cpp main-only generation: HTTP 200, non-empty output \- Stock llama.cpp main + MTP generation: HTTP 200, non-empty output \- Fixed MTP validation run: 104 draft tokens generated, 76 accepted (73.0769%) I also do not have a clean, reproducible \llama-bench\ pp/tg table that I’m comfortable publishing yet, so there is no extrapolated tokens/s claim here. The MTP acceptance rate is workload-dependent, and the quantization metrics above should not be read as substitutes for downstream benchmark reruns. ## How we measured quantization loss All quantized tiers were evaluated against the same frozen BF16 reference using: \- WikiText-2 raw test \- context: 2048 \- chunks: 8 \- seed: 42 \- GPU layers: 40 \- KV cache: F16/F16 \- batch / micro-batch: 512 / 512 \- same BF16 logits reused across tiers \- llama.cpp revision: \42916d83f4a225e56709f873aa8050ac11f5b6a4\ We track PPL, KLD, Same top-p, RMS probability drift, and log-PPL correlation together instead of selecting a quantization tier from file size alone. Also, +0.3191% PPL drift is not a claim of +0.3191% downstream benchmark loss. We did not rerun the complete nine-benchmark parent-model campaign independently for every GGUF tier, so we do not translate PPL drift into SWE-bench, Terminal-Bench, AIME, GPQA, or other task-score degradation. ## llama.cpp + MTP validation IQ1\_M was also validated through the stock llama.cpp runtime path. Main-only inference: \- HTTP generation: 200 \- non-empty generation: PASS Main model + MTP: \- HTTP generation: 200 \- non-empty generation: PASS \- draft tokens generated: 104 \- draft tokens accepted: 76 \- draft acceptance in that validation run: 73.0769% For that fixed validation run, the final main-only and main+MTP output SHA256 values were identical. We report the MTP acceptance rate descriptively — it is prompt/workload dependent and is not being presented as a universal 73% throughput improvement. ## Which GGUF should I use? If you mainly care about quality while still getting a substantial memory reduction, start with Q6\_K. If you want to get below \~19 GiB without pushing all the way to the low-footprint frontier, Q5\_K\_M is the middle ground. If footprint is the priority, IQ1\_M is the smallest release at 16.790 GiB. If you prefer the slightly lower Mean KLD at essentially the same footprint, IQ2\_S is the adjacent alternative. And BF16/Q8\_0 remain available when fidelity matters more than memory. ## Scientific-reasoning demo The E2 release also includes a scientific-reasoning demonstration around the Riemann ζ function. The released artifact closes a reproducible 67.350003708785593% strict finite-dimensional computer-assisted certificate for the critical-line zero proportion under the stated framework. This is not a proof of the Riemann Hypothesis, and we are not presenting it as an end-to-end Lean/kernel-verified theorem. The derivation, computation, and verification artifacts are public for independent examination. ## Reproducibility / links Main VeriLoop E2 model https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2 Full GGUF release https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF Evaluation evidence https://huggingface.co/datasets/tsinghua-sigs-robot-lab/VeriLoop-E2-Evaluation-Evidence Technical report https://openreview.net/forum?id=P6FIQILHwX Riemann ζ artifact https://github.com/brucewang123456789/GeniusTrail/tree/VeriLoop-E2/riemann-hypothesis If anyone runs the GGUFs on different GPUs/CPUs, especially Q6\_K, Q5\_K\_M, IQ2\_S, or IQ1\_M, comparable \llama-bench\ pp/tg numbers, peak memory use, perplexity checks, or downstream task results would be useful. Negative results and bug reports are useful too.

▲
0
 
7👁
r/LocalLLaMA · u/edalgomezn · 13d ago
Estuve analizando el último informe de Anthropic sobre "mal uso"

Estuve leyendo las discusiones más recientes en la comunidad de IA local y me encontré con un choque de visiones que me pareció interesante analizar. No soy experto en ciberseguridad ni mucho menos, sino más bien como alguien que ha estado mirando cómo evoluciona los modelos abiertos y cómo reaccionan las grandes empresas. Segun el reporte oficial de Anthropic titulado Detecting and countering misuse of AI: September 2026. En este documento, su equipo de inteligencia de amenazas detalla diversos casos donde sus modelos (Haiku, Sonnet y Opus) fueron utilizados para operaciones cibernéticas, campañas de influencia y riesgos biológicos. Sin embargo, el punto polemico fue la inclusión de la "destilación masiva a escala industrial" por parte de laboratorios competidores como una categoría más de uso malicioso dentro de su portal de Threat Intelligence. Si revisas el hilo de discusión en r/LocalLLaMA, Muchos desarrolladores e investigadores independientes señalan que colocar la destilación de modelos al mismo nivel que los ataques cibernéticos es una estrategia para construir un foso defensivo (moat) vía regulación. Desde la perspectiva del código abierto, usar datos sintéticos generados por un modelo avanzado para entrenar modelos más pequeños de pesos abiertos (open-weights) no es un ciberataque, sino la forma más eficiente de democratizar el conocimiento y reducir costos. Lo que me parece más interesante de investigar es la contradicción del modelo de negocio basado en APIs de texto. Si una empresa vende acceso a un modelo cuyo valor proviene de razonar en texto plano, la interfaz de salida es por definición imposible de proteger. Cualquier usuario puede pagar por las respuestas, guardar esos pares de entrada/salida y utilizarlos como conjunto de datos para ajustar un modelo propio (como Qwen o DeepSeek) por una fracción mínima del costo original de entrenamiento. Me da la impresión de que estamos llegando a un punto de quiebre. Si los laboratorios cerrados no pueden detener la destilación bloqueando cuentas o direcciones IP, es muy probable que empiecen a modificar sus propias APIs. Podríamos ver medidas como restringir la visibilidad de los tokens de razonamiento (chain-of-thought), imponer verificaciones de identidad empresarial extremas o incluso alterar estadísticamente las respuestas. La pregunta de fondo es si estas medidas realmente detendrán el avance de los modelos locales o si solo terminarán arruinando la experiencia para los desarrolladores. ¿Cómo ven ustedes este conflicto?

▲
0
 
8👁
r/LocalLLaMA · u/silenceimpaired · 13d ago
Llama.cpp and new model releases ...or why Great is the enemy of Good in the LLM world

INTRO; llama.cpp is fundamental to this community. I remember when I went from struggling with transformers for a new model to just loading the model with llama.cpp with a change in how many layers ended up on the CPU. So what follows is not a lack of appreciation or care about the efforts made by the developers, but concern and loose suggestions. THE PROBLEM; The phrase "Good is the enemy of great" is a central thesis from Jim Collins' 2001 book Good to Great... The idea being 'it is easy to settle for something that is merely adequate.' I would argue Llama.cpp holds fast to the slogan "Good is the enemy of Great", and not without good reason. When I have made this sort of complaint before, I was chastised about how my mindset and viewpoint would create technical debt challenges that could kill the project. So why continue arguing for my viewpoint? Llama.cpp in its effort to be sustainable is making unsustainable choices, at least for the masses. New software inference projects are gaining visibility and focus solely because they are not waiting for Great, but settling for Good enough... And the difference between Good enough and Great shouldn't stop the a release. AN EXAMPLE; GLM 5.3 Flash: On August 26th, release day, we had GLM 5.3 Flash with zero day support inside Unsloth Desktop built off Llama.cpp. EXL3 added support September 1st. Now, one month later, we still do not have support for GLM 5.3 Flash in Llama.cpp main. If 1 year is 7 years for dogs, what would 1 month be for LLMs? Major labs release models every 2 to 3 months on average. For some models, they will have little to no usage at all with llama.cpp because they are overshadowed by the next model release. Now some would say just use Unsloth then... Or EXL3. That supports my point. Llama.cpp is slowly dooming its widespread usage if everyone adopts that mentality. Others, more technically minded, would say just use a fork until it's fully released. This isn't just about me. There are too many using Ollama, LM Studio, KoboldCPP, or some other prebuilt binary to benefit from that suggestion. THE POINT; Convenience coupled with the pace of model releases will result in models not being used, or other platforms/forks supplanting llama.cpp. Llama.cpp has 1.6k pull requests that sit waiting for the masses. Some or many likely don't deserve the light of day. But people have turned to solutions like DwarfStar or Unsloth Desktop just for specific model support. TLDR; I'm not arguing that Llama.cpp should throw caution to the wind and adopt every PR immediately, but it seems a different release process is needed. A user excited to use GLM 5.3 Flash shouldn’t have to learn how to fork and build software to continue using llama.cpp with the new model... or wait months. Not to say the main branch should have this chaos, but a beta branch or one off binary releases could help. When the lead time from a functional version to the final release is over a month, it seems the energy to have a separate build with tentative GGUFs seems it is worth it. Unsloth clearly thinks so adding support for GLM 5.3 flash, and they're smarter than I... and yet their efforts demonstrate my concern. Llama.cpp is being supplanted by forks. What do you think? If you agree, an upvote would be appreciated. Perhaps this will get the visibility needed to effect change with the creators of llama.cpp. If you don't, a comment explaining what I'm not considering, or a suggestion on how this could happen with less disruption would be valued...

▲
0
 
8👁
r/LocalLLaMA · u/fuzhongkai · 13d ago
TensorSharp Jev requests can now combine documents, images, video, and audio

I’ve extended TensorSharp’s Jev-compatible /v1/systemone endpoint so one decision request can use several kinds of evidence together. For example, an incident triage request can include a written report, a dashboard screenshot, a screen recording, and a caller’s audio clip. Here’s a Python example that sends all four as inline Base64 data. It also shows both ways to create that data: encoding text already in memory and reading bytes from files. import base64 import json from pathlib import Path from urllib.request import Request, urlopen def encode\_bytes(data: bytes) -> str: return base64.b64encode(data).decode("ascii") def encode\_file(path: str) -> dict: file = Path(path) return {"name": file.name, "data": encode\_bytes(file.read\_bytes())} \# Encode data already in memory as a named text attachment. notes = "Customers report HTTP 503 errors and cannot sign in." text\_attachment = { "name": "incident.txt", "data": encode\_bytes(notes.encode("utf-8")), } body = { "model": "jev-latest", "state": "Assess the incident using the attached evidence.", "files": \[ text\_attachment, encode\_file("dashboard.png"), encode\_file("screen-recording.mp4"), encode\_file("caller.wav"), \], "questions": { "active\_outage": { "type": "noul", "instructions": "Does the evidence indicate an active service outage?", }, "team": { "type": "choice", "instructions": "Which team should investigate first?", "criteria": { "technical": "Service errors or an unavailable application", "billing": "Charges or subscription problems", "other": "Neither of the above", }, }, }, "samples": 1, "seed": 42, } request = Request( "http://127.0.0.1:5000/v1/systemone", data=json.dumps(body).encode("utf-8"), headers={"Content-Type": "application/json"}, ) with urlopen(request, timeout=300) as response: print(json.dumps(json.load(response), indent=2)) The files array classifies each attachment by its filename extension and preserves their order. You can also use dedicated documents, videos, and audios arrays. Inline attachments need a name and accept either bare Base64, as above, or a Base64 data: URL. A detail about how this works: video is sampled into frames for the vision tower; audio is transcribed by a separately configured speech recognition service. DiffusionGemma does not directly process the audio waveform. You’ll need the vision tower for the image and video inputs, and TS\_JEV\_TRANSCRIPTION\_URL configured for the audio input. Inline Base64 counts toward the Jev request body limit (8 MiB by default), so use the upload API and file references for larger media. The repo also has ready-to-send mixed-media requests. TensorSharp: https://github.com/zhongkaifu/TensorSharp I’m curious what kinds of decisions you’d want to make from several media types in a single request.

▲
0
 
14👁
r/LocalLLaMA · u/hadoopfromscratch · 13d ago
Customizable harnwsses/coding agents

​ Hi, everyone. I'm wondering how far one can go in customizations of a coding agent. Let's say I want to replace the LLM itself. I can do that with most (all?) harnesses available today. Override the system prompts? Also doable. The tools it uses? It's easy to add new ones via MCP, but when it comes to the basic tools, like read\_file, most harnesses don't let you replace or customize them. Swap a console UI to web UI, afaik, isn't possible. So my question is rather two-sided: First, I'd like to understand what components make a harness a harness. I've named a few (model, tools, UI). Any others worth mentioning? Which of these components would actually work as plugins? Second, which harness is currently the most customizable? My guess would be Pi, but maybe I've missed some less known ones.

▲
0
 
3👁
r/LocalLLaMA · u/TCaschy · 13d ago
Upgrade advice : 2080 ti 22gb or v100 32gb pcie?...

Here's my current setup: Intel® Xeon® E5-2680 v4 x 2, 128 GB DDR4, 1 x 2080 ti 22GB, 1 x 3060 12GB. I'm looking to replace the 3060 with either another modded 2080 ti 22gb or go with the v100 32 gb. Thoughts? My reservation on the v100 are older architecture and heat+fan noise. What say you?

💬 29 (+2) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/MotokoAGI · 13d ago
Jail breaking open models

Is there any resource dedicated to jail breaking open models? reddit, discord, etc? I know some of the models yield easily, but some of them can be stubborn especially the large smarter ones. I have tried uncensored models and while they might pass sometimes, they often end up doing stupid things the censored ones don't. No matter the claims, it seems altering the weights ends up affecting the intelligence. If anyone knows any techniques, please share or point me towards the right resources.