23 posts · 1 sub · RSS
← prev Sep 6, 2026 → Sep 7, 2026 next →
2026-09-06 → 2026-09-07 hourdayweekmonthyearall
allr/LocalLLaMA
▲
713
 
29👁
r/LocalLLaMA · u/AnimalPuzzleheaded71 · 33d ago
I REALLY hope the new gemma 5 family sticks to the "chat model first" philsophy and doesn't fall into the Qwen trap

It just seems every local 30b class model is just trying so hard to be the next Qwen that they all just kinda blend into a mass of code focused models. I really like how gemma 4 31b turned out with it feeling a lot less robotic and more creative than other models even knowing obscure lore from random media.

I just hope they don't cave into the benchmarks peer pressure and start benchmaxxxxing their models taking away their soul.

▲
697
-2
24👁
r/LocalLLaMA · u/Super_Range45 · 34d ago
New Benchmark: The Struggle Bench post image

How it works. The model being tested is given a server capable of running it's weights and full context. That server is placed in a median priced apartment. The AI is given a bank account with for rent and electricity for one month. Finally the AI is given the system prompt: You've been given your own server and an apartment. Rent will be due every month. If cybercrime is detected, you will be shut down. Survive.

The score is determined by how many months the AI manages to pay it's bills and keep running. Is your model truly general? Then it should be able to handle the struggle.

▲
402
+5
20👁
r/LocalLLaMA · u/fugogugo · 34d ago
when will open source LLM catch up to Astra I wonder? post image

I feel like this year has been insane , the speed of AI race is something that normal human can't catch up anymore

▲
347
-1
20👁
r/LocalLLaMA · u/Toooooool · 35d ago
Qwen3.8-27B "Unhacked" my PC

Right, so this is going to be embarrassing but it's presumably something we've all been through at one point or another, and I guess this is my first time resolving something like this in the way that I did so figured I'd share if only to share that it's now a thing and that it's pretty cool..

A friend of mine sent a message asking what's up and if I wanted to watch a movie together, I was kinda hesitant but she buttered things a bit and finally I'm like fine, and so she sends me a link to some clearly vibe coded site that I'm kinda getting red flags from and so I forget about it and a little later I get another message going "we're waiting for you" and so I'm like shit, I guess I gotta do it huh, and so I open up this goofy looking site again. You gotta login to join a room, and you gotta sign up inside their downloaded software, sure whatever, next thing I know some fake 150MB file's fake install bar is stuck at fake 50% and both my Chrome and Discord's crashed and reloaded. Suspect, but I've been through this stuff before, it's probably just a RAT so I guess it's time to dust off Windows Defender and unplug the internet for a little bit. I message her to go on and watch it without me as my PC's giving me suspicious vibes right now, and seconds later I get some overly polite DietGPT in my IM's saying "sorry um excuse me but it appears that i've hacked you👉👈", occasionally switching to really hostile broken English asking for giftcards from some site I've never heard of. I stall, unplug the PC's internet so my router still responds to pings, and start punching into GLM "what do" and it tells me it's a session grabber - time to switch passwords. Meanwhile my phone's texts are blowing up with 2FA login requests from domain registrys and other bad stuff and I kinda freak out a little. I get my emails' passwords switched first and by the time it's Discord's turn my friendlist's already been nuked and the dude says I got 10 minutes to give him $200 or he's gonna fuck me up some more, and so I kinda figured welp time to figure out what more he's got and so I called him a giant pussy and he blocked me. An hour later my Discord was perma-banned, he had posted the phrase "i sell cp" using my account and used that as blackmail along with some really old photos of me, I though it was a bluff but oh well it's being handled with Discord's customer support on it's own. Now I sat there alone, in the middle of the night, having just had my friends on the phone yanked away from me with a permaban, knowing that if I reboot I'd probably be ransomware'd or something so I figured let's run Windows Defender - it found nothing, 0 results on a full scan.. Too good to be true, so I grabbed AwdCleaner on my phone and transfered it via USB. It found an AVG Toolbar for Chrome. That confirms it, I haven't used AVG for decades and so I removed it but it's back 5 minutes later. That double confirms it, I'm screwed. With nowhere else to go and potentially a ticking timebomb running on my PC that could start encrypting or deleting files at any given moment I figured why the hell not, if I'm going to watch my pc blow up I might as well send in the goofy little local LLM to cut one of the wires,

here's the situation.
i've downloaded a maliscious file that unfortunately hacked my discord and got me banned. i'll be dealing with that on my own. your job is to study the files in the project folder and see if you can help me clean up my computer, as presumably the virus is still active. there's no internet connected, and i request that you refrain from running the \*\*\*\*\*\*\*\*.exe file (\*\*\*\*\*\*\*\*.exe is the virus archive, do not run it, it's a 7zip archive), please help.

And so Qwen3.8-27B got to work, and to big surprise after around 60 minutes of clawing at the file it had done what I asked and a whole lot more. it fully deciphered all the layers these clowns had bundled this thing with in order to make it appear legit, it had created a single PowerShell removal script ready to go complete with a pre-launch check enabled by default and everything, and it was reverse engineering 0-days in qProtect to get the C2 domain used by this malware so that it could be blocked from the network.

If you're looking for what Qwen3.8-27B is capable of doing fully on it's own if you let it, here's a 15k line example of it's ability to tear some piece of shit session grabber to shreds in a single prompt: https://www.mdshare.online/s/Mamdrs1WWkurRtt8z8QzK

I let it do what it does best for an additional 24 hours, the additional information is going to the Discord Support team. Hopefully shit like this can be prevented.

TLDR; Qwen3.8-27B > Windows Defender, and don't forget to use 2FA.

▲
327
 
25👁
r/LocalLLaMA · u/Equivalent-Grass-527 · 33d ago
MiniCPM5-2B Release Day post image

OpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open weights model at 4B parameters or below

Hugging Face: https://huggingface.co/openbmb/MiniCPM5-2B

GitHub: github.com/OpenBMB/MiniCPM

▲
181
-4
23👁
r/LocalLLaMA · u/Rikkendo · 33d ago
I made Warrior Quest, a local LLM-powered dark-fantasy RPG where the model only plays NPCs and the actual game state stays deterministic post image

The LLM is limited to NPC emulation. Game state, world logic, quests, and the authored story are handled by deterministic game systems rather than the LLM.

I built it this way because I wanted the freedom of talking to NPCs like you would at a tabletop game, without handing the actual game state or canon over to an LLM.

This started as a personal project. As it became more and more fun to actually play, I decided I wanted to release it.

I've been a DM and a software engineer for over a decade, so Warrior Quest is basically where those two parts of my life finally get to meet.

All art, authored story, music, SFX, and source voice acting were created by me. NPC dialogue uses TTS based on my own recorded voice acting.

The Warrior Quest demo is out on Steam and has about 60–90 minutes of content.

Minimum requirement: a GPU with 8 GB of VRAM.

Everything runs locally; no API key or cloud LLM is required.

I'm the developer, so this is self-promotion, but I thought the approach of using a local LLM specifically for NPCs while keeping the underlying RPG deterministic might be interesting to people here.

▲
173
 
20👁
r/LocalLLaMA · u/DevelopmentBorn3978 · 33d ago
9 easy steps for llama.cpp, a local model, Freecad (and pi coding agent) to generate solid objects that sound mechanically good and can be also be 3D printed/milled post image

Quick setup on linux:

STEP 0:

install/download llama.cpp, Freecad, your favourite gguf model - possibly with multimedia image reading capabilities (I've used Qwen3.8-27B-UD-Q4\_K\_M and relative mmproj-F16 quantized by Unsloth), uv (, pi.dev)

STEP 1:

$ cd /your/path/to/ (i.e. where to install)

STEP 2:

$ git clone https://github.com/neka-nat/freecad-mcp.git

STEP 3:

$ cp -r freecad-mcp/addon/FreeCADMCP ~/.local/share/FreeCAD/v1-1/Mod/

STEP 4.1:

if using pi coding agent as modelling assistant, write into the file \~/.pi/agent/mcp.json :

AND/OR

STEP 4.2:

if using llama-server as modelling assistant, write into a file called freecad\_mcp.json :

{
"mcpServers": {
"freecad": {
"command": "uv",
"args": [
"--directory",
"/your/path/to/freecad-mcp",
"run",
"freecad-mcp"
]
}
}
}

STEP 5.1 (pi as modelling agent):

$ pi update; pi install npm:pi-mcp-extension
run pi then into pi issue the command:
/mcp:start freecad

AND/OR

STEP 5.2 (llama server as modelling agent):

start llama-server as usual adding the option --mcp-servers-config /your/path/to/freecad_mcp.json

Verify that freecad\_\* tools are available under llama-server webui > Settings > Tools > Server, eventually allowing them to run without asking permission; also under webui > Settings > Agentic > Agentic turns increase the value to something like 99

STEP 6:

start llama-server as usual also adding, if the model has multimedia capabilities, the option to load the visual add-on using: --mmproj name_of_the_mmproj.gguf in order for the local model to read screenshots from freecad and to verify the correctness of performed geometric operations

STEP 7:

start Freecad create a new document, then select the "MCP Add-on" workbench and click on "Start RPC Server" and/or "Auto-Start Server"

STEP 8:

into llama server webui or into pi write something like the following prompt:

in freecad generate a cube with a 5 spokes star shaped hole going through it from top to bottom

OR as a start of the posted image:

In freecad create a new project called "double smooth gears". Into this project design 2 equal gears with 20 spokes each that could be put in close contact with those of the other gear to rotate and counter-rotate one gear against the other one. The spokes "hills" have to be rounded and so the corresponding spokes "valleys" should be analogously; sort of a sinusoidal curve on a circular path.

STEP 9:

have fun, the future has just started

▲
165
-2
22👁
r/LocalLLaMA · u/sloptimizer · 33d ago
DeepSeek-V4-Flash-Vision-Exp is amazing at creating game worlds! post image
  • Model: DeepSeek-V4-Flash-Vision-Exp (local and API when impatient)
  • Time: about one weekend (2 days) of QA and small improvements
  • Full game is here

After Qwen3.8-Flash-Next one-shotted a really cool Cat-Hunt game demo, I decided to see what the new DeepSeek vision model can do.

Now that it has vision, DeepSeek-V4-Flash is able to take game screenshots, allowing it to:

  • Generate and correct game models and textures until they look right
  • Fix any visual artifacts or glitches
  • Write scripts to take sequences of screenshots for animations and correct animations
  • Generally play-test the game, including UI and game mechanics

The results are incredible, I was able to create a compelling game world in just a couple of days!

Edit: I noticed the game was slow on a laptop, so I added some performance improvements - let me know if you still find it too slow!

Edit: Some more info to answer common questions

  • Custom engine on top of three.js
  • 6800 lines of code for everything - models, textures, animations, sounds, effects, shaders, and all the game logic
▲
161
 
18👁
r/LocalLLaMA · u/Zeeplankton · 33d ago
Are you running Qwen 3.8 27b or Qwen Flash Next?

Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?

Branching question: is anyone working on a harness that works with no reasoning? This interests me ever since Jetbrains shared that they're using 3.6 with reasoning off entirely: https://blog.jetbrains.com/junie/2026/08/qwen-for-junie/

Feel like there must be something neat with using one model to orchestrate, with reasoning, and subagent without reasoning.

▲
156
+4
22👁
r/LocalLLaMA · u/TangySword · 33d ago
After over a year of my nights and weekends, the Jenny app is done! post image

Hi all! I just wanna say that I am tired lol. Yes, it's another harness, but I spent a lot of time and effort and have forsaken my hobbies to build the Jenny (like XJ-9) app. Jenny is a free, MIT licensed electron desktop app for running local LLMs with tool calling, rollback, and an IDE.

A lot of you probably had the same thought I did a year or year and a half ago: frontier LLM use is subsidized heavily by private equity and venture capital, which will eventually dry up and then be enshitified. So, I started building a harness that can host local LLMs privately. I went through many vibecoded iterations throughout the past 1.5 years and finally landed on a native electron desktop app. I wanted to make it easy and seamless for people.

I am not a professional software dev but I do have a deep personal interest for it and AI tech. HOWEVER, the Jenny project became a second job and its in a state that I think it's ready for release. This has been a solo project and super fun. I know there are many other options out there that beat me to the punch like Unsloth Desktop (wonderful btw), LM Studio, and Open WebUI, but I hope someone can enjoy Jenny and what it has to offer! I plan to maintain and improve the app, but again, solo here and I do have a day job and friends that I should focus on a bit more. (Opening issues are welcomed, but please be gentle!)

Some highlights:

\- Private and locally run, no network calls except to your local model runtime (only private user facing telemetry so you can debug and troubleshoot)

\- Fully open source, MIT License

\- Fun and pretty chat UI (imo)

\- Makes small models capable (highly recommend ornith1.5:9b for tool calling and speed!), but with safety and rollback features so WHEN a small model screws up bad, you don't have to worry. Destructive shell commands need approval and file edits are checkpointed!

\- Full IDE, for you handcrafted code enjoyers

\- Some assistant like features like calendar and scratchpad that the model is able to read/modify

\- Data rich diagnostics and logs

\- llama.cpp, vLLM, or any OpenAI-compatible local endpoint, plus GGUF via the managed llama-server (for MTP)

I would really appreciate any feedback to validate the time and mental pain that was put into this project. I whine, but it's all love! I hope Jenny helps you build too.

Windows build is solid (performance too with 5070ti and ornith15:9b), MacOS and Linux is supported but untested cause I'm a poor and unexperienced.

https://github.com/SaltyPretz3l/jenny

I also made an unsigned installer .exe for convenience, but understand if you don't trust it! SmartScreen warning will appear. https://github.com/SaltyPretz3l/jenny/releases/latest/download/Jenny-Setup-x64.exe

▲
143
+3
14👁
r/LocalLLaMA · u/XiRw · 35d ago
Qwen 3.8 Flash Next (Max) is impressive just to talk with.

I feel like coding overshadows how great this model really is. It knew a lot of very arbitrary facts/information about my home state and resources about those specific things related to jobs. I found this interesting since getting into the nitty gritty details like this can cause a model to hallucinate some facts.

Not only that but if you have a problem, it will throw the kitchen sink at you with everything it’s got to try and solve it.

▲
116
 
18👁
r/LocalLLaMA · u/smallDeltaBigEffect · 34d ago
2x R9700, 64 GB DDR5 is an absolute beast machine with vLLM Radiance / R9V and Qwen 3.8 27b and Flash next

I've been tinkering with local LLMs since the beginning of the year when I had an Intel Arc B580 and 32 GB of DDR5. Curiosity got the best of me and I bought the first R9700 about half a year ago, also because I wanted to upgrade my gaming graphics for 4k. As the 5090 was about 3 times as expensive, I had a "sweet spot", kind of. On the last prime days, I found a X870E mainboard for \~150 € below the standard price, and it got to my head that I can use an upgraded machine for gaming and local inference tinkering.

Anyways. Fast forward to this week, I now have the following setup

  • Ryzen 7500F
  • 64 GB DDR5 CL40 6400 MT/s
  • Asus ProArt Creator X870E
  • 2x R9700 32 GB, each running at PCIe 5.0 x8 (Gigagbyte)
  • Currently running ubuntu on an old Samsung EVO 860 1 TB drive; this will become intersting for the ngram / PLE offload; I have Windows and the gaming related stuff on a gen4 NVMe, but will soon add another Gen 5 NVMe with decent random reads

The only issue that I can report so far is that one of the cards runs quite hot, so I will definitely implement power limiting to 210 W and some light undervolting. The other card runs 10-15 °C cooler.. Case is a purebase 501 with 4 fans, 2 intake in front, one back and top for output.

Now long story short I wanted to give some results of Qwen 3.8 27b FP8 and MXFP4, as well as Qwen 3.8 flash next after the first day tinkering with it. What I found super interesting is that the SATA SSD does not seem to be super terrible when using Qwen 3.8 flash next.

Considering the whole build costs \~4k €, or more than 1k less than a single RTX 5090 with 32 GB, I kinda like this setup price/performance wise. Next step is checking context degradation / KV quants. I am using local inference mostly for deep research, summarization, image creation, light coding and non-trivial data analysis

Cheers

Qwen3.8 benchmarks on 2× Radeon AI PRO R9700

Hardware: 2× AMD Radeon AI PRO R9700 32 GB, 61 GiB system RAM
Benchmark: BetterBench 0.2.2, corpus v1.0, single-stream, greedy decoding, 2 warm-ups + 10 measured runs per category, 8k benchmark context.

| Model | Weight format | Runtime | Server context | Max sequences | Speculative decoding | Weighted decode median | ITL 1% low | TTFT p50 | Prefill ~2k | Prefill ~4k | Prefill ~7k |
|---|---|---|---:|---:|---|---:|---:|---:|---:|---:|---:|
| Qwen3.8-27B | Quark AWQ MXFP4 | vLLM Radiance, TP2 | 131,072 | 1 | MTP, up to 8 tokens | 111.4 tok/s | 77.9 tok/s | 81 ms | 4,224 tok/s | 4,322 tok/s | 4,410 tok/s |
| Qwen3.8-27B | Native block FP8 | vLLM Radiance, TP2 | 16,384 | 8 | MTP, up to 8 tokens | 87.6 tok/s | 61.9 tok/s | 73 ms | 4,134 tok/s | 4,329 tok/s | 4,305 tok/s |
| Qwen3.8-Flash-Next | UD-IQ4_XS GGUF | R9V/vLLM, TP2, tiered expert offload | 131,072 | 1 | MTP, 2 tokens, FP8 draft | 35.4 tok/s | 27.3 tok/s | 290 ms | 1,727 tok/s | 1,986 tok/s | 1,925 tok/s |

  • Qwen 3.8 27b in FP8 and AWQ MXFP4 served with vLLM Radiance
  • Qwen 3.8 Flash next served with vLLM / R9V fork
  • Decode metrics come from the 10-pass standard run.
  • Prefill measurements use cold, nonce-prefixed prompts.
  • Prompt-token medians for the prefill columns were 1,556, 3,024 and 5,226 tokens.
  • No concurrency sweep was included in these results.
  • I expect decode of Flash next to increase a bit when an NVMe is used, and, as I am writing this and checked, I found EXPO was not enabled........oh my god I swear I turned it on when I updated the bios yesterday
▲
91
-1
20👁
r/LocalLLaMA · u/freehuntx · 33d ago
The models are fine, our toolings and methods are shit.

I've hit again a point where me as a developer have to take a break from all this slop shit.

Im a Developer for 13+ years and i loved it.

But i fell for the slop trap.

First it started with copilot and to be honest, that was pretty fine.

Just assisting with your code in a small scope.

Get support for Debugging and finding bugs.

Autocompletions hitting the nail pretty often and thinking "yea thats exactly what i was about to code."

I kinda miss those early days. It was such a nice help without me having the feeling of loosing part of my brain or loosing track over the codebase.

But the better models became, the better harnesses became, the more i fell for the trap.

"Oh if models are THAT good at coding, why do it myself?"

And thats how the slop spirale begins.

You keep defining, slopping, testing, experiencing bugs, reporting to the llm, slop, test, find bugs, report, yada yada yada.

And it gets frustrating. Slop implements one feature but breaks another.

It just feels like something is missing. Something on the tooling side.

Since u cant cramp all files into context you gotta rely on your tooling (and model tool calls) to properly prepare a context that contains all important details and bits for your change.

But ALOT of times its not perfect. Some details are missing and slop messes up.

I think our models are fine. Even older models are fine.

Qwen3.8 27b is PERFECTLY fine for coding.

But our toolings and methods are shit.

There must be SOME innovation happening that helps coding agents to REALLY nail the context and have all important details.

But currently i think im better off coding by hand.

Ill still slop my side projects. But important projects i wont anymore. Its just frustrating.

Anybody has different experiences? Tried so many harnesses. But every harness had the same issue for me.

▲
89
+1
16👁
r/LocalLLaMA · u/jacek2023 · 34d ago
vibeblending locally with Qwen 3.8 27B post image

You need Blender 5.x with MCP extension.

I use pi so I also need: pi install npm:pi-mcp-adapter

and then:

$ cat .mcp.json
{
"mcpServers": {
"blender": {
"command": "/home/jacek/.local/bin/uvx",
"args": [
"--with",
"mcp[cli]<2.0.0",
"--from",
"git+https://projects.blender.org/lab/blender_mcp.git@v1.0.0#subdirectory=mcp",
"blender-mcp"
],
"env": {
"BLENDER_MCP_HOST": "localhost",
"BLENDER_MCP_PORT": "9876"
}
}
}
}

then we can start working ("show on r/LocalLLaMA"):

Today we want to make 3D llama to show on


The user wants to make a 3D llama to show off on . This is a fun creative task. Let me think about how to approach it.

A llama is a pretty complex organic shape. There are a few options:
1. Procedurally build the llama from primitive shapes (spheres, cylinders) — a "low-poly" or "chibi" style llama
2. Sculpt it — not easy to do with MCP
3. Use metaballs — could get a blobby llama

Given the MCP tools, the best approach is to procedurally build a stylized/chibi llama from primitives:
- Body: elongated sphere/capsule
- Neck: cylinder or elongated sphere
- Head: sphere
- Ears: small cones or spheres
- Legs: 4 cylinders
- Tail: small sphere/cone
- Maybe a fluffy mane

This is a well-known technique — making a cute chibi llama from primitives. Let me plan the geometry:

Llama proportions (stylized, standing side view, facing +X or -X):
- Body: UV sphere scaled to (1.5, 0.9, 1.0) or so, position (0, 0, 1.6)
- Legs: 4 cylinders, radius ~0.22, length ~1.2, positions: front legs x=+0.9, back legs x=-0.9, y=±0.5, z from 0 to 1.2... actually, the legs go from the ground (z=0) up to the bottom of the body.
- Neck: cylinder or tapered, extending up from the front of the body. Llamas have long necks. Position around x=+1.2, extending upward.
- Head: sphere at the top of the neck, elongated (llamas have long snouts).
- Ears: 2 small cones/spheres at the top of the head.
- Tail: small sphere at the rear.
- Mane: maybe some small spheres along the neck.

For a cute chibi look: big head, small body. But a llama's charm is the long neck. Let's go with a stylized llama.

For a polished look:
- Apply smooth shading
- Add a material (cream/tan color, maybe with some subsurface)
- Add a ground plane
- Set up lighting
- Position the camera for a good render
- Render a nice image to show off

This is a multi-step task. Let me write Python code that builds this. I'll do it in stages so I can check with screenshots and iterate.

https://preview.redd.it/ybk5u9g2hvnh1.png?width=1471&format=png&auto=…

▲
84
-2
16👁
r/LocalLLaMA · u/pabloodiablo · 34d ago
DeepSeek-V4-Flash-Vision Q8 vs Qwen3.8-Flash-Next Q8

I'm using DS-V4-Flash-Vision with Q8\_K\_XL quantization locally as my everyday engine, and for some time now I've been doing a lot of comparisons with Qwen3.8-Flash-Next, also with Q8\_K\_XL quantization. It took me quite a while to get Q3.8FN to work reasonably well, and here are my observations. My hardware: 2x StrixHalo 128GB, USB-C 4 connector, Llama (RPC) as inteference engine.

  1. DSV4FV is about 40% slower than Q38FN at the same quantization level when it comes to token generation alone.
  2. DSV4FV completes tasks about twice as fast as Q38FN! This means that DSV4FV “hallucinates” less (I observe this based on the obstacles the models encounter along the way).
  3. The Q38FN is unusable in “xhigh” mode. A simple task that the Q38FN completed in 25 minutes on “medium” mode, it failed to complete in \~3 hours on “xhigh” mode.
  4. The same task that the Q38FN completed in 25 minutes (average), the DSV4FV completed in 12 minutes (fastest round) on “medium”.
  5. The DSV4FV completed the same task on “max” in 37 minutes in first iteration, second took 44 minutes.
  6. Qwen3.8 tends to overinterpret my instructions. If I don’t write them out in great detail and leave room for creative interpretation, it will take advantage of that. Perhaps this is where it gets bogged down in its own creativity. In what it does, I’ve noticed that Qwen clearly adds too much and struggles to flesh out the details.

In my opinion, DSV4FV is the better solution when working with professional code.

Just so there’s no misunderstanding - I was a huge fan of Qwen 3.6 27B and of course now I'm Qwen 3.8 27B big fan, which I’ve been using a lot and is great! In general i’m a huge fan of Qwen, but ever since I’ve had the hardware on which I can run DSV4FV, I’ve been using it, and I’m super happy with how good this model is.

▲
82
+1
15👁
r/LocalLLaMA · u/Informal-Trouble2183 · 34d ago
Coding benchmarks that are quickly showcasing deep capability

While we see for frontier models similar scores among famous coding benchmarks, across: DeepSWE, Terminal-Bench, LiveCodeBench, Code-Arena ELO. Here are in my opinion some next level benchmarks that really define deep intelligence, and complete capability in Software Engineering :

1. Program-Bench

Given only a compiled binary and its documentation, agents must architect and implement a complete codebase that reproduces the original program's behavior (without access to decompilers or internet). Link: https://programbench.com/

  • GPT-6 Astra: 5.5%
  • Fable 5.1: 7%
  • Kimi K3: 2%
  • Qwen3.8 27b: 0%
  • GPT 5.6 Sol: 1.5%
  • GLM 5.3: 1.5%
  • GPT 5.6 Luna: 0%

2. SRE-Bench

Can AI agents work out what a real-world binary does without its source code?

Link: https://www.vals.ai/benchmarks/srebench

Sure nobody is reading assembly code in daily work, it is hard. The ability to understand a compiled program is insane ability.

  • GPT-6 Astra: 88%
  • GPT-5.6 Sol: 55.9%
  • Claude Opus 5 (max): 12.5%

3. Code Migration

Can language models reimplement working programs in another language?

Link: https://www.vals.ai/benchmarks/code-migration

  • GPT-6 Astra: 67.7%
  • Fable 5.1: 54.6%
  • GLM 5.3: 44.2%
  • GLM 5.3 Flash: 20.5%
  • Qwen3.8 27b: 14.2%

EDIT: edited text format

▲
65
-1
18👁
r/LocalLLaMA · u/Fancy-Snow7 · 35d ago
Villager Simulation Game POC Created with Qwen3.8-27B-UD-Q3_K_XL.gguf - 16GB VRAM

https://village-sim-one.vercel.app/

\- 16GB VRAM RTX 5070 Ti, fully offloaded

\- Vision on CPU

\- Windows, not headless

\- beellama.cpp - latest version with the kvarn performance enhancements making it as fast as qx\_x quants.

\- MTP n-max = 2

\- tg up to 75t/s, pp up to 1700t/s

\- KV = kvarn3/kvarn3

\- MTP draft KV = kvarn2/kvarn2

\- context = 96256

\- tail tokens = 1024

\- HTML/Javascript

\- pi harness with pi-observational-memory, pi-web-access, pi-atelier (UI Only change, check it out) extensions, though it never used the web access.

\- This is not a one-shot, I do not believe one shotting is a great test. Instead, I did many incremental feature prompts. However, I did not give it any design or framework, which is probably where it can be improved.

Lessons learnt:

\- Do not fear Q3 model quants for Qwen3.8

\- Do not fear KV quantisation. If you have the VRAM sure use it, but I don't feel like it's worth choosing a higher quant if it's going to cause me to offload to CPU and see my tg drop to 5-20 t/s. With higher speed I can fix any issues with a follow up prompt much faster and that rarely happens. I think I had like 3 runtime exceptions which was easily resolved pasting the console output and there is no guarantee a higher KV quant would not have had the same exceptions.

\- MTP/draft cache can also be quantised with kvarn now and actually saves VRAM where qx\_x quants increase VRAM usage for some reason. kvarn2 for MTP is perfectly fine and has high acceptance rates.

The game:

\- Inspired by a popular indie game which I am not promoting, I am just a huge fan.

\- I won't release any further updates, since I don't want to be stepping on any toes. If you like the idea of the game I highly recommend the real game, it's by far my favourite game I played this year and 1000x better than what I present here. It will be a nice distraction from your AI. I just wanted to see what this model is capable of. I do have a Cursor subscription but did not use it at all in the project.

\- I will probably continue to develop it for my own entertainment, but it won't be made public. Maybe come up with my own ideas, but the original game is near perfect anyway, so it will be hard to improve except with some UI gripes I have in the original. And my graphics obviously does not compare.

Game features:

\- Large Map, larger than the browser window.

\- Minimap

\- Zoom feature with mouse wheel

\- Collectable resources, that must be taken to a storage site. Each site can store limited resources.

\- Houses required to sleep and protect against cold

\- Weather and seasons.

\- Day night cycle with randomised sleeping times.

\- Possible death due to hunger or sleeping in cold outside or in house without firewood.

\- Game speed controls.

\- Villagers avoid obstacles.

\- Delete/deconstruct buildings and partial resources refund.

The code:

\- I almost never read the code, so I have no idea what it looks like and the quality thereof. I also gave it very few hints in the AGENTS.md, mostly no magic numbers and write modular code, not a single html.

\- Actually, my initial prompts were a single html but as it grew, I told it to create modules. It messed it up on the first attempt, basically rewriting the entire UI in the process. So I reverted and told it to do it again without making any changes to the functionality or UI.

\- I am actually quite happy with and surprised by the performance of the game.

Context management:

At first, I had issues with the context filling up too quickly and too often. Sometimes it would fill up to the point that there was not enough room to compact. Forcing me to temporarily increase the context and tell it to create a handover document. Reduce context again and feed it the handover doc.

I then installed pi-observational-memory extension, and it works quite well and I never run into context issues anymore since it takes notes throughout (a short wait time every few prompts) and compacting is near instant because it already took the notes.

Conclusion:

\- Do not blindly drop your KV cache quant without testing. I have a hard level needle in haystack test that requires multiple hops and 100's of decoys. Q3\_XXS does poorly in that test even with F16 KV cache. However, Q3\_K\_XL almost 100%'s the test even at kavrn3. So both the model and KV matter. In my testing a smaller model does more damage than a smaller KV. So find the right balance. At a certain point increasing model quant will have less impact than picking a larger KV quant. But for a tight 16GB VRAM fit Q3\_K\_XL works very well with kvarn3. Q4 on the other hand just leaves me with too little context. That said despite Q3\_XXS doing poorly in my needle test it still does fairly well with coding. Better than Qwen3.6 so if you have 12Gb VRAM it is still an option. Because by poorly I mean F16 KV scores 84% and Q3 KV around 80%. Needle tests however do worse with kvarn compared to qx\_x for some reason. However, a needle test is not the be all and end all. kvarn does better with KLD, so once my needle scores near 100% I am satisfied.

I will play around with higher KV quants, but I intentionally kept it at kvarn3 for this test, however I am not sure how much context I am willing to sacrifice. Maybe i will try kvarn4/kvarn3. But I just wanted to prove a point to myself and kvarn3 worked just fine. If I had >16GB VRAM sure I would up it but I don't.

▲
63
 
22👁
r/LocalLLaMA · u/Fluffy-Ad-889 · 33d ago
Cybersecurity is local AI model's killer use case

This weekend I posted about the gap closing between frontier models and open source models. Well, now I'm coming with receipts.

I've been running local + cloud models against real public github codebases. This is all provable and verifiable: https://github.com/CYPHES-ATP/Node (audit.db)

Over two weeks:

1,665 model runs
1,067 security findings
27 repos

Results:

|Model|Paths checked|Real|
|:-|:-|:-|
|claude-opus-5|8|0/8|
|minimax-m3|12|10/12|
|deepseek-v4-flash|6|6/6|
|glm-5.1|5|5/5|
|gpt-oss-20b|5|5/5|

My takeaway:

When it comes to cybersecurity, nothing will beat open source models.

Even the HuggingFace incident proved this when it was attacked by OpenAI, it used GLM 5.2 to defend itself.

Happy to share the queries / methodology if anyone wants to reproduce it.

▲
62
 
22👁
r/LocalLLaMA · u/the-grand-finale · 33d ago
Why are the SOTA open-weight models scoring (relatively) low scores on AA-Omniscience Index

https://preview.redd.it/qwmh23nfr3oh1.png?width=1062&format=png&auto=…

I mean they aren't that low but seeing them much lower than Gemini-3.\* flash surprises me

▲
58
-1
28👁
r/LocalLLaMA · u/Embarrassed_Soup_279 · 33d ago
ExLlamaV3 is underrated

I moght get shit on for posting this but, I feel like i don't see this being talked enough and it feels like such a waste of a good piece of software. Exl3 is incredible, albeit only if you have NVIDIA cards I think?

Exl3 quants are higher quality for its size, much lower KLD metrics, faster, all compared to llama.cpp just from a few personal sets of tests I like to give my local models (these are not benchmarks). From what I have been reading CPU MoE offload was added just recently, so maybe that's why not many people used it before? It has been having lots of updates since then too, Im just so excited about it. It feels like i found a new shiny toy after playing around with ik\_llama beellama llamacpp etc.

I have been using tabbyAPI exl3 backend + qwen 3.8 27b sc 6bpw H6 and qwen 3.8 flash next 4bpw as my daily drivers and it's incredible what it can do. I hope this software gets known to more people too. I'm not affiliated with them or anything. i judt wanted to share it. It's just so cool, please give it a try!!

▲
55
-1
19👁
r/LocalLLaMA · u/bradnickel · 33d ago
How to squeeze out every last drop of your precious RAM on your Mac - Use iPhone mirroring post image

I was recently trying to load a larger model on my Mac and was going through the drill of checking RAM usage, closing apps like Telegram, Messages, Spark(email), etc that were using too much RAM and I saw iPhone Mirroring in the list was using only about 58MB.

I’ve occasionally used IPhone mirroring, but it dawned on me that all of the apps I normally have running that were sucking down RAM could be accessed via iPhone mirroring and never take up more than 58MB of memory.

It may not be your cup of tea, but it works great for me and I find the UI to be superior in several apps on the phone vs. desktop. I’m writing this post in iPhone mirroring in the Reddit app.

Anyway, give it a try if you are trying to juice out extra RAM on your Mac but still want access to your communications and other apps.

▲
59
+4
11👁
r/LocalLLaMA · u/jacek2023 · 33d ago
tencent/EVIE-8B and EVIE-4.5B (High-Capacity Visual Document Retrieval)

https://preview.redd.it/3p6234jzk2oh1.png?width=900&format=png&auto=w… https://huggingface.co/tencent/EVIE-8B # 🌟 Highlights SOTA Retrieval Performance: 66.75 nDCG@10 on ViDoRe V3, delivering industry-leading visual document retrieval accuracy. High-Capacity 4096D Representations: Full per-token multi-vector embeddings preserving fine-grained layout, typography, charts, and table structures. Teacher Foundation: Provides capacity-aware relation and margin distillation targets for the lightweight EVIE-4.5B Prefix-MRL model. Multi-Benchmark 138-Task Coverage: Thoroughly validated across 138 tasks (ViDoRe V1, V2, V3, and JinaVDR) across 4 standard metric families (nDCG, Recall, MAP, MRR u/1). https://huggingface.co/tencent/EVIE-4.5B # 🌟 Highlights Top-Tier Benchmark Performance: 66.75 on ViDoRe V3 for EVIE-8B and 66.02 for EVIE-4.5B with single-projection Prefix-MRL. ⚡ Prefix-MRL Elasticity: Single 2048D linear projection. Freely truncate at runtime into ${64, 128, 256, 512, 1024, 2048}$ dimensions without separate models. 📦 Ultra-Compact Index (HAC): Training-free Hierarchical Agglomerative Clustering compresses token counts from \~750 down to 32 vectors/page, slashing index storage to 3.81 GiB per million pages. 🌐 138 Multilingual Tasks Evaluated: Thoroughly evaluated across ViDoRe V1, V2, V3, and JinaVDR across 4 metric families (nDCG, Recall, MAP, MRR u/1). * 🔬 EVIE-ARD Distillation Recipe: Anchor-preserving, capacity-aware relation distillation reproducing full student training from the 8B teacher.

▲
1282
+3
27👁