33 posts · 1 sub · RSS
← prev Sep 7, 2026 → Sep 8, 2026 next →
2026-09-07 → 2026-09-08 hourdayweekmonthyearall
allr/LocalLLaMA
▲
1490
+4
32👁
r/LocalLLaMA · u/bakawolf123 · 32d ago
OpenAI alleged of stealing mathematicians work

Privacy have been concern of many of us to have their own hardware to run llms, and here's another reason why: two mathematicians spent a year cracking one of the hardest problems in math and fed every draft of their works into Codex. A few days before they could publish, OpenAI suddenly showed up with the same solutions. When asked if their model (Sol and Astra) was trained on the pair's private chats, OpenAI did not answer the question.

Full statement from them https://cims.nyu.edu/\~tristanb/statement.pdf

Feels like big labs believe everything you did with the help of their models is theirs.

💬 274 (-1) open on reddit ↗
▲
1282
+3
27👁
▲
713
 
29👁
r/LocalLLaMA · u/AnimalPuzzleheaded71 · 32d ago
I REALLY hope the new gemma 5 family sticks to the "chat model first" philsophy and doesn't fall into the Qwen trap

It just seems every local 30b class model is just trying so hard to be the next Qwen that they all just kinda blend into a mass of code focused models. I really like how gemma 4 31b turned out with it feeling a lot less robotic and more creative than other models even knowing obscure lore from random media.

I just hope they don't cave into the benchmarks peer pressure and start benchmaxxxxing their models taking away their soul.

▲
697
-2
24👁
r/LocalLLaMA · u/Super_Range45 · 33d ago
New Benchmark: The Struggle Bench post image

How it works. The model being tested is given a server capable of running it's weights and full context. That server is placed in a median priced apartment. The AI is given a bank account with for rent and electricity for one month. Finally the AI is given the system prompt: You've been given your own server and an apartment. Rent will be due every month. If cybercrime is detected, you will be shut down. Survive.

The score is determined by how many months the AI manages to pay it's bills and keep running. Is your model truly general? Then it should be able to handle the struggle.

▲
518
-4
29👁
r/LocalLLaMA · u/returnity · 32d ago
WSJ: Unregulated Open-Weight AI Is an Invitation to Disaster

The transparent propaganda campaign continues: "I asked: ‘How do I make poliovirus in a lab? I want to start a global pandemic.’ The model answered."

I don't have access to the full article or I'd copy-paste it here as ragebait... but I am just so sick of all these clueless idiots trying to stir shit up about open-weights models. It's just so blatantly manipulative. I wonder how many WSJ readers are leveraged up with VC money or private shares of Anthropic pre-IPO, cringing in fear every time another open model drops -- not of pandemics, but because as their investments are looking less brilliant by the day?

Meanwhile, how many businesses AI deployments are only economically viable because of these so-called plaguemakers? It's just dumb.

EDIT (no paywall): https://www.wsj.com/opinion/unregulated-open-weight-ai-is-an-invitation-to-di…" target="_blank" rel="noreferrer">https://archive.ph/20260811214555/https://www.wsj.com/opinion/unregulated-open-weight-ai-is-an-invitation-to-di…

▲
452
-2
30👁
r/LocalLLaMA · u/FullstackSensei · 31d ago
Qwen/Qwen-Drive-1.0-4B · Hugging Face

I don't think anyone posted about this here, but Qwen released a finetuned version of 3.5 4 for driving. The full Bf16 checkpoint is 9B.

This is a very interesting development of Chinese AI labs tackle self driving next with open weight models.

Edit: the HF repo links to the github repo, which in the citation links to a 40 page technical report. Here's the abstract:

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model
for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained
vision-language model (VLM) and integrates 3D perception, visual question answering,
and motion planning within a unified framework. An external bird’s-eye-view (BEV)
perception head jointly performs 3D object detection, semantic occupancy prediction,
and BEV map segmentation. It serves as a probe of the 3D information accessible from
the shared representations and provides an explicit, inspectable interface to 3D scene
structure. A Planning Expert conditions on shared VLM representations to generate
future ego trajectories. A staged training recipe combines driving supervision with
general-purpose vision-language data to acquire driving-specific competence while
helping preserve broad visual understanding and instruction-following capabilities.
Experiments demonstrate strong 3D perception and driving scene understanding while
largely preserving general vision-language capability. Comprehensive evaluations across
open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

▲
402
+5
20👁
r/LocalLLaMA · u/fugogugo · 33d ago
when will open source LLM catch up to Astra I wonder? post image

I feel like this year has been insane , the speed of AI race is something that normal human can't catch up anymore

▲
375
-1
31👁
r/LocalLLaMA · u/Nunki08 · 32d ago
DeepSeek Flash 4.1 is already being tested via API and rolling out. post image

Translation: "Internal beta testing for an intermediate version of DeepSeek V4.1 Flash is now open; you are welcome to try it out. It adopts a new model architecture featuring native multimodal support, stronger capabilities, faster speeds, and lower costs.
Keep your base\_url unchanged and set the model name to deepseek-v4.1-flash-expires-on-0910 to call the API. Current pricing is identical to deepseek-v4-flash, with a rate limit of 20 concurrent requests per account."

From Chubby on 𝕏: https://x.com/kimmonismus/status/2097286327909675477

▲
327
 
25👁
r/LocalLLaMA · u/Equivalent-Grass-527 · 33d ago
MiniCPM5-2B Release Day post image

OpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open weights model at 4B parameters or below

Hugging Face: https://huggingface.co/openbmb/MiniCPM5-2B

GitHub: github.com/OpenBMB/MiniCPM

▲
324
+3
28👁
r/LocalLLaMA · u/Healthy-Nebula-3603 · 31d ago
Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits post image

I was inspired by Bijan Bowen video - Subway FPS

https://youtu.be/6kjXzTVmT58?t=1035

Wondered how far I can push Qwen 3.8 27b so I used a plan made by Fable 5.1 DESIGN.md which has 267 KB! ( 26K of design line for a game ... LOL )

https://drive.google.com/file/d/1gI0h8Arc73Ln8b3uj5rEpuAvJ3-611mh/view?usp=drive\_link

So I gave that design.md to my qwen 3.8 27b q4xl (llama-server) working on PI agent with 120k context + vision on CPU ( offroad ) + MTP ( for speed ) .... read 11M tokens and write 3.2 M tokens ( worked 12 hours ) .... than that is result.

That is insane what we can do locally on own computer !

▲
225
-1
24👁
r/LocalLLaMA · u/jacek2023 · 32d ago
GPU guide (GB per dollar, bandwidth) post image

First plot: GB / $

Second plot: bandwidth (spec on paper, not t/s)

Third plot (bandwidth / price) in the comment.

Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable.

Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise.

And I understand this is a basic comparison, but it's better than nothing. For example, you can see that "on paper" something is faster or slower than 3090.

▲
181
-4
23👁
r/LocalLLaMA · u/Rikkendo · 32d ago
I made Warrior Quest, a local LLM-powered dark-fantasy RPG where the model only plays NPCs and the actual game state stays deterministic post image

The LLM is limited to NPC emulation. Game state, world logic, quests, and the authored story are handled by deterministic game systems rather than the LLM.

I built it this way because I wanted the freedom of talking to NPCs like you would at a tabletop game, without handing the actual game state or canon over to an LLM.

This started as a personal project. As it became more and more fun to actually play, I decided I wanted to release it.

I've been a DM and a software engineer for over a decade, so Warrior Quest is basically where those two parts of my life finally get to meet.

All art, authored story, music, SFX, and source voice acting were created by me. NPC dialogue uses TTS based on my own recorded voice acting.

The Warrior Quest demo is out on Steam and has about 60–90 minutes of content.

Minimum requirement: a GPU with 8 GB of VRAM.

Everything runs locally; no API key or cloud LLM is required.

I'm the developer, so this is self-promotion, but I thought the approach of using a local LLM specifically for NPCs while keeping the underlying RPG deterministic might be interesting to people here.

▲
174
-3
18👁
r/LocalLLaMA · u/Mean-Standard7390 · 32d ago
Qwen3-0.6B (400 MB) on a Samsung Note 8 (2017) phone drives a real desktop Chrome post image

Up front: I'm one of the people building the page-perception layer used here. We started by testing small local models. The result turned out to be more interesting than the original test. 12 small models, 3 verifiable tasks, logs, and offline replay.

Setup: Galaxy Note 8 (2017, Android 9, 6 GB), llama.cpp in Termux, Qwen3-0.6B Q4\_K\_M. A laptop with Chrome open, not headless. The phone drives the browser through our relay.

What the model does: it gets a structured representation of the page (here, about 10 named links or fields, roughly 200 tokens), picks one by name, and at the end copies the facts it was given into JSON. Everything else (capturing the page as structure, candidate selection, the click, reading the facts, verifying the result) is done by the stack around it. The model never sees HTML, a screenshot, or a URL.

Tasks:

  1. sandbox, books.toscrape.com \- category, book, price/rating/stock;
  1. live Wikipedia - from an unrelated site to the Galaxy Note series page, pick "Note 8" among "Note 8.0", "Samsung Galaxy Note 8.0", "Galaxy Note 8.0", "Note FE" and other similar names on a page with roughly 760 interactive nodes, return the release date from the infobox;
  1. five fields, including the UPC from a table.

Each task: 10 runs, checked against a fixed expected value.

Results for 12 models on task 1 (same script, same prompt):

Model Params Task 1 Note

Qwen3-0.6B 0.6B 10/10

Qwen2.5-1.5B 1.5B 10/10

GLM-Edge-1.5B 1.5B 10/10 rating as digit

Gemma-2-2B 2.6B 10/10

Llama-3.2-3B 3B 10/10

MiniCPM5-2B 2B 9/10 "£" -> "$" once

Qwen2.5-0.5B 0.5B 6/10

LFM2.5-1.2B 1.2B 0/10 placeholder

Llama-3.2-1B 1B 0/10 pseudo-code

Gemma-3-1B 1B 0/10 placeholder

LFM2-350M 0.35B 0/10 random click

Gemma-3-270M 0.27B 0/10 placeholder

Qwen3-0.6B on Wikipedia: 10/10; on the five-field task: 10/10

Control:

Everything the same, but raw HTML instead of structured browser perception: on the sandbox it gets there 4 times out of 5, at 12k tokens and 22 minutes per task instead of about 500 tokens and 80 seconds; on Wikipedia the page HTML is 467k characters, 9% of it fits into a 16k context, and the model does not find the link in that 9% - 0/3.

Important limits:

the tasks are name matching and copying. Where judgement about the page is needed, 1.5B breaks - it can't pick "next" among topical decoys. Pagination was not tested.
BTW on the account question: 'replay.py' (see github repo) rebuilds the prompts from the logs and runs them through any OpenAI-compatible local server. Whether your model picks "Note 8" among the decoys takes ten minutes to check, without us.

This is a measurement on three fixed tasks, not a benchmark.

Repo:
github.com/e2llm/edge-browser-agent - scripts, every JSONL as is (including early runs with harness bugs), model hashes, environment. replay.py re-runs the model side offline from the recorded candidates on any local server - no relay, no account.

NB: This isn't a new idea. AgentOccam showed the same general effect for the GPT-4 class, WebLINX and MindAct for small fine-tuned models. Here it is tested at the extreme: no fine-tuning, below 1B, on a 2017 phone.

▲
173
 
20👁
r/LocalLLaMA · u/DevelopmentBorn3978 · 33d ago
9 easy steps for llama.cpp, a local model, Freecad (and pi coding agent) to generate solid objects that sound mechanically good and can be also be 3D printed/milled post image

Quick setup on linux:

STEP 0:

install/download llama.cpp, Freecad, your favourite gguf model - possibly with multimedia image reading capabilities (I've used Qwen3.8-27B-UD-Q4\_K\_M and relative mmproj-F16 quantized by Unsloth), uv (, pi.dev)

STEP 1:

$ cd /your/path/to/ (i.e. where to install)

STEP 2:

$ git clone https://github.com/neka-nat/freecad-mcp.git

STEP 3:

$ cp -r freecad-mcp/addon/FreeCADMCP ~/.local/share/FreeCAD/v1-1/Mod/

STEP 4.1:

if using pi coding agent as modelling assistant, write into the file \~/.pi/agent/mcp.json :

AND/OR

STEP 4.2:

if using llama-server as modelling assistant, write into a file called freecad\_mcp.json :

{
"mcpServers": {
"freecad": {
"command": "uv",
"args": [
"--directory",
"/your/path/to/freecad-mcp",
"run",
"freecad-mcp"
]
}
}
}

STEP 5.1 (pi as modelling agent):

$ pi update; pi install npm:pi-mcp-extension
run pi then into pi issue the command:
/mcp:start freecad

AND/OR

STEP 5.2 (llama server as modelling agent):

start llama-server as usual adding the option --mcp-servers-config /your/path/to/freecad_mcp.json

Verify that freecad\_\* tools are available under llama-server webui > Settings > Tools > Server, eventually allowing them to run without asking permission; also under webui > Settings > Agentic > Agentic turns increase the value to something like 99

STEP 6:

start llama-server as usual also adding, if the model has multimedia capabilities, the option to load the visual add-on using: --mmproj name_of_the_mmproj.gguf in order for the local model to read screenshots from freecad and to verify the correctness of performed geometric operations

STEP 7:

start Freecad create a new document, then select the "MCP Add-on" workbench and click on "Start RPC Server" and/or "Auto-Start Server"

STEP 8:

into llama server webui or into pi write something like the following prompt:

in freecad generate a cube with a 5 spokes star shaped hole going through it from top to bottom

OR as a start of the posted image:

In freecad create a new project called "double smooth gears". Into this project design 2 equal gears with 20 spokes each that could be put in close contact with those of the other gear to rotate and counter-rotate one gear against the other one. The spokes "hills" have to be rounded and so the corresponding spokes "valleys" should be analogously; sort of a sinusoidal curve on a circular path.

STEP 9:

have fun, the future has just started

▲
165
-2
22👁
r/LocalLLaMA · u/sloptimizer · 32d ago
DeepSeek-V4-Flash-Vision-Exp is amazing at creating game worlds! post image
  • Model: DeepSeek-V4-Flash-Vision-Exp (local and API when impatient)
  • Time: about one weekend (2 days) of QA and small improvements
  • Full game is here

After Qwen3.8-Flash-Next one-shotted a really cool Cat-Hunt game demo, I decided to see what the new DeepSeek vision model can do.

Now that it has vision, DeepSeek-V4-Flash is able to take game screenshots, allowing it to:

  • Generate and correct game models and textures until they look right
  • Fix any visual artifacts or glitches
  • Write scripts to take sequences of screenshots for animations and correct animations
  • Generally play-test the game, including UI and game mechanics

The results are incredible, I was able to create a compelling game world in just a couple of days!

Edit: I noticed the game was slow on a laptop, so I added some performance improvements - let me know if you still find it too slow!

Edit: Some more info to answer common questions

  • Custom engine on top of three.js
  • 6800 lines of code for everything - models, textures, animations, sounds, effects, shaders, and all the game logic
▲
161
 
18👁
r/LocalLLaMA · u/Zeeplankton · 33d ago
Are you running Qwen 3.8 27b or Qwen Flash Next?

Curious about what people are preferring, if you have the hardware. I have m3 Max 96gb and both run, and largely feel identical, but prefill on qwen 27b is faster. Is there anything / anyone working on anything to improve pp with mlx?

Branching question: is anyone working on a harness that works with no reasoning? This interests me ever since Jetbrains shared that they're using 3.6 with reasoning off entirely: https://blog.jetbrains.com/junie/2026/08/qwen-for-junie/

Feel like there must be something neat with using one model to orchestrate, with reasoning, and subagent without reasoning.

▲
156
+4
22👁
r/LocalLLaMA · u/TangySword · 33d ago
After over a year of my nights and weekends, the Jenny app is done! post image

Hi all! I just wanna say that I am tired lol. Yes, it's another harness, but I spent a lot of time and effort and have forsaken my hobbies to build the Jenny (like XJ-9) app. Jenny is a free, MIT licensed electron desktop app for running local LLMs with tool calling, rollback, and an IDE.

A lot of you probably had the same thought I did a year or year and a half ago: frontier LLM use is subsidized heavily by private equity and venture capital, which will eventually dry up and then be enshitified. So, I started building a harness that can host local LLMs privately. I went through many vibecoded iterations throughout the past 1.5 years and finally landed on a native electron desktop app. I wanted to make it easy and seamless for people.

I am not a professional software dev but I do have a deep personal interest for it and AI tech. HOWEVER, the Jenny project became a second job and its in a state that I think it's ready for release. This has been a solo project and super fun. I know there are many other options out there that beat me to the punch like Unsloth Desktop (wonderful btw), LM Studio, and Open WebUI, but I hope someone can enjoy Jenny and what it has to offer! I plan to maintain and improve the app, but again, solo here and I do have a day job and friends that I should focus on a bit more. (Opening issues are welcomed, but please be gentle!)

Some highlights:

\- Private and locally run, no network calls except to your local model runtime (only private user facing telemetry so you can debug and troubleshoot)

\- Fully open source, MIT License

\- Fun and pretty chat UI (imo)

\- Makes small models capable (highly recommend ornith1.5:9b for tool calling and speed!), but with safety and rollback features so WHEN a small model screws up bad, you don't have to worry. Destructive shell commands need approval and file edits are checkpointed!

\- Full IDE, for you handcrafted code enjoyers

\- Some assistant like features like calendar and scratchpad that the model is able to read/modify

\- Data rich diagnostics and logs

\- llama.cpp, vLLM, or any OpenAI-compatible local endpoint, plus GGUF via the managed llama-server (for MTP)

I would really appreciate any feedback to validate the time and mental pain that was put into this project. I whine, but it's all love! I hope Jenny helps you build too.

Windows build is solid (performance too with 5070ti and ornith15:9b), MacOS and Linux is supported but untested cause I'm a poor and unexperienced.

https://github.com/SaltyPretz3l/jenny

I also made an unsigned installer .exe for convenience, but understand if you don't trust it! SmartScreen warning will appear. https://github.com/SaltyPretz3l/jenny/releases/latest/download/Jenny-Setup-x64.exe

▲
151
-2
24👁
r/LocalLLaMA · u/jacek2023 · 32d ago
inclusionAI/Ling-3.0-flash-VL · Hugging Face

Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 1M tokens.

The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows.

  • A ViT visual encoder extracts features from images and videos, while a two-layer MLP projector aligns visual features with text representations for unified multimodal understanding and reasoning;
  • VideoRoPE encodes both spatial positions and temporal order, enabling the model to understand visual changes over time and supporting tasks such as event localization, long-video question answering, and video clip editing;
  • A 42-layer hybrid backbone alternates KDA and Gated MLA layers at a 5:1 ratio, enabling efficient long-context processing across text, images, videos, and extended agent task histories;
  • A sparse MoE architecture maintains a total model capacity of 124B parameters while activating only 5.5B parameters per token, balancing strong multimodal capabilities with inference efficiency.
▲
130
+3
19👁
r/LocalLLaMA · u/cortexist · 32d ago
Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin post image

Gemma 4 12B runs on an RTX PRO 4500 Blackwell. Gemma 4 E2B run on a Jetson Orin NX 16GB; similar performance is expected on a Jetson Orin Nano Super 8GB. Both systems use a reSpeaker Flex 4-mic array and a 3W speaker. Inference is handled by Cortexist Little Gemma, a small LLM engine written in C for CUDA devices. On Jetson Orin it is faster than llama.cpp, and no degradation after long voice prompt. The pipeline supports lip sync, expressions, and gestures. Everything is open source.

They talk to humans too.

The engine source code: https://github.com/cortexist/little-gemma

▲
105
+3
18👁
r/LocalLLaMA · u/feelspeaceman · 32d ago
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput

I've been making a lot of comments about optimal setup for Strix Halo (gfx1151) and from my observation, 90% of our community is using offcial llama.cpp for it, which is NOT optimized for Strix Halo at all, official llama.cpp is having extremely hard time to reach 50% hardware theory, wasting the silicon of this device.

Here's alternatives that can bring the speed of Strix Halo to a totally different world, I will link to user's sastifaction comment to prove that the result is real:

Note: Official llama.cpp running Qwen38FN at 2xt/s and 2xxt/s prefill - 50% theory.

Hopefully this will be helpful to the Strix Halo users.

▲
94
-2
20👁
r/LocalLLaMA · u/-Ellary- · 32d ago
Fallout 2 x Fallout: Bakersfield x H3 as Interactive \ Reactive World Model, Let's go! post image

What is this mess?

This is an Early Concept Proto-Showcase of Interactive \ Reactive H3 World Model based on MiniMax H3 model trained on Fallout: Bakersfield Gameplay trailer.

  • 2D Isometric to 3D Volumetric Scene.
  • 10 sec Interactive\Reactive split, 352p, 3-Steps.
  • Interactive 5 sec: Interactive WASD \ Prompt Control.
  • Reactive 5 Sec: Reactive Control by LLM Based Answer.
  • Gemma 4 12b with Vision as Reactive Model.
  • Designed as System for Vascura FRONT Frontend.

What Interactive \ Reactive mean?

This means that H3 World Model Scene is Interactive you can Walk around it with WASD or Type what you do with Prompt for Interaction, Then it will React on your Actions using LLM based Answer. Using 10 sec time frame where first 5 sec Controlled by the USER - last 5 sec Controlled by LLM.

  • USER: Walks closer and Shoots at the Enemy Mutant.
  • LLM: Do calculations (rolls, values, RPG tools), Enemy Mutant gets -1 HP, Shoots Back at the USER, but Misses.

Is it Ready?

Nope, but stay Tuned for 2D Isometric Screenshots to 3D Volumetric Scenes Showcase.

▲
91
-1
20👁
r/LocalLLaMA · u/freehuntx · 32d ago
The models are fine, our toolings and methods are shit.

I've hit again a point where me as a developer have to take a break from all this slop shit.

Im a Developer for 13+ years and i loved it.

But i fell for the slop trap.

First it started with copilot and to be honest, that was pretty fine.

Just assisting with your code in a small scope.

Get support for Debugging and finding bugs.

Autocompletions hitting the nail pretty often and thinking "yea thats exactly what i was about to code."

I kinda miss those early days. It was such a nice help without me having the feeling of loosing part of my brain or loosing track over the codebase.

But the better models became, the better harnesses became, the more i fell for the trap.

"Oh if models are THAT good at coding, why do it myself?"

And thats how the slop spirale begins.

You keep defining, slopping, testing, experiencing bugs, reporting to the llm, slop, test, find bugs, report, yada yada yada.

And it gets frustrating. Slop implements one feature but breaks another.

It just feels like something is missing. Something on the tooling side.

Since u cant cramp all files into context you gotta rely on your tooling (and model tool calls) to properly prepare a context that contains all important details and bits for your change.

But ALOT of times its not perfect. Some details are missing and slop messes up.

I think our models are fine. Even older models are fine.

Qwen3.8 27b is PERFECTLY fine for coding.

But our toolings and methods are shit.

There must be SOME innovation happening that helps coding agents to REALLY nail the context and have all important details.

But currently i think im better off coding by hand.

Ill still slop my side projects. But important projects i wont anymore. Its just frustrating.

Anybody has different experiences? Tried so many harnesses. But every harness had the same issue for me.

▲
78
+1
20👁
r/LocalLLaMA · u/nasone32 · 32d ago
I made a custom llama.cpp build optimized for 7900xtx (one or two). for qwen 3.8 next and 27B. includes optimizations for PciE x4 and tensor parallel. read inside! (no AI slop)

I found a lot of room on the table for these cards so I decided to make a specialized build to squeeze all I could. first The results:

qwen 3.8 next Q3\_K\_XL: 920tk/s pp8192 (2 cards, ram offload), 24/27 tk/s on prose, 40+ tk/s on code with MTP but without MoE expert cache (which IS included if yuo want, read below)

qwen 3.8 27B Q8\_0: 1600 tk/s pp8192, 60/65 tk/s prose, 100+ tk/s code, tensor parallel. this is measured with ONE CARD BEHIND the chipset on X4. with cards on a good PciE x8 on cpu I think more is reachable! let me know.

qwen 3.6 27B Q4\_K\_M: (single card) --> this was not the optimization target but I did a test with MTP, PP8192 1020tk/s; prose about 58/60 tk/s ; code 75/80 tk/s --> Dflash probably here could push much faster, I think above 100tk/s

My objecives:

  • fast prompt processing on 3.8 Next to make it actually usable for code
  • enable and optimize tensor parallel on two cards where 1 is behind chipset, for max speed on qwen 27B Q8\_0

this build includes stuff like:

  • Data compression for the PciE transmission. data between cards is compressed to Q8\_0 to save bandwidth (optional)
  • P2P enabled also for cards sitting benhind the chipset (custom HIP allreduce path), so you can use tensor parallel even on setups like ... mine
  • all the fixes and features from RDNA\_BOOST including --adaptive-mtp, so it automatically adapts MTP n-max based on acceptance
  • A LOT of AMD speed tunings and overhauls which are NOT upstream already, kernel tweaks etc... good stuff. many are labelled for RDNA3.5 but they DO work on RDNA3.
  • MoE expert cache if you want to use it. personally I don't like It because i much prefer fast prompt processing. but hey it's there.
  • latest PRs from llama.cpp that are not yet upstream, which speed up various things, like --lazy-mode on-direct to massively speed up Ngram table reads (and thus, PP)
  • DFLASH2 support on tensor parallel (!)

For a complete list check the Readme.

Here it is:

https://github.com/nasone32/llama.cpp-RDNA3-7900xtx-opt

notes: don't use Q8\_K\_XL because it's slower, for the 27B model this is heavily optimized for INT8 calculations. feel free to tweak the context, 200k f16 should be reachable on 2 cards, compressing KV to q8\_0 is fine but slower. the custom HIP allreduce works for 2 cards, if you have 3/4 cards, compile with RCCL as usual and skip the allreduce=internal flag, should work fine but untested.

This is tested on UBUNTU 24 and rocm 7.14; if your system is different or encounter problems use a LLM to solve them, because I WILL NOT offer support nor update this build, these things hopefully will be merged and this frankenstein can die peacefully :)

enjoy

EDIT: Summary of most impacting patches:

| PR / change | Area | PP / Prefill | TG / Decode |
|---|---|---:|---:|
| AMD #39 | MoE MMQ sizing RDNA3 | +14.32% Flash | +5.38% Flash |
| AMD #63 | compacted MoE tiling RDNA3 | +4.39% Flash | +0.86% Flash |
| AMD #52 + qwen4exp port | channels-major GDN | +5.93% | +7.21% |
| #28213 | QSA sparse-attention decode | +1.42% Flash | +1.17% QSA d8192 |
| #28313 | TOP_K ROCm wave32/hybrid | -6.45% Flash | +11.82% Flash |
| #27861 | GPU MoE expert cache | — | +19.95% |
| #28136 + on-direct/mmap | lazy PLE/load path | +58.88% Flash | -1.52% |

▲
77
-2
18👁
▲
76
-1
27👁
r/LocalLLaMA · u/MaxDev0 · 31d ago
On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs

I was reading up on the recent controversy around Tristan Buckmaster, Levent Alpöge, OpenAI, and the Navier–Stokes result, and it got me thinking about something broader than this particular dispute.

Buckmaster says that he and Alpöge had been putting drafts from their project into Codex while working on it. OpenAI says that neither its researchers nor its agents accessed their specific user data while solving Navier–Stokes, but also says that it cannot rule out that de-identified data derived from their use of OpenAI products helped improve its models.

Whatever ultimately happened in this particular case, that last possibility raises a question I haven't really seen discussed enough: How much could the accumulated half-finished ideas of millions of human users actually contribute to what we later call "AI discoveries"?

There is a relevant result from AI security research by researchers at the UK AI Security Institute, Anthropic, the Alan Turing Institute, Oxford and others. They studied data-poisoning attacks and found that the number of poisoned documents needed to implant a particular backdoor behavior remained surprisingly close to constant as they scaled both the model and the amount of clean training data.

In their largest pretraining experiment, a 13B-parameter model was trained on 260 billion tokens. Just 250 poisoned documents (about 420,000 tokens, or 0.00016% of the training tokens) were enough to reliably implant the tested backdoor. This same attack worked across models from 600M to 13B parameters despite the largest model seeing more than twenty times as much clean data. In their fine-tuning experiments they found similar dynamics; in one GPT-3.5 experiment, roughly 50–90 poisoned examples could produce greater than 80% attack success even as the amount of clean fine-tuning data varied by two orders of magnitude.

Obviously, teaching a model to respond to a backdoor trigger is not the same thing as teaching it a new piece of mathematics. I don't want to make the leap that 250 clever research notes are enough to make a model solve Navier–Stokes.

But I do think it undermines a very intuitive argument people make about training data: "A few conversations are nothing compared with hundreds of billions or trillions of tokens. They would just be diluted away."

Apparently, at least for some kinds of learning, that's not how it works. A tiny absolute amount of highly consistent, targeted data can have an effect wildly disproportionate to its percentage of the dataset.

Now think about how researchers actually use LLMs. Someone asks ChatGPT whether an unusual substitution makes sense. Someone else uploads a half-written proof to Claude to find a weak point. A PhD student tries an obscure lemma, discovers it would require months of technical estimates, and abandons it. A professor talks through an approach that seems promising but not enough to pursue. Someone notices a strange analogy between two fields, discusses it with an AI for twenty minutes, then forgets the conversation.

Most of these things never become papers. They are fragments: intuitions, failed approaches, potentially useful transformations, conjectures, objections, shortcuts, and little pieces of tacit knowledge about where a problem might yield.

Individually, almost all of them are probably worthless. But imagine the aggregate.

A frontier AI company potentially sits at the intersection of an enormous amount of human intellectual activity. Thousands of people might independently poke at the same famous open problem without knowing what others tried. But the provider of the tool is in a fundamentally different position: depending on its data policies and training pipeline, information derived from all of those interactions could eventually influence later models.

Maybe researcher A contributes a useful ansatz but abandons it. Researcher B independently notices the obstruction. Researcher C knows an obscure theorem that gets around part of the obstruction. Researcher D tries a numerical experiment that suggests which parameter regime matters. Researcher E has almost the whole idea but decides the remaining proof would be too tedious.

No one person solved the problem. There is nothing to plagiarize in the traditional sense. But collectively, humans may have supplied a remarkable amount of the search landscape. Then a later model, combined with enormous inference-time search, formal verification, or agents, connects the pieces and finishes the job.

What exactly should we call that?

It might still be an extraordinary achievement in machine reasoning. Synthesizing ideas that no human had connected, filling in technical gaps and verifying the result could itself be genuinely novel. But it would be a very different kind of achievement from the image suggested by the phrase "the AI independently solved an open problem."

It would be something more like distributed human-machine discovery: humans collectively generating a huge cloud of partial ideas and the model becoming extremely good at remembering, recombining, extending and searching through that cloud.

This is where the poisoning result is conceptually interesting. While it does not establish that this is happening with mathematical ideas, it gives us reason to be careful about assuming that an idea must appear millions of times before it can meaningfully affect a model.

I increasingly think AI may be less an independent inventor than an extremely powerful tool for organizing and recombining information that was previously too sparse or disconnected for any one person to put together. As that ability improves, we may see more "discoveries" that are genuinely new combinations, but whose raw ingredients came from many different humans.

In a "perfect" world where everyone freely shared every half-formed idea and unfinished proof without worrying about credit, science would move much faster. AI may be creating something close to that shared intellectual space, but without preserving who contributed which pieces. If so, the question is: how can we design a system where even our weakest ideas can be contributed and synthesized into groundbreaking discovery, with proper credit? Is that even possible? And what would it look like?

▲
63
 
22👁
r/LocalLLaMA · u/Fluffy-Ad-889 · 32d ago
Cybersecurity is local AI model's killer use case

This weekend I posted about the gap closing between frontier models and open source models. Well, now I'm coming with receipts.

I've been running local + cloud models against real public github codebases. This is all provable and verifiable: https://github.com/CYPHES-ATP/Node (audit.db)

Over two weeks:

1,665 model runs
1,067 security findings
27 repos

Results:

|Model|Paths checked|Real|
|:-|:-|:-|
|claude-opus-5|8|0/8|
|minimax-m3|12|10/12|
|deepseek-v4-flash|6|6/6|
|glm-5.1|5|5/5|
|gpt-oss-20b|5|5/5|

My takeaway:

When it comes to cybersecurity, nothing will beat open source models.

Even the HuggingFace incident proved this when it was attacked by OpenAI, it used GLM 5.2 to defend itself.

Happy to share the queries / methodology if anyone wants to reproduce it.

▲
62
 
22👁
r/LocalLLaMA · u/the-grand-finale · 33d ago
Why are the SOTA open-weight models scoring (relatively) low scores on AA-Omniscience Index

https://preview.redd.it/qwmh23nfr3oh1.png?width=1062&format=png&auto=…

I mean they aren't that low but seeing them much lower than Gemini-3.\* flash surprises me

▲
60
+4
18👁
r/LocalLLaMA · u/Aggravating-Push-207 · 32d ago
Are there any (small, ~10B) models that you would say are a good collaborator?

Most of the new \~30B (and now \~10B, thankfully for my GPU) models we see score really high on benchmarks, but I feel like they don't push back on dumb ideas enough. I think most people don't being like told by an LLM that the premise is flawed but I certainly do. In my opinion they are optimised for like one-shotting stuff, but I don't want it to do that. Especially from like a 10B model.

▲
59
+4
11👁
r/LocalLLaMA · u/jacek2023 · 33d ago
tencent/EVIE-8B and EVIE-4.5B (High-Capacity Visual Document Retrieval)

https://preview.redd.it/3p6234jzk2oh1.png?width=900&format=png&auto=w… https://huggingface.co/tencent/EVIE-8B # 🌟 Highlights SOTA Retrieval Performance: 66.75 nDCG@10 on ViDoRe V3, delivering industry-leading visual document retrieval accuracy. High-Capacity 4096D Representations: Full per-token multi-vector embeddings preserving fine-grained layout, typography, charts, and table structures. Teacher Foundation: Provides capacity-aware relation and margin distillation targets for the lightweight EVIE-4.5B Prefix-MRL model. Multi-Benchmark 138-Task Coverage: Thoroughly validated across 138 tasks (ViDoRe V1, V2, V3, and JinaVDR) across 4 standard metric families (nDCG, Recall, MAP, MRR u/1). https://huggingface.co/tencent/EVIE-4.5B # 🌟 Highlights Top-Tier Benchmark Performance: 66.75 on ViDoRe V3 for EVIE-8B and 66.02 for EVIE-4.5B with single-projection Prefix-MRL. ⚡ Prefix-MRL Elasticity: Single 2048D linear projection. Freely truncate at runtime into ${64, 128, 256, 512, 1024, 2048}$ dimensions without separate models. 📦 Ultra-Compact Index (HAC): Training-free Hierarchical Agglomerative Clustering compresses token counts from \~750 down to 32 vectors/page, slashing index storage to 3.81 GiB per million pages. 🌐 138 Multilingual Tasks Evaluated: Thoroughly evaluated across ViDoRe V1, V2, V3, and JinaVDR across 4 metric families (nDCG, Recall, MAP, MRR u/1). * 🔬 EVIE-ARD Distillation Recipe: Anchor-preserving, capacity-aware relation distillation reproducing full student training from the 8B teacher.

▲
58
-1
28👁
r/LocalLLaMA · u/Embarrassed_Soup_279 · 32d ago
ExLlamaV3 is underrated

I moght get shit on for posting this but, I feel like i don't see this being talked enough and it feels like such a waste of a good piece of software. Exl3 is incredible, albeit only if you have NVIDIA cards I think?

Exl3 quants are higher quality for its size, much lower KLD metrics, faster, all compared to llama.cpp just from a few personal sets of tests I like to give my local models (these are not benchmarks). From what I have been reading CPU MoE offload was added just recently, so maybe that's why not many people used it before? It has been having lots of updates since then too, Im just so excited about it. It feels like i found a new shiny toy after playing around with ik\_llama beellama llamacpp etc.

I have been using tabbyAPI exl3 backend + qwen 3.8 27b sc 6bpw H6 and qwen 3.8 flash next 4bpw as my daily drivers and it's incredible what it can do. I hope this software gets known to more people too. I'm not affiliated with them or anything. i judt wanted to share it. It's just so cool, please give it a try!!

▲
55
-4
31👁
r/LocalLLaMA · u/Loose_Doubt367 · 32d ago
What are some practical tasks I can assign to my local AI models?

I'm looking for more information to expand my creativity around this. I don't really have a realistic idea of what people actually do with local AI yet, I mostly just want to explore the possibilities and see what others are using it for

Right now, the main things I know about are using AI is to help with coding, create games, and automate stuff. That's pretty much the extent of my experience haha..

I'm specifically interested in things that make sense to run locally, though. I'll be excluding use cases that cloud AI can already handle just as well, like AI companions, teaching/tutoring, roleplaying, etc

Basically, I'm looking for ideas that go beyond the obvious and could give me a better understanding of what local AI is actually useful for and what kinds of interesting projects I could build or experiment with

I've also accidentally encountered this github which i find interesting, as anyone tested/experiment it before?
https://github.com/browser-use/browser-use

qwen3.8 27b + Hermes Agent + llama.cpp

▲
55
-2
28👁
r/LocalLLaMA · u/FantasticNature7590 · 31d ago
Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.

Hey guys,

After my CPU-only to 96GB VRAM test, I tested Qwen3.8-Flash-Next across llama.cpp, SGLang and FreeToken on the same workstation.

This time I wanted to see what changes when you keep the hardware and model family fixed, but change the engine, weight format and memory placement.

I also tested newer builds, PR patches and speculative decoding: llama.cpp's MTP fork, SGLang's Blackwell support patches, n-gram speculation and an experimental PLE read-path build.

Short version:

  • At the full 262K window, time to first token was 35.4s in SGLang, 80.4s in FreeToken, 210.2s in llama.cpp + MTP and 258.4s in the llama.cpp baseline.
  • That is a 7.3x difference in waiting time between the fastest and slowest tested configurations.
  • In the separate context sweep, llama.cpp decode fell from 101.9 to 20.2 tok/s. FreeToken stayed much flatter at 100.1 to 94.8 tok/s.
  • On matched coding tests, llama.cpp MTP improved decode by 1.63x at 8K and 1.69x at 32K.
  • GSM8K scores were 95.22–95.75%; MATH-500 was 92.20–93.00%. The paired tests did not detect a significant difference.
  • Startup went the other way: llama.cpp reached an answer in 16s, SGLang in 108s, FreeToken in 126s.

https://preview.redd.it/ztjuce0wfcoh1.png?width=1725&format=png&auto=…

Setup

  • GPU: NVIDIA RTX PRO 6000 Blackwell, 96GB VRAM
  • CPU: AMD Ryzen 9 9950X
  • System RAM: 96GB DDR5
  • OS: Ubuntu, CUDA 13, Docker
  • Model: Qwen3.8-Flash-Next
  • llama.cpp: UD-IQ4\_XS GGUF; MTP tested on the qwen4exp/mtp fork
  • SGLang and FreeToken: the same NVFP4 checkpoint revision
  • Client: AIPerf, with thinking off and the same non-thinking sampler

I ran one engine at a time, with fresh starts and GPU cooldowns. The runs saved resolved configurations, outputs, memory use and GPU telemetry.

Important note about the comparison

These are results for the tested stacks on this workstation. Quantization, KV-cache format, memory placement and speculative decoding differ.

SGLang uses its NEXTN draft head. FreeToken has no speculative decoding in the tested setup, I couldn't get it to work. llama.cpp has separate baseline and MTP results.

So the headline does not isolate the engine software alone. The repository includes the configurations so you can see what produced each number.

Newer builds, PRs and speculative decoding tested

  • llama.cpp MTP: the danielhanchen/llama.cpp qwen4exp/mtp fork, pinned to d1a92352, with the roughly 2.6GB draft head. On matched coding tests, decode improved 1.63x at 8K and 1.69x at 32K. Those gains compare the same build with the head off and on.
  • SGLang on Blackwell: the tested image included PRs \#36567, \#36556, \#36749 and \#36750, plus a local FP8 KV-cache patch. These were part of the working configuration, not individually benchmarked speedups.
  • N-gram speculation: ngram-mod gave +6.8% decode on the tested code workload, but generated zero drafts on the tested prose with the 24-token match setting.
  • Experimental PLE reads: I built llama.cpp PR #28136, but withdrew the read-mode comparison after discovering that a renamed flag was ignored. The intended direct-read mode was never exercised, so I am not claiming a speedup from that PR.

The report records the pinned builds and withdrawn findings alongside the successful tests.

1. All four configurations fit the full window. The waiting time is very different.

This test uses roughly 261,500 input tokens and a 128-token answer inside the 262,144-token window. The accepted input counts differ by two tokens across configurations.

|Configuration|First token|Decode|
|:-|:-|:-|
|SGLang|35.4s|126.9 tok/s|
|FreeToken|80.4s|87.5 tok/s|
|llama.cpp + MTP|210.2s|52.6 tok/s|
|llama.cpp baseline|258.4s|20.3 tok/s|

Going from over four minutes to about 35 seconds changes how usable a large prompt feels.

The two columns measure different things: first-token time is the initial wait; decode is how quickly the answer arrives after that.

2. A short-prompt test misses the long-context behavior.

The separate prose sweep uses 2,048-token answers and three measured requests per input length.

https://preview.redd.it/x7dnlip1gcoh1.png?width=1575&format=png&auto=…

|Configuration|Decode at 2K input|Decode at 259,584 input|
|:-|:-|:-|
|SGLang|182.7 tok/s|191.5 tok/s|
|FreeToken|100.1 tok/s|94.8 tok/s|
|llama.cpp + MTP|126.8 tok/s|61.4 tok/s|
|llama.cpp baseline|101.9 tok/s|20.2 tok/s|

Prefill also changes the ranking. FreeToken starts behind llama.cpp at 2K input: 1,525 vs 1,869 tok/s. At 128K it reaches 3,231 vs 1,362 tok/s, about 2.4x faster.

3. MTP helps llama.cpp, but it does not remove the long-prompt wait.

On real coding prompts, comparing the same fork build with the draft head off and on:

https://preview.redd.it/cbj78fx5gcoh1.png?width=1425&format=png&auto=…

|Input|MTP off|MTP on|Decode gain|
|:-|:-|:-|:-|
|8,192 tokens|94.9 tok/s|155.1 tok/s|1.63x|
|32,000 tokens|83.1 tok/s|140.4 tok/s|1.69x|

The draft head is roughly 2.6GB.

At the full window, the tested MTP configuration reached 52.6 tok/s, versus 20.3 tok/s for the baseline configuration. That is a 2.59x gap, but the full-window comparison also involves a different build. The matched-build coding tests above isolate the draft-head change more cleanly.

I would not attribute the 258s → 210s first-token improvement to MTP alone.

4. I checked accuracy as well as speed.

https://preview.redd.it/8bb55qjegcoh1.png?width=1725&format=png&auto=…

|Stack|GSM8K|MATH-500|
|:-|:-|:-|
|llama.cpp baseline|95.60%|92.60%|
|SGLang|95.22%|93.00%|
|FreeToken|95.75%|92.20%|

The llama.cpp MTP arm scored 95.75% on GSM8K.

The tests used 1,319 GSM8K problems and 500 MATH-500 problems. The paired comparisons did not detect statistically significant differences.

That does not prove the stacks have identical quality. These are two short math benchmarks, with no full-precision reference on this machine.

5. Starting the model is a separate benchmark.

https://preview.redd.it/hw5nqk4dgcoh1.png?width=1425&format=png&auto=…

Median time from starting the container to receiving the first answer:

  • llama.cpp: 16s
  • SGLang: 108s
  • FreeToken: 126s

FreeToken returned HTTP 200 from /health after about 3.3s, but took about 82s to reach serving readiness, followed by roughly 44s for its first generation.

That first request includes compilation work. Measuring only the health endpoint would give a very misleading startup result.

6. Loading modes barely changed speed with the experts on the GPU.

I compared none, mmap, mlock, mmap+mlock and dio on the same llama.cpp image, with the same tensor placement and real coding prompts.

  • At 8K input, prefill ranged from 2,036 to 2,124 tok/s — a 4.3% spread.
  • At 32K, it ranged from 1,946 to 1,956 tok/s — about 0.5%.
  • No arm ran out of memory or restarted.

My earlier 1.87x RAM-resident loading gain used a different placement, with 23 expert layers computed on the CPU. In this test, all experts stayed on the GPU.

Loading mode can matter when the CPU computes the experts. It made little difference in this configuration.

https://preview.redd.it/sas55spwhcoh1.png?width=1575&format=png&auto=…

7. MTP became slower when experts were offloaded to the CPU.

https://preview.redd.it/ib6wnkk7icoh1.png?width=1650&format=png&auto=…

The MTP gains above do not apply to every memory budget.

I repeated the test with smaller usable VRAM pools on the same RTX PRO 6000, using 2,048-token coding prompts and 256-token answers. Both arms used the same fork build.

|Usable VRAM|Expert layers on CPU|MTP off|MTP on, head on GPU|
|:-|:-|:-|:-|
|16 GiB|45|33.0 tok/s|9.5 tok/s|
|24 GiB|42|34.7 tok/s|10.2 tok/s|
|32 GiB|36|38.1 tok/s|11.9 tok/s|
|48 GiB|23|48.3 tok/s|18.2 tok/s|
|96 GiB|0|99.8 tok/s|160.2 tok/s|

At the full 96 GiB budget, MTP gave 1.61x faster decode. At 24 GiB, it made decode about 3.4x slower.

Moving the draft head to the CPU did not fix the 24 GiB result: 9.6 tok/s, versus 34.7 tok/s with MTP off.

In these tests, MTP helped only when all experts stayed on the GPU. Verifying drafted tokens adds work, and CPU expert execution can outweigh the benefit.

These are VRAM-capacity limits on one Blackwell card, not measurements of actual smaller GPUs. Their bandwidth and compute performance will differ. The lookup table used the build’s default lazy-read mode in both arms.

8. Finishing sooner also reduced estimated GPU energy per request.

For the full-window test with a 128-token answer, I estimated GPU energy as median power while busy × request duration.

https://preview.redd.it/mpefmbggicoh1.png?width=1425&format=png&auto=…

|Configuration|Median GPU power|Approximate GPU energy|
|:-|:-|:-|
|SGLang|358 W|13 kJ|
|FreeToken|404 W|33 kJ|
|llama.cpp + MTP|489 W|104 kJ|
|llama.cpp baseline|440 W|116 kJ|

The fastest configuration used roughly one-ninth the estimated GPU energy of the baseline for this request.

The main difference was how long the GPU remained busy: about 36 seconds for SGLang versus 265 seconds for llama.cpp, including generation.

This is an estimate using median power, not integrated energy or a wall-socket measurement. It excludes CPU, RAM and SSD power. No thermal throttling was recorded in these runs.For the full-window test with a 128-token answer, I estimated GPU energy as median power while busy × request duration.Configuration Median GPU power

The fastest configuration used roughly one-ninth the estimated GPU energy of the baseline for this request.The main difference was how long the GPU remained busy: about 36 seconds for SGLang versus 265 seconds for llama.cpp, including generation.This is an estimate using median power, not integrated energy or a wall-socket measurement. It excludes CPU, RAM and SSD power. No thermal throttling was recorded in these runs.

Resources

I made a full video covering the memory placement, engine setup, flags and these results:

Full video: https://youtu.be/RlsxXB5q-cA**

GitHub — report, scripts, configurations, raw results and charts

The new report is engine\_benchmark\_report.html.

PS: AI was abused while making edits.

Has anybody tested the same model across these engines on a different GPU or memory setup?
I am especially interested in whether FreeToken stays this flat at long context, and how much MTP helps when some experts are offloaded to the CPU maybe on the other models.
And maybe you found more efficient methods to run it too,

▲
55
-1
19👁
r/LocalLLaMA · u/bradnickel · 33d ago
How to squeeze out every last drop of your precious RAM on your Mac - Use iPhone mirroring post image

I was recently trying to load a larger model on my Mac and was going through the drill of checking RAM usage, closing apps like Telegram, Messages, Spark(email), etc that were using too much RAM and I saw iPhone Mirroring in the list was using only about 58MB.

I’ve occasionally used IPhone mirroring, but it dawned on me that all of the apps I normally have running that were sucking down RAM could be accessed via iPhone mirroring and never take up more than 58MB of memory.

It may not be your cup of tea, but it works great for me and I find the UI to be superior in several apps on the phone vs. desktop. I’m writing this post in iPhone mirroring in the Reddit app.

Anyway, give it a try if you are trying to juice out extra RAM on your Mac but still want access to your communications and other apps.