58 posts · 1 sub · RSS
← prev past day next →
day hourdayweekmonthyearall
allr/LocalLLaMA
▲
619
+528
7👁
r/LocalLLaMA · u/dasbin · 18h ago
Strata rewrote their Github history to wipe evidence of Claude-authoring

Just noticed this today when I went to run the built-in "UPDATE" script and git failed because there was no common ancestor.

Looked into why, and apparently every historical commit has been re-written to strip the "Co-Authored by Claude" text from the descriptions.

Personally I think that's pretty gross. I'm struggling to think of any reason to do this other than an intention to be dishonest about the origins of the project.

💬 383 (+324) open on reddit ↗
▲
615
+429
3👁
r/LocalLLaMA · u/Mr_BETADINE · 22h ago
chatgpt's new intelligent ui was reverse engineered in less than 24 hours, and apparently you can recreate it with local llms post image

came across a pretty interesting technical breakdown of chatgpt's newly launched "intelligent ui" feature, and thought this subreddit might find it interesting.

for anyone unfamiliar with the concept, intelligent ui is essentially openai's take on generative ui. instead of restricting llm responses to plain text or markdown, the model can compose actual interactive interfaces in real time.

there are different approaches to making this work. some systems let the model choose and compose elements from a predefined component library, while others allow it to generate entire interfaces on the fly (basically writing html/react code and rendering it inside an iframe).

it's more of a spectrum than a single technique. projects like openui, vercel's json-render, google's a2ui, and now chatgpt's intelligent ui all sit somewhere along this spectrum, with different trade-offs in flexibility, reliability, performance, and how much freedom the model gets.

but that's not even the most interesting part.

These folks managed to reverse engineer chatgpt's implementation in less than 24 hours after launch!

what's particularly impressive is that they claim to have done this entirely through publicly observable behavior, without access to openai's internal codebase.

from their write-up:

“All observations come from our own ChatGPT accounts, from the traffic the ChatGPT web app generates, and from the JavaScript that chatgpt.com serves publicly.”

found this pretty fascinating from an engineering perspective, especially considering how quickly they managed to put together a breakdown of how the system works.

and then there's the funnier part.

the same team released something called open intelligent ui, which is a pretty obvious jab at how openai isn't really "open" anymore. the joke works even better when you realize these guys actually own the domain openui.com lol.

the idea they're pitching is that you can recreate experiences similar to chatgpt's new intelligent ui inside your own applications using their open source framework.

and here's where it gets particularly interesting, you can technically do all of this with local llms.

since openui is model agnostic, you can integrate it with local models through ollama, lm studio etc. it's not necessarily a one click, out of the box recreation of chatgpt's experience, but from what i understand, the underlying pieces are there to build something similar that runs entirely locally.

i initially came across these folks through a viral twitter post comparing chatgpt's intelligent ui with openui's generative ui, and ended up going down a rabbit hole reading about the different approaches to generative ui.

some helpful links for anyone interested:

would love to know what everyone here thinks about generative ui in general.

is this actually a useful direction for llm interfaces or is it another one of those things that looks amazing in demos but doesn't translate particularly well to real world applications?

i'm especially curious about the local inference angle. with smaller models getting increasingly capable, do you see a future where something like this becomes practical entirely on device? or is the additional complexity, latency and structured output overhead simply not worth it compared to a conventional ui?

local llama has been my go to subreddit for years whenever i come across something interesting in the llm space, so genuinely curious what the general opinion here is.

would love to hear your thoughts, especially if you've tried building something similar with local models!

💬 70 (+39) open on reddit ↗
▲
294
+237
9👁
r/LocalLLaMA · u/Fun-Meaning-6474 · 24h ago
Running decision model locally on an RTX 4090 to find out which one is the fastest post image

recently saw a bunch of open decision models pop out of nowhere in the last two weeks (laya, liquid's d1, cloudflare's clef-flash, interfaze's lev), so I wanted to see how far apart they actually are on the same GPU(yes, model size is a huge factor, but still isn't the only factor). all four had the same task of reading nine wikipedia articles about centipedes (9,534 words) word by word and flag every word that names a centipede. one /v1/systemone call per word, the next word goes out the second the answer comes back

request for every word:

{"state": "Word: \"Scolopendra\".", "questions": {"centipede": {"type": "noul", "instructions": "Does this word name a kind of centipede?"}}}

|model|weights|engine|per word (p50)|words in 32s|accuracy|centipede names caught|wrong picks|
|:-|:-|:-|:-|:-|:-|:-|:-|
|Laya|Laya-BF16.gguf|llama.cpp b11495|3.9 ms|7,980|97.4%|70%|98|
|d1 3B|d1-3B-AD-Q4\_K\_M.gguf|llama.cpp b11495|6.0 ms|5,306|96.5%|51%|51|
|Clef-Flash 9B|Clef-Flash-Q8\_0.gguf|llama.cpp b11495|24.4 ms|1,292|97.2%|36%|2|
|Lev 4B|interfaze-ai/lev, bf16|lev serve (PyTorch)|51.0 ms|626|98.9%|83%|4|

laya and d1 gap the other models in speed, though not so much on accuracy (yes, it does say 95%, but even saying "no" counts as a correct answer, so that's where the high acc comes from). what everyone might care about more is how well each one did their respective task and lev catches the most while being 13x slower than laya, partly because it runs in its own pytorch server instead of llama.cpp (it measured 68 ms on a different 4090, so it's CPU-sensitive too). but in the end Laya is the fastest model overall, and considering how easily it can be fine-tuned for any use case I'd say that be my go to pick

setup:

  • GPU: rented RTX 4090 (driver 580.119.02, 32 vCPU)
  • engine: llama.cpp b11495 (commit 37ac63456, CUDA 12.8 release build), -ngl 99, everything else default
  • Laya, Clef-Flash: the ggml-org GGUFs
  • d1: our own AD-Q4\_K\_M quant (atomic.chat), runs natively on /v1/systemone since the lfm2-d1 support landed in #30110
  • Lev: interfaze's LoRA on Qwen3.5-4B in its own lev serve, default settings (--compile never finished warming up)
  • latency: end to end from a Python client on the same box over localhost
💬 51 (+34) open on reddit ↗
▲
344
+209
15👁
r/LocalLLaMA · u/_TheWolfOfWalmart_ · 23h ago
$2800 rig with 8x Radeon Pro V620 (256 GB VRAM) + custom vLLM fork = Qwen3.8-Flash-Next at 60 to 100 t/s decode and 3000+ t/s prefill post image

Post title is slightly misleading, I don't think you can get these for $350 each anymore but they're still pretty cheap all things considered. They're Radeon Pro V620's which are older RDNA2 enterprise cloud gaming cards with 32 GB VRAM.

(Ignore the RTX 4090 on the side, it's just used for stuff like image/video gen models, no LLMs)

But I bought these cards a couple months ago as a gamble to see if I could build a big VRAM rig with usable speed for relative peanuts.

I was struggling with llama.cpp for a long time, but the prefill was pretty bad (around 350-450 t/s average with this same model) and vLLM just didn't work on the cards. Plus llama.cpp just sucks at concurrency.

I'd been planning to sell the cards lately because this wasn't going to work for my use case, but then decided to see if I (Claude) could make a vLLM fork that both works with the cards and actually gets good speeds out of them. I had it build/test/iterate on custom RDNA2 kernels.

Problem solved! It worked out way better than I expected. I thought maybe I'd hit 1000 t/s prefill with QFN at best, but this is something like 800% faster than llama.cpp was managing.

Couldn't be happier with the results! GPU sale plan canceled lol.

I'm going to have it continue optimizing and see how it goes, and make sure DeepSeek and GLM-5.3-Flash work as well.

llama-benchy results below with concurrency = 1 and vLLM running with PP=4 (no tensor parallel here) with orcarouter's uncensored QFN which I quantized. Routed experts are W4A16 and everything else remains at BF16. MTP enabled with 3 token drafting.

It gets 40 to 50 t/s decode with MTP disabled.

https://preview.redd.it/4223w82yz9uh1.png?width=666&format=png&auto=w…

💬 153 (+66) open on reddit ↗
▲
133
+131
7👁
r/LocalLLaMA · u/TheOriginalG2 · 15h ago
Qwen3.8-27B: 159 tok/s on R9700, 64 tok/s on Strix Halo

LemonSeed Studio is an iPad editor/IDE with on-device inference on an AMD GPU in a Thunderbolt enclosure. It embeds the unmodified upstream Linux amdgpu + amdkfd driver (mac\_linuxgpu) as a PCIDriverKit extension, and runs LemonSeed Engine (LSE) on the GPU. In the photos: iPad Pro + Sapphire Radeon AI PRO R9700, Qwen3.8-27B Q4 with a Q8 DFlash2 draft, 131k context.

LSE is the same engine on every platform: Linux, macOS and iPadOS. It records the model's forward pass as a graph, fuses ops and generates kernels for the GPU it's on. Where a kernel has several layouts, it measures each on the device and keeps the fastest. Speculative decoding (MTP and DFlash2) picks how many draft tokens to verify from measured acceptance and cost.

Decode, code prompt, Qwen3.8-27B Q4 (baseline / MTP=3 / DFlash2 tok/s):
\- R9700 on iPadOS (LemonSeed Studio): 31.2 / 108.7 / 158.5
\- R9700 on macOS: 32.2 / 111.9 / 159.4
\- R9700 on Linux: 32.0 / 111.0 / 144.6
\- Strix Halo (Radeon 8060S) on Linux: 14.0 / 47.8 / 64.4

Prefill at 4K tokens: 1,636 (iPadOS), 1,634 (macOS), 1,418 (Linux R9700), 517 (Strix Halo). It holds up at 32K: 1,395 / 1,401 / 1,226 / 462.

New in 0.5.8:
\- DFlash2 draft trees on by default: one target pass verifies a tree of draft candidates when that's measured to be faster than a chain
\- Strix Halo (gfx1151) prefill and decode work: fused gate/up GEMM, two query tiles per workgroup in prefill attention
\- Fixed a startup GPU fault
\- Linux archive bundles its own HSA runtime, so you only need the amdgpu driver with /dev/kfd access
\- lse-server models / pull: grab Hugging Face models with their MTP or DFlash2 companions

OpenAI-compatible server, CLI, and a C library (libLSE).

Engine: https://github.com/Geramy/LSE
Studio: https://github.com/Geramy/LemonSeed-Studio

Happy to answer questions, especially about getting an eGPU working on iPadOS.

💬 63 (+61) open on reddit ↗
▲
99
+78
4👁
r/LocalLLaMA · u/pmttyji · 10h ago
[Paper] EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.
💬 6 (+6) open on reddit ↗
▲
152
+63
3👁
▲
51
+49
8👁
r/LocalLLaMA · u/Mr_Moonsilver · 19h ago
Bois, there's now a waterblock for the R9700. Quiet 4x or 6x builds are now possible.

Seems 1-slot design, so you could cram in quite a lot into a case

💬 24 (+24) open on reddit ↗
▲
50
+46
6👁
▲
63
+42
7👁
r/LocalLLaMA · u/Mrinohk · 16h ago
Qwen 3.6 35B appreciation post

Referring to specifically Unsloth's UD\_Q4\_K\_XL quant because that's going to be a question, and is relevant regardless.

It's old now. It's not great at coding medium sized or even small-ish projects. I wouldn't hand my codebase to it by any means. It hallucinates, like any other model. It's not perfect, by any means.

I don't like the fine tunes; they're almost all coding focused, and lose general assistant capability to a strange degree.

What it is good at is general, broad agentic action.

Set the lights in the living room, and bash into this machine to get a movie going.

Send my grocery list to my phone/watch when I get to walmart.

Remind me to clock in at work each day because it's becoming a problem.

Tell my husband to come here because I'm under the car and I can't drop this thing that I finally got in just the right spot but need a third hand to get this wrench in the correct spot.

Tell my dad about the pets or any one of my projects because I'm showing off your memory to him.

This is all shit that it can do, consistently. Sometimes it'll thrash a bit, but it stays on task and is fast enough that little mistakes are a non-issue.

I recently got a couple of Tesla P100s to power the model, and it's made this model, that I already had going at a pretty good clip on limited hardware, to run at speeds that are genuinely conversational. \~120-140 t/s generation, \~1000PP at 0ctx, \~700 by 13k.

It's more than smart enough to know how to do these tasks and, with the right sampler settings, thinks ridiculously efficiently for them. Actually awesome.

Fucker got a minecraft mod pack installed and running on a machine it wasn't even running on through the prism launcher appimage. Don't worry, it only has ssh towards computers on my tailnet. It's probably fine.

It's bad at holding a persona. My TTS model is trained on Paul Bettany's MCU Jarvis. When the model emits the right phrasing, it feels like magic. It doesn't do that very often. Gemma4 is great at that.

I am actually so scared that if they do release a Qwen4 model in this class (30-40B parmeters, 2-5b active) that it's going to be a coding focused, overthinking, genuinely capable but not at all fast, mess. Gemma4 26b feels like it would be so close if it wasn't dumb as rocks. As it stands, I pay anthropic for that shit coding shit. Maybe when I can run Flash next at good speed (current \~30-40 t/s rn with strata, but at q3. It's ok.) I can kick claude to the curb.

I want a model like what 3.6 35b is, but smarter. Just as fast, but thinks of the little things. Memory updates. I ask it to add something to my calendar, but maybe it also sets up a dedicated notification for the specific time separately. That'd be nice. Not more coding focused. There are so many coding assistant bots. They're great. But the general AI assistant isn't solved yet for the normal person. Maybe it could have better general knowledge, but I think an n-gram table would solve that. It said the P100s used ROCm the other night. Silly bot.

💬 31 (+21) open on reddit ↗
▲
47
+41
3👁
r/LocalLLaMA · u/pmttyji · 9h ago
[Paper] DLoop: Looped Speculative Decoding
Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at this https URL.
The code is currently under internal review and will be released soon. Stay tuned!
💬 9 (+9) open on reddit ↗
▲
38
+28
5👁
▲
23
+11
8👁
r/LocalLLaMA · u/Express_Quail_1493 · 21h ago
Currently having high success with this little niche finetune i found sitting in the corner of huggingface

Currently having high success with this little niche finetune i found sitting in the corner of huggingface

If you want to try it out here is a smaller quantisation iq3\_s works really well in my codebases.

Original Model:
https://huggingface.co/tsinghua-sigs-robot-lab/VeriLoop-E2-GGUF


Smaller Quant:

https://huggingface.co/tahaalam2009/VeriLoop-E2-GSQ-RCO-GGUF

💬 16 (+7) open on reddit ↗
▲
11
+9
4👁
r/LocalLLaMA · u/Distinct-Pie2389 · 11h ago
Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s

HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format.

Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does:

  • 262K context (the model's full native window)
  • \~130 tok/s decode with MTP3, 70.8% acceptance
  • 3,591 tok/s prefill on a 9K prompt
  • Perplexity within 1.3% of the official artifact, so the conversion is clean
  • Vision and tool calls still work

• E8 4Bit quant

Converter + writeup here: https://github.com/T-Crypt/ninfer-4090/pull/4

Questions welcome.

Repo: ninfer-uncensored

💬 23 (+23) open on reddit ↗
▲
10
+7
6👁
r/LocalLLaMA · u/Felladrin · 16h ago
Decisions in a js13kGame post image

I'm having AI models play Cat Goric, a 2D platformer I made in 2021 during js13kGames: https://js13kgames.com/games/cat-goric-escape-from-the-warp-chamber

The best result so far is 8 of the 14 playable levels cleared in one continuous run. Two models have cleared eight levels so far:
\- Qwen/Qwen3.8-Flash-Next
\- Cloudflare/clef

If you'd like me to try a specific model, please name it in the comments.

You can also try beating the game with a model of your choice. The project with instructions for running the challenge with different models/engines is here:
https://github.com/felladrin/ai-plays-cat-goric (PRs welcome!)

And attached is a 2-minute recording of Clef clearing eight levels (the game pauses while the model thinks, so the video is time-warped).

💬 2 (+2) open on reddit ↗
▲
24
+7
2👁
r/LocalLLaMA · u/pmttyji · 9h ago
[Paper] Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that conditions on both the context and target efficiency specifications, enabling fine-grained control over the accuracy-efficiency trade-off at inference time. The model learns to activate task-relevant parameters within elastically-nested sub-networks, allowing a single model to span multiple capacity points while maintaining input-adaptive routing. Through experiments we demonstrate that we can create a model that allows the flexibility to use 1,2,3,4 billion parameters while being more accurate than their dense counter-parts (2-5\\% on knowledge-intensive benchmarks) and at par with their static versions while delivering similar latency metrics as dense models. Overall, we save on device disk space by sharing the model parameters, allow flexibility of serving based on DRAM and compute available while delivering more accurate results.
▲
9
+6
7👁
r/LocalLLaMA · u/Brilliant-Hall1387 · 19h ago
Staging quantized weights to FP8 instead of fp16: 2× M6 matrix path, +40% MLX prefill (+ int8 on M5) post image

Regular MLX QMM (quantized matmul) stages operands (dequant) to FP16 before matrix multiplication. But if you modify MLX to stage to FP8 instead, you can use the 2x faster FP8 matrix path on the new Apple Silicon M6 hardware!

The same idea works on the M5 family: stage 4-bit affine weights to int8 before the matmul and you get a similar prefill boost (M5 and M6).

Benefits apply to prefill (+40% prefill Qwen3-8B or +50% Qwen3.8-27B). Decode is bandwidth limited so not much difference on decode side and better to let it use default FP16 staging on decode side.

It's an experimental fork, not upstream (mlx#4627) and quality was measured: worst-case perplexity increase under \~1%. FP8 performance improvement needs an M6 on macOS 27 with a deployment-target-27 build; int8 works on M5 and later.

This was quite an interesting research project and it helped me get a deeper understanding of the math behind LLMs on the hardware side. 😄

All details, code and evidence available on my blog: https://precisit.com/en/blog/apple-matrix-formats/

The MLX fork itself with FP8 + Int8 QMM patch: https://github.com/precisit/mlx/tree/staged-8bit-qmm

Disclosure: the blog is from Precisit, where I work.

💬 6 (+3) open on reddit ↗
▲
12
+6
4👁
r/LocalLLaMA · u/woct0rdho · 13h ago
LoRA over GGUF: Train Qwen3.8-Flash-Next in 40G VRAM

https://github.com/woct0rdho/transformers5-qwen3.5-recipe

An update to my LoRA over GGUF series: Now we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM, with no CPU offloading, with engram on disk that does not reduce training speed.

On Strix Halo it trains context chunk size 2048 at 9.5 s/it. That's 200 token/s. There is still room to optimize, compared to > 1600 token/s PP we've achieved, and the common sense that LoRA training takes 2-3x work of PP.

Since transformers 5.18, initial support for modern GGUF has been merged, and we can expect more work in this direction.

Spoiler: In the torch-ggml-ops repo there is something called GGTensile. Basically it's Tensile-like asm-level optimization on MMQ kernels. We already see it's faster than HIP in many cases. I'll make a new post when I have something to show on this.

I guess I'll skip DeepSeek-V4.1, unless someone can quantize or prune it to < 125 GiB.

💬 12 (+3) open on reddit ↗
▲
8
+6
3👁
r/LocalLLaMA · u/jjusko20 · 10h ago
Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M)

Hey guys - primarily a research release with working models,

Not so much a model as a psuedo-new quantization technique. I've been experimenting with a modification of ISTALab's RCO algorithm that can quantize models on a strict VRAM budget.

It's at its core an approximation algorithm that attempts to make up the difference with a few different strategies. I'm quite happy with how well it's working at low bits. I've seen significant KLD improvements, particularly in the 1 bit and 2 bit range. This should be model agnostic and it doesn't require loading the FP16 teacher into VRAM. I've included a little more detail on the model card and will probably publish the full recipes. I plan to apply this method to the larger models in the family next.

Unfortunately Unsloth does not have their KL divergence table available, but I've created a table for you to see the difference between these quants and standard llama.cpp imatrix quants. For accuracies sake, both my quants and the llama.cpp quants are calibrated on the sane wikitext imatrix dataset, and tested on two held out sets.

The KLD table

As you can see, particularly at extremely low BPW, these quants vastly outperform their standard imatrix KLD - particularly drastically cutting IQ1\_M divergence at nearly the same size budget.

These are more proof of concept quants, as 2B is a fairly small model and suffers from such low BPW, but here's a fun example of how these are still relatively functional at extreme compression

4\/5 right on an IQ1\_M version of Qwen 3.5 2B \(green, not purple\) - I think it's kinda crazy it can answer anything

and here's the standardized imatrix IQ1\_M quant (not using dynamic RCOL)

yikes

and the file sizes

GGUFs:

https://huggingface.co/trubisky/Qwen3.5-2B-RCOL

▲
24
+6
2👁
r/LocalLLaMA · u/falconandeagle · 7h ago
Stepping away from Benchmarks and Code, what models are you using for Creating writing projects

All anyone ever talks in this sub is about x new model having x benchmark scores and how well it can code. Or some new tech to increase t/s for qwen.

So I wanted to go back to the roots of this sub and talk about how well local models can write.

I am not going to mention closed source models, this discussion is for open weight models only.

I don't have the best machine, so I am limited in what I can use locally.

So far I have mostly been running Gemma 31b finetunes. I have a finetuned system prompt for creative writing that I have refined over time. This is mostly for long form story writing and not RP. I use it to write Sci Fi and Fantasy novels and short stories and sometimes fanfiction.

I also have created a creative writing harness using Pi. It's mostly for self correcting and getting rid of AI slopisms.

Gemma 31B Mero-mero-v2 has so far been my favourite. It writes well, its does not feel lobotomized and its mostly uncensored. It's instruction following is okay, sometimes it makes mistakes and gets the character traits mixed up but my creative writing harness takes care of this.

Qwen 3.8 27b finetunes, Disappointing, I guess the base model is really fucking bad for creative writing purposes and even with finetuning you can only do so much. It does follow instructions quite well but its writing is just terrible.

Muse Glimmer 30b, this one is an interesting one, sometimes it will give really good prose but sometimes it will respond like a 2b model, the variance seems to be super high for some reason. Maybe something to do with chat template, I am not sure. Anyway I think even when bad it's still better than Qwen.

Gemma 4 26b4a, pretty good but gets beaten by its bigger bother Gemma 4 31b. I do use it when I want faster responses and to check for issues with my initial scene beat prompts.

Gemma 31B Artemis, drummer's finetune, eh, I was a bit disappointed with this. Seems like its more focused on RP rather than long form story writing. Also got refusals which I did not get with the mero finetune.

So that has been my experience writing with some local models. What are you guys using for local creative writing projects? Have you guys tried a creative writing harness?

I just want to mention quickly my writing process before I end this post, I write using an outline and scene beat prompts that are hand written. So far AI has been extremely disappointing in being truly creative and still requires a ton of hand holding to not write the most cliché ridden drivel.

💬 29 (+8) open on reddit ↗
▲
12
+4
6👁
r/LocalLLaMA · u/demonicpigg · 23h ago
Open sourcing Game Summoner, my prompt to game site

tldr: Open sourced my prompt to game suite: https://github.com/ndamiano/ai-agent-test, it's MIT licensed, and this runs well on my 5090 with 64gb ram, but any model that can handle tool calls can manage.

Edit: I can't believe I forgot to share a game... This is one shot with qwen 3.8 flash next! https://gamesummoner.com/g/siWy5VIMxF-H

Hey everyone, I recently launched https://gamesummoner.com. You may have also seen that google just released https://playground.google. I cannot compete with that, and honestly, I wanted to figure out how I could give back to the people here who, probably unknowingly, helped me get from idea to implementation.

This isn't a nice clean repo for you to trivially run with something like python run.py, as it is tailored to my specific setup on digital ocean, and using runpod and aws as my hosted GPUs. That said, there is a script called scripts/local_gpu.py that will get you most of the way to running this locally. It manages the worker boxes, spinning up instances of inference (I use ninfer with Qwen 3.8 27B locally and a modified sglang https://github.com/ndamiano/sglang-rtxpro6000 with Qwen 3.8 Flash-Next on a rented rtx 6000 in prod, and comfyui with a buncha models), but this can be modified by your agent to launch however you need.

The architecture is pretty straightforward. Anytime a request comes in, it throws it into a queue (sql, I am cheap, and it works just as well at this scale as something like kafka), a worker long polls for work, picks it up, and returns the value.

This requires there to be workers, and so I built a simple autoscaler, it walks up the cost ladder from aws / runpod to try to get the cheapest GPU available (I probably should add more sources, but eh, that's work on the least interesting part.)

And for how the actual generation goes, I've done a ton of iterations (you can see many of them in https://github.com/ndamiano/maestro-labs, as I said.. several iterations on the name), and settled on creating a design doc with a team of agents. The first agent creates a high level design, second and third in parallel are visual and engineering, fourth is an integrator that puts them all together as the "holy grail" of the design.

Once we've got the design, in it goes to the same model, with a new prompt, that is, effectively, build the game described in the design. We give it access to tools that let it test the game, take screenshots, etc. and wait for it to call done. Once it's finished, we validate the build and give it a quick "play", where the model looks at a photo, tries some input, and sees what happens. We return any exceptions and inputs that do nothing (the model has notoriously been AWFUL at "is this good"...), and once there are none, we say "complete" and return to the user.

There are a couple other repos that are necessary:
```
https://github.com/ndamiano/gamesummoner-workers
https://github.com/ndamiano/gamesummoner-images
(I told you, the name went through some iterations...)
```

All said and done, I'm releasing this with an MIT license. This is a full, scalable, deployable website that generates games. I made sure all of the models used are well licensed, and so should probably not be an issue if you want to stand it up. There's quite a bit of setup, but like, you could get this up and running in a couple days with an agent. If you do and somehow make a few million, I'm currently unemployed, so I'd love a job lol.

▲
5
+3
5👁
r/LocalLLaMA · u/ailee43 · 16h ago
Best method to train a persona model

Hardware available:

2080ti 11gb

5070ti 16gb

Goal: I've downloaded my whole internet history (reddit posts, emails, chats, etc) and i'd like to train a model to think, act, and talk like me.

I'm currently running a qLORA pass with Qwen3.5-9B as a base, but when i run A/B tests, its unfortunately still pretty clear its not me.

Whats the best method to do this these days? Obviously would like to stay as local as possible since the corpus contains tons of personal info

Arch document to show the path i've already taken: https://markdownpastebin.com/?id=2493bd33b50b48ebbe0a49f437f573a6

💬 11 (+7) open on reddit ↗
▲
3
+2
9👁
r/LocalLLaMA · u/Reasonable-Height704 · 21h ago
blackwell gpus have PCIe5 issues?

Let's preface this with the fact I have 3090, 4090, 2080ti - and lots of stable long running compute heavy workloads.

I recently got a 5070ti (I am not willing to pay ridiculous money for 5090)

And it's been working reasonably well, until I left it running on a 6 hour CUDA job.

Near the end, it died, with system journal message:

NVRM: krcWatchdog\_IMPL: RC watchdog: GPU is probably locked! Notify Timeout Seconds: 7

So I tried to reproduce, but no success. My code is fine.

Then I investigate...

Apparently this is just problem with Blackwell we just accept?

https://en.gamegpu.com/news/zhelezo/rtx-5070-rtx-5080-i-rtx-5090-prodolzhayut…

I searched this sub and reddit, and previously people have mentioned it, but surprised there isn't more noise about it. Seems like Nvidia have only in the last few months officially acknowledged the problem.

https://www.reddit.com/r/LocalLLaMA/comments/1tifo1o/anyone_else_fighting_bla…

https://www.reddit.com/r/nvidia/comments/1wiaj9e/nvidia_acknowledged_the_blac…

💬 18 (+14) open on reddit ↗
▲
2
+1
3👁
r/LocalLLaMA · u/dreamyrhodes · 22h ago
Help with EXL2/3 pls

I am trying to run EXL2/3 models using Silly Tavern. Normally I am running GGUF but I wanted to see if EXL2/3 format could provide a better lore coherence than Q4 quants.

My rig is running a 4060 with 16GB.

As an API provider I tried TabbyAPI (know a better one for EXL2/3?).

I tried it with this template (and various adjustments, tempereture etc) https://huggingface.co/Nitral-AI/Violet\_Magcap-12B/blob/main/ST%20Presets/ChatML\_Master-Import.json
But it generates gibberish only. Sometimes it runs halfway ok but there will still be grammatical errors, half words and sometimes loops (doesn't get a stop token), often it's just entire word salad.

Now wtf am I doing wrong? How do I run EXL2 or 3 locally?

Screenshot is default bot's response to "Hello".

https://preview.redd.it/k6q7nbylcauh1.jpg?width=1159&format=pjpg&auto…

💬 10 (+1) open on reddit ↗
▲
4
+1
6👁
r/LocalLLaMA · u/pmttyji · 23h ago
Probably I'm doing something wrong using PR#29887 (Add a GPU cache for MoE experts kept in host memory)

I have 8GB VRAM(4060) + 32GB RAM(DDR5 5600). Tried this feature with b11491. Experimented with both cmoe & fit. Not getting expected t/s.

Please fix this for me.

And others, what are you getting for your limited VRAM? Share your t/s stats.

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -cmoe
3.25.566.603 I slot print_timing: id 3 | task 0 | prompt eval time = 7546.59 ms / 342 tokens ( 22.07 ms per token, 45.32 tokens per second)
3.25.566.621 I slot print_timing: id 3 | task 0 | eval time = 74878.16 ms / 1750 tokens ( 42.81 ms per token, 23.36 tokens per second)
3.25.566.624 I slot print_timing: id 3 | task 0 | total time = 82424.75 ms / 2092 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -cmoe --moe-cache-mib 1536
2.07.904.789 I slot print_timing: id 3 | task 0 | prompt eval time = 11915.27 ms / 342 tokens ( 34.84 ms per token, 28.70 tokens per second)
2.07.904.798 I slot print_timing: id 3 | task 0 | eval time = 86812.44 ms / 1305 tokens ( 66.57 ms per token, 15.02 tokens per second)
2.07.904.800 I slot print_timing: id 3 | task 0 | total time = 98727.70 ms / 1647 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -cmoe --moe-cache-mib 2048
2.45.980.222 I slot print_timing: id 3 | task 0 | prompt eval time = 15102.25 ms / 342 tokens ( 44.16 ms per token, 22.65 tokens per second)
2.45.980.408 I slot print_timing: id 3 | task 0 | eval time = 124490.76 ms / 1669 tokens ( 74.63 ms per token, 13.40 tokens per second)
2.45.980.412 I slot print_timing: id 3 | task 0 | total time = 139593.01 ms / 2011 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -cmoe --moe-cache-mib 4096
1.45.860.593 I slot print_timing: id 3 | task 0 | prompt eval time = 19690.18 ms / 342 tokens ( 57.57 ms per token, 17.37 tokens per second)
1.45.860.818 I slot print_timing: id 3 | task 0 | eval time = 58251.17 ms / 1143 tokens ( 51.01 ms per token, 19.60 tokens per second)
1.45.860.821 I slot print_timing: id 3 | task 0 | total time = 77941.35 ms / 1485 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -cmoe --moe-cache-mib 8192
6.12.049.848 I slot print_timing: id 3 | task 0 | prompt eval time = 27804.55 ms / 342 tokens ( 81.30 ms per token, 12.30 tokens per second)
6.12.049.864 I slot print_timing: id 3 | task 0 | eval time = 311652.61 ms / 3753 tokens ( 83.06 ms per token, 12.04 tokens per second)
6.12.049.866 I slot print_timing: id 3 | task 0 | total time = 339457.17 ms / 4095 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -c 131072 -cmoe --moe-cache-mib 2048
1.28.103.097 I slot print_timing: id 3 | task 0 | prompt eval time = 13046.26 ms / 341 tokens ( 38.26 ms per token, 26.14 tokens per second)
1.28.103.169 I slot print_timing: id 3 | task 0 | eval time = 52369.09 ms / 802 tokens ( 65.38 ms per token, 15.30 tokens per second)
1.28.103.171 I slot print_timing: id 3 | task 0 | total time = 65415.35 ms / 1143 tokens

Above ones with -cmoe while below ones without -cmoe & fit is on by default

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf --moe-cache-mib 2048
1.11.969.119 I slot print_timing: id 3 | task 0 | prompt eval time = 5196.27 ms / 341 tokens ( 15.24 ms per token, 65.62 tokens per second)
1.11.969.129 I slot print_timing: id 3 | task 0 | eval time = 38764.96 ms / 939 tokens ( 41.33 ms per token, 24.20 tokens per second)
1.11.969.131 I slot print_timing: id 3 | task 0 | total time = 43961.23 ms / 1280 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -b 2048 -ub 2048 --moe-cache-mib 2048
1.11.944.661 I slot print_timing: id 3 | task 0 | prompt eval time = 5403.22 ms / 341 tokens ( 15.85 ms per token, 63.11 tokens per second)
1.11.944.672 I slot print_timing: id 3 | task 0 | eval time = 39398.70 ms / 971 tokens ( 40.62 ms per token, 24.62 tokens per second)
1.11.944.674 I slot print_timing: id 3 | task 0 | total time = 44801.92 ms / 1312 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf --moe-cache-mib 4096
1.05.634.810 I slot print_timing: id 3 | task 0 | prompt eval time = 6284.73 ms / 341 tokens ( 18.43 ms per token, 54.26 tokens per second)
1.05.634.821 I slot print_timing: id 3 | task 0 | eval time = 32617.75 ms / 533 tokens ( 61.31 ms per token, 16.31 tokens per second)
1.05.634.822 I slot print_timing: id 3 | task 0 | total time = 38902.48 ms / 874 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -c 131072 --moe-cache-mib 2048
1.18.552.620 I slot print_timing: id 3 | task 0 | prompt eval time = 6213.00 ms / 341 tokens ( 18.22 ms per token, 54.88 tokens per second)
1.18.552.627 I slot print_timing: id 3 | task 0 | eval time = 47349.21 ms / 859 tokens ( 55.19 ms per token, 18.12 tokens per second)
1.18.552.629 I slot print_timing: id 3 | task 0 | total time = 53562.22 ms / 1200 tokens

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -c 262144 --moe-cache-mib 2048
2.09.088.373 I slot print_timing: id 3 | task 0 | prompt eval time = 14680.06 ms / 341 tokens ( 43.05 ms per token, 23.23 tokens per second)
2.09.088.387 I slot print_timing: id 3 | task 0 | eval time = 88642.28 ms / 997 tokens ( 89.00 ms per token, 11.24 tokens per second)
2.09.088.389 I slot print_timing: id 3 | task 0 | total time = 103322.34 ms / 1338 tokens

Below one is from past without this PR. 20 t/s for 128K context is not bad with 8GB VRAM + RAM.

llama-server -m E:\LLM\models\MOE\Qwen3.6-35B-A3B-IQ4_XS-00001-of-00002.gguf -fa 1 -ctk q8_0 -ctv q8_0 -kvu --cache-ram 24576 --cache-idle-slots -np 1 -cb -fit on -fitt 512 -t 8 --mlock --no-mmap --no-warmup -ctxcp 64 --no-mmproj -c 131072
4.39.110.891 I slot print_timing: id 0 | task 0 | prompt eval time = 1717.32 ms / 35 tokens ( 49.07 ms per token, 20.38 tokens per second)
4.39.110.903 I slot print_timing: id 0 | task 0 | eval time = 178110.45 ms / 3448 tokens ( 51.66 ms per token, 19.36 tokens per second)
4.39.110.905 I slot print_timing: id 0 | task 0 | total time = 179827.77 ms / 3483 tokens

Tried Q2 of Qwen3.8-Flash-Next just for fun.

llama-server -m E:\LLM\models\MOE\Qwen3.8-Flash-Next-GSQ-RCO-Q2_0-00001-of-00002.gguf -ctk q8_0 -ctv q8_0 --load-mode none
5.19.601.200 I slot print_timing: id 3 | task 0 | prompt eval time = 54472.12 ms / 379 tokens ( 143.73 ms per token, 6.96 tokens per second)
5.19.601.214 I slot print_timing: id 3 | task 0 | eval time = 186902.08 ms / 1500 tokens ( 124.68 ms per token, 8.02 tokens per second)
5.19.601.216 I slot print_timing: id 3 | task 0 | total time = 241374.19 ms / 1879 tokens

💬 16 (+3) open on reddit ↗
▲
4
+1
5👁
r/LocalLLaMA · u/tabletuser_blogspot · 20h ago
GPU - Vulkan llama.cpp benchmarks sorted by price to performance

This table to help anyone looking to build a budget Data Center homelab. I copied the bulk of value based, mid level, decent speed results GPUs and feed it to AI or SI and here are the recommended results. Data taken from Llama.cpp discussion thread: Performance of llama.cpp with Vulkan #10879 There are 76 different GPU models listed in the benchmark.

"Testing the 'Llama 2 7B model' and use Q4\_0 as it's simple to compute and small enough to fit on a 4GB GPU"

Based on the specific llama-bench baseline data provided, running local LLM inference via the Vulkan backends shifts the value hierarchy drastically. Modern mid-range consumer cards are severely bottlenecked by narrow bus widths (128-bit or 192-bit) during decoding (tg128), whereas enterprise components and older massive-bus flagships dominate performance-to-cost value. By analyzing the current 2026 secondary market pricing (collating active trends across secondary platforms like eBay and specialized tech hardware communities) against your baseline metrics, here is the performance-to-cost value ranking. The cost-to-performance efficiency formula balances the entry price against prefill speeds (pp512), decoding throughput (tg128), and total accessible VRAM.

Top 20 GPU Performance-to-Cost Ranking (Used Market)

|Rank|GPU Model|Est. Used Price|pp512 (t/s)|tg128 (t/s)|VRAM Capacity|Performance-to-Cost Architecture Profile|
|:-|:-|:-|:-|:-|:-|:-|
|1|Nvidia P102-100|\~$40 - $50|\~510|\~62.8|10 GB|Absolute Value King: Stripped mining card with a 320-bit bus. Yields \~1.3 tokens/sec per dollar spent on decode cycles.|
|2|AMD Instinct MI50|\~$110 - $130|\~1,119|\~108.5|16 GB|tg128 Efficiency King: Full 1,024 GB/s HBM2 bandwidth. Best cost-per-token decode engine on the secondhand market.|
|3|AMD Radeon VII|\~$140 - $160|\~1,059|\~101.1|16 GB|Same elite HBM2 memory substrate as the MI50 but packaged with consumer display outputs.|
|4|Nvidia GTX 1080 Ti|\~$110 - $130|\~585|\~67.7|11 GB|Legacy consumer warrior. Its wide 352-bit bus regularly out-decodes modern architecture under $300.|
|5|Nvidia Tesla P100|\~$90 - $110|\~678|\~63.1|16 GB|Budget HBM2 alternative. Slower core processing bounds its prefill, but decode values are incredibly high.|
|6|AMD Radeon RX 6800|\~$220 - $240|\~1,593|\~101.4|16 GB|Exceptional balance. Clean driver architecture yields massive decode velocity relative to modern hardware tiers.|
|7|AMD Radeon RX 7900 GRE|\~$400 - $430|\~2,336|\~116.1|16 GB|Modern value standout. RDNA3 architecture scales beautifully on compute tasks with excellent memory throughput.|
|8|Nvidia RTX 3060 (12GB)|\~$180 - $200|\~1,815|\~75.9|12 GB|The entry-level standard for consumer setups. Ample VRAM budget for small models at a highly accessible price tier.|
|9|Nvidia Tesla V100 (16GB)|\~$180 - $220|\~1,391|\~129.5|16 GB|Combined Enterprise Pick: Volta core structure provides blisteringly reliable generation and prefill baselines.|
|10|Nvidia RTX 2080 Ti|\~$200 - $230|\~1,888|\~97.5|11 GB|Highly efficient Turing flagship layout. Out-paces newer equivalents due to an aggressive 352-bit bus framework.|
|11|AMD Radeon RX 7800 XT|\~$350 - $380|\~2,017|\~118.2|16 GB|Clean, highly competitive RDNA3 compute engine displaying great out-of-the-box Vulkan metrics.|
|12|AMD Radeon RX 7900 XT|\~$500 - $550|\~2,941|\~123.1|20 GB|Massive 20GB framework buffer size. Excellent performance scale, though commands a higher price footprint.|
|13|Nvidia Tesla P40|\~$120 - $140|\~488|\~59.3|24 GB|The cheapest entry to 24GB allocation. Let down by poor FP16 computing speeds, keeping context loading sluggish.|
|14|Nvidia RTX 4070 Super|\~$480 - $520|\~4,608|\~108.7|12 GB|Blistering prefill speed bounds. Highly performant cores make up for the standard 192-bit bus structure.|
|15|Intel Arc A750|\~$90 - $110|\~1,075|\~42.6|8 GB|Phenomenal raw bandwidth per dollar, but tightly restricted by a fixed 8GB VRAM ceiling.|
|16|Nvidia RTX 4070 Ti Super|\~$680 - $730|\~6,099|\~129.4|16 GB|Outstanding raw throughput benchmarks, but hits a higher tier of up-front investment cost.|
|17|Nvidia RTX 5060 Ti|\~$420 - $460|\~3,460|\~93.5|12 GB / 16 GB|Blackwell mid-tier layout. Offers highly robust processing bounds, though carries a modern market premium.|
|18|AMD Radeon RX 580|\~$40 - $50|\~258|\~39.3|8 GB|Dirt cheap entry floor. Delivers text processing capability at the lowest possible cost parameter.|
|19|Nvidia P104-100|\~$30 - $40|\~311|\~46.1|8 GB|Low-profile budget node. Useful for multi-card distributed matrices where base components must be inexpensive.|
|20|AMD Radeon RX 9070|\~$550 - $600|\~3,164|\~119.7|16 GB|Next-gen RDNA4 architecture architecture layout. High performance density but subject to lower hardware-to-cost scaling.|

Key Strategic Takeaways from Vulkan results

  • Lowest Cost for Token Generation (tg128): The AMD Instinct MI50 and P102-100 completely distort the curve. The MI50 nets you over 100 t/s on a Llama-7B architecture for roughly $120, a metric that consumer desktop tiers require twice the budget to replicate.
  • Lowest Cost for Prompt Processing (pp512): Modern architectures rule prefill metrics due to hardware tensor capabilities. If prompt processing latency is your critical bottleneck, look at the RTX 4070 Super or RTX 5060 Ti, which punch far above their weight class on ingest speeds.
  • Best Combined Balancer: The GTX 1080 Ti and AMD Radeon RX 6800 hit the absolute "sweet spot" for standard desktop nodes. They avoid the strict cooling modifications or specialized software handling required by headless data center units (like the Tesla series) while maximizing bandwidth-to-dollar efficiency.

Chart and summary provided by Gemini and myself. I currently own RX 7900 GRE, MI50, P102-100, GTX-1080Ti, GTX 1070, RX 480/580.

Here is the breakdown of the cost-per-token-per-second (\\(\\div \\text{t/s}\\)) for each metric across the top 20 GPUs.

Lower cost values ($/t/s) mean you get more performance out of every dollar spent. Combined throughput represents a balanced arithmetic baseline of both prefill and generation.

|GPU Model|Est. Used Price|pp512 Cost per t/s|tg128 Cost per t/s|Combined Cost per t/s|
|:-|:-|:-|:-|:-|
|Nvidia P102-100|$45|$0.0881|$0.7162|$0.1569|
|Nvidia P104-100|$35|$0.1122|$0.7579|$0.1955|
|AMD Instinct MI50|$120|$0.1072|$1.1059|$0.1954|
|AMD Radeon RX 580|$45|$0.1744|$1.1445|$0.3027|
|Intel Arc A750|$100|$0.0929|$2.3441|$0.1788|
|Nvidia RTX 3060|$190|$0.1046|$2.5020|$0.2009|
|Nvidia RTX 4070 Super|$500|$0.1085|$4.5981|$0.2120|
|Nvidia RTX 2080 Ti|$215|$0.1139|$2.2033|$0.2165|
|Nvidia RTX 4070 Ti Super|$705|$0.1156|$5.4461|$0.2264|
|Nvidia RTX 5060 Ti|$440|$0.1271|$4.7054|$0.2476|
|AMD Radeon VII|$150|$0.1416|$1.4824|$0.2585|
|Nvidia Tesla V100|$200|$0.1437|$1.5434|$0.2630|
|Nvidia Tesla P100|$100|$0.1475|$1.5833|$0.2698|
|AMD Radeon RX 6800|$230|$0.1443|$2.2667|$0.2713|
|AMD Radeon RX 7900 GRE|$415|$0.1776|$3.5742|$0.3384|
|AMD Radeon RX 7800 XT|$365|$0.1809|$3.0862|$0.3418|
|AMD Radeon RX 7900 XT|$525|$0.1785|$4.2621|$0.3426|
|AMD Radeon RX 9070|$575|$0.1817|$4.8033|$0.3502|
|Nvidia GTX 1080 Ti|$120|$0.2049|$1.7712|$0.3674|
|Nvidia Tesla P40|$130|$0.2664|$2.1900|$0.4750|

The 10 worst GPUs based on performance-to-cost value are ranked below using the provided benchmark dataset and current secondhand market value trends. These values represent the highest cost per token per second ($/t/s). A higher number means you are paying significantly more money for every unit of inference speed generated.

|Rank|GPU Model|Est. Used Price|pp512 Cost per t/s|tg128 Cost per t/s|Combined Cost per t/s|Primary Bottleneck Profile|
|:-|:-|:-|:-|:-|:-|:-|
|1|Nvidia Tesla M40|$60|$0.6488|$1.5248|$0.9103|Worst Overall Value: Outdated Maxwell architecture yields critically low processing throughput across both prefill and generation.|
|2|Nvidia Titan V|$350|$0.4395|$3.3314|$0.7766|Premium Collector Tax: Despite HBM2 memory, a high up-front market premium makes its performance-to-dollar ratio poor.|
|3|AMD Radeon Instinct MI60|$150|$0.4062|$1.9191|$0.6705|Severely low prefill scaling limits its deployment utility relative to the much cheaper MI50 framework.|
|4|AMD Radeon RX 7600 XT|$260|$0.3092|$4.9038|$0.5817|Extreme Decode Bottleneck: A very narrow 128-bit bus forces an incredibly inefficient $4.90 per token/sec on decode loops.|
|5|AMD Radeon RX 6600 XT|$160|$0.2784|$2.9674|$0.5091|Limited by entry-tier bandwidth configurations that fail to translate into meaningful compute value.|
|6|Nvidia Tesla P40|$130|$0.2664|$2.1900|$0.4750|While popular for cheap 24GB capacity, missing native FP16 compute hardware tanks its relative speed value.|
|7|AMD Radeon RX 5700 XT|$130|$0.2454|$1.8379|$0.4330|Older RDNA1 compute layers drop performance significantly compared to modern secondhand equivalents under $150.|
|8|AMD Radeon RX 6900 XT|$400|$0.2104|$3.7037|$0.3982|Commands a high premium on the used market but struggles to scale its text generation speeds efficiently.|
|9|AMD Radeon RX 6750 XT|$220|$0.2114|$2.6836|$0.3920|Tightly squeezed by low raw compute density relative to its market price window.|
|10|Nvidia GTX 1070|$70|$0.2177|$1.6876|$0.3856|The basement floor of the Pascal generation. Replaced entirely by the vastly superior cost-to-performance curve of the P102-100.|

Tesla P40 made both charts.

💬 11 (+9) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/devshore · 22h ago
Accounting / Tax Filing (48GB VRAM)

Question 1: Which model? They used to make specialty-related models, like a model just for knowledge about plant-life or cars etc. For tax filing, is there a tax knowledge model to use, or should we use a non-specialized model like qwen something?

Question 2: Obviously one of the points of failure would be having it tally numbers by looking at CSV files, but we can avoid that by using other software for that. The question is: what software should that be? Maybe 2 different softwares are needed: 1 that is used for fetching bank info for tracking income and expenses (the AI would be used to categorized the transactions), and a software for tax filing based on the values from the first software etc. Which two softwares would work? Self-hosted preferably, and obviously would need some way for the AI to interact with (api, or mcp).

Has anyone set something like this up?

💬 9 (+5) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Miserable-Dare5090 · 23h ago
This is too true, I had to share post image

It’s just interesting to me that everyone ends in the same loop: There is NEVER enough VRAM.

▲
0
 
6👁
r/LocalLLaMA · u/Low-Future-9387 · 23h ago
Running a 3B roleplay finetune fully on iPhone: what we measured about keeping a small model in character (I make the app)

I make Castmates, a closed-source iOS app (free tier, paid Pro) that runs a 3B roleplay model entirely on the phone. Posting for the engineering notes, not to sell it. The app is mentioned once, at the bottom.

Setup: Impish Llama 3B (a Llama 3.2 3B RP finetune), plus our own rank 16 LoRA, fused and requantized to Q4\_K\_M from the fp16 base. About 2.0 GB, downloaded after install. llama.cpp with Metal, all layers offloaded, KV cache at q8\_0 (roughly 60 KB per token). No server, no account, works in airplane mode.

Limits first, because they shape everything:
- 4 GB phones are the floor. Weights x1.6 plus KV plus \~350 MB for compute buffers has to fit, otherwise we refuse to load rather than get jetsammed. Those phones stay at 4096 context.
- 6 GB phones get 6144 context and 8 GB phones get 8192. Llama 3.2 is natively 128K so no rope scaling is needed, it only costs KV RAM.
- It's slow. Early on we measured around 6 tok/s on an A18. Newer chips are faster but I'm not going to quote a number I haven't re-measured on the current build.
- The 2 GB download is the biggest drop-off in the app, so we use Background Assets to start it before first launch. It's non-essential on purpose (essential blocks launch).

What we learned about staying in role:
- Bigger window alone does nothing. Our history trim budget was the binding constraint, not n\_ctx. Scaling the trim to \~0.68 x n\_ctx took long-conversation fact recall from 25% to 75% in our lab. Verbatim history beat the lossy summarizer by a lot.
- Retrieval can hurt. Our BM25 memory retrieval re-injected superseded facts when the window still contained the newer one (says-stale +20.8pp vs no retrieval). Dropping any hit whose rare entities still appear in the live window fixed it (-18.8pp, CI \[-35.4, -6.2\]) without losing facts on the recall benchmark.
- Prompt tweaks mostly measured as zero. Single 3-4 seed runs were noise. Fixed-history micro-tests with N=16 and same-seed controls were the only thing we trusted. Example: "say my name" went 4/12 vs 11/12 purely from history length, no prompt change.
- Placement matters more than wording. A scene direction is ignored in the reminder slot (0/16) but lands 16/16 as its own block after the final user turn. A reminder after the user turn makes the model answer the reminder instead of the user.
- Guards beat prompts for the 3B's habits: rerolls on third-person drift about the user, invented names, and bare role-label output. Thinking mode didn't help: one-pass is impossible with this finetune and two-pass was noise at 2x latency.

What it still can't do: override a fact it can still read in context, so contradictions in a long scene stay a model ceiling.

The lab is a Python port of the production prompt and guard pipeline run against llama-server with fixed seeds, so every claim above came from a run, not a vibe. Happy to go into any of it.

The app is Castmates on the App Store if you want to try it. I'd rather get criticism of the approach than installs.

💬 4 (+1) open on reddit ↗
▲
4
 
1👁
r/LocalLLaMA · u/Disastrous-Work-1632 · 23h ago
chat with llm model with zero setup! post image

Hey folks!

Aritra here from Hugging Face. We introduce a no config, no key, no setup way to directly chat with a model hosted with the Hugging Face Inference Providers.

\ssh chat.hf.co\

And you are good to go. 🔥

Let us know what you think about this.

▲
0
 
8👁
r/LocalLLaMA · u/Aggravating-Push-207 · 23h ago
LFM 2.5 5.4B

would be good for laptops, 8B A1B is a bit worse than 2.6B dense imo, not worth the speed bump ime as you can't even use it for subagents with low vram/ram

💬 9 (+1) open on reddit ↗
▲
0
 
8👁
r/LocalLLaMA · u/RA2B_DIN · 24h ago
Eron v1.4: A native iOS client for Ollama & local models with zero-buffer streaming, thinking tokens, and local Apple Home/Calendar tools

Hey everyone,

Most mobile LLM setups for iOS suffer from two issues:

  1. Web UIs in mobile Safari tend to drop streaming the second your screen locks or you switch apps, with zero access to native iOS APIs.
  2. Most App Store clients push aggressive $15/month subscriptions and route your private prompts through their own cloud proxies.

I built Eron as a clean, native iOS companion specifically for people running their own local hardware (Ollama, vLLM, LM Studio) or using their own API keys (BYOK).

Technical details & v1.4 architecture:

  • Direct Socket / Zero Proxy: Direct HTTP/WebSocket connection straight to your local IP or Tailscale/WireGuard node. No intermediate servers, no telemetry, no account required.
  • Zero-Buffer Streaming: Rewrote the streaming pipeline from scratch. Instead of waiting for sentence buffers, tokens render as raw chunks as fast as your GPU outputs them.
  • Reasoning Stream: Native streaming and collapsible rendering for <think> reasoning blocks (DeepSeek R1, Qwen reasoning, etc.).
  • Local iOS Tool Calling: If your local model supports function calling, Eron provides native bridges to Apple Reminders, Calendar events, and HomeKit smart home control directly from your prompt.
  • Workspaces: Isolated project workspaces with persistent custom system prompts to keep coding contexts separate from daily chats.
  • v1.4.1: Native dual-screen layout ready for the upcoming iPhone Duo form factor.

Pricing & Community Codes:
It’s a $2.99 one-time purchase on the App Store

To get feedback from this community, I have 20 App Store promo codes to give away to anyone running a local setup who wants to test it for free.

Just drop a comment with your setup (what models/hardware you’re running) and I’ll DM you a code!

App Store: https://apps.apple.com/app/eron/id6760043923
Setup docs: https://henningwinter.com/app/eron

Self-promotion disclosure: I am the sole developer.

💬 14 (+1) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/sachasayan · 19h ago
I put together some text-only Qwen3.5 2B, 4B and 9B MLX 4-bit packages (including abliterated variants) perfect for local use — have at 'em. :)

Hey folks — I put together some text-only MLX 4-bit packages of Qwen3.5 great for local inference because nothing else quite exists in those specific configurations/sizes. Sharing them here in case they’re useful to anyone else running models on Apple Silicon.

Brief summary: There are six packages: 2B, 4B and 9B, each in the original Qwen version and the corresponding Huihui abliterated version. They’re text-only (vision-removed) and 4-bit, which makes them incredibly svelte (Only 1GB for the 2B version!) and great for running passively with resources to spare.

I've got them currently working on my writing software Minstrel doing summarization tasks and making contextual decisions (more on this later!), but they're of course free for everyone to use. Hopefully someone finds them useful!

Models and download links on Hugging Face

Note: These build on existing upstream models and community conversions, so my work here was mostly the text-only packaging and MLX conversion. The Huihui 4B GGUF source was already text-only, so all it needed was conversion.

Credit to Qwen, Huihui, and the community conversion authors linked in each model card.

Let me know if you run into issues. ✌️

▲
0
 
7👁
r/LocalLLaMA · u/sleight42 · 21h ago
Qwen 3.8 Flash Next is smarter than the new Siri

... and almost no one was surprised. At least that's what I imagine.

I gave Siri a PDF of medical provider statements and asked for a sum of the payments. It defaulted to finding the value on the first page. I pointed out it was wrong. Then it summed payments across a few pages. Still wrong. I gave up.

I handed the same document to QFN through Hermes. It extracted the text and summed. Then it used vision to doubl-checked itself.

So...

  1. Siri is still an idiot
  2. QFN 3\_xxs is quite good at administrative agent tasks
  3. There goes another reason to want to buy a new iPhone.

UPDATE: Evidently, many commenters are unaware that Apple leverages cloud-hosted AIs (Gemini) and uses more than just on-device AI.

UPDATE 2 (for the less than generous commenters): Please consider stopping for a moment, before commenting and ask yourself, "Will this comment make the world better or am I just trying to make someone else feel bad?" If the latter, it's probably best for the world, and your own psyche, to exercise forbearance.

UPDATE 3: The point of the post was to *celebrate* what many of us have access to now that most people buying high end phones do not.

💬 23 (+6) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/texasdude11 · 18h ago
GLM 5.3 Flash decided it's Claude, then lectured me about admitting when it doesn't know something lol post image

So yesterday night I was testing my local setup, GLM 5.3 Flash running through some custom pipeline thing. Asked it a simple question only, "what is your knowledge cutoff?"

First line of the thinking itself it says "I'm jarvis-thinker, custom model name, but based on Claude." Based on Claude?? Nobody told it that. There is nothing anywhere saying that. It just made up its own identity and moved on like it's normal thing. Full confidence. I understand that so much of the training traces have that in it, that it all has gotten polluted :) that's not the point tho... Keep reading.

Then next it's estimating the cutoff date. "Claude models typically have early 2025 cutoffs, I should say roughly early 2025." Note the word "should say". It knows it's guessing. It literally wrote "I don't know the exact date with certainty" and then went ahead and gave the date anyway.

Now the best part. The final answer it gave me, it has one bullet point like this:

"I know what I don't know. If you ask about something recent and I'm not sure, I'll tell you instead of confidently inventing an answer."

Dude! You invented an answer 2 seconds back. About yourself. The one thing you should actually know. Your whole identity is a hallucination and the very next output you're telling me you never hallucinate.

I thought it was funny and maybe a couple others here will get a chuckle out of it too.

💬 10 (+7) open on reddit ↗
▲
0
 
7👁
r/LocalLLaMA · u/Jromagnoli · 19h ago
I have potato laptops, which cannot run many models. Would "hosting"/using via a cloud service work?

E.g. hosting a cloud server/GPU rental, and using "huge" models which otherwise would be impossible for me to run, would it technically work? (e.g. I boot it up from a "site" or address) How would I "save" my work, and the cost to host/run? And is it "worth it"? any experiences from those who have used the services?

(also does anyone know of any good/private cloud-server/GPU service?)

----

(my laptop specs if anyone is wondering:

  • Acer swift 5 SF514-55TA (main, budget laptop)

| .| . |
| --- | --- |
| Installed Physical Memory (RAM) | 16.0 GB |
| Total Physical Memory | 15.8 GB |
| Available Physical Memory | 4.49 GB |
| Total Virtual Memory | 25.3 GB |
| Available Virtual Memory | 6.33 GB |

  • Acer Nitro 5 AN515-53 (not used currently)

| . | . |
| --- | --- |
| Installed Physical Memory (RAM) | 8 GB |
| Total Physical Memory | 7.85 GB |
| Available Physical Memory | 4.96 GB |
| Total Virtual Memory | 9.72 GB |
| Available Virtual Memory | 5.89 GB | )

💬 18 (+5) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Loose_Doubt367 · 15h ago
what can i do with my local (qwen3.6 35b) model inside pi harness or any other harness

I've been playing a lot with the models and have settled with the qwen3.6 35b and pi, but im not sure on what to do other than code html games. Any suggestions?

▲
0
 
4👁
r/LocalLLaMA · u/Distinct-Pie2389 · 14h ago
llm-tune: getting local models that actually perform post image

Made a little agent skill called llm-tune to help find the best settings for running local LLMs on your hardware.

Still working on it, but I'm looking to test it across more GPU setups. If you try it out, feedback and benchmark results are welcome.

Supported:

  • Architectures: Dense, MoE, hybrid MoE/Mamba
  • GPUs: NVIDIA, AMD, Intel Arc
  • Apple Silicon: M-series Macs via MLX
  • Engines: llama.cpp, Ollama, vLLM
  • Tuning: Quantization, context/KV cache, GPU offloading, MTP, sampling, reasoning, and agent harness settings
  • Benchmarks: VRAM/RAM usage, tokens/sec, context recall, and output quality

Currently measured on an RTX 4090; other hardware and backends are documented but need more real-world testing.

💬 7 (+3) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/MrCatberry · 9h ago
Reverse Engineering Web Application/Service

Hi Guys!

Is there a known harness/workflow that makes it easier to Reverse Enginerring a Web Application/Service thats behind a payed subscription?

My current problem is that I use a service that costs me quite a lot of money but still does not have all the features or tweaking options I need.

Now I'm asking myself, if it would be best to describe every feature myself or if there is a way to let a LLM "explore" the Web Application/Service by itself and write it's own notes what it needs to code.

Did somebody here do something familiar and has some tips?

Thanks in advance!

💬 8 (+5) open on reddit ↗
▲
0
 
3👁
r/LocalLLaMA · u/General-Spite1222 · 10h ago
NVFP4 is the GOAT, prove me wrong.

It being on par with bf16: https://huggingface.co/nvidia/Qwen3.8-27B-NVFP4#evaluation

It being on par with FP8: https://huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4#evaluation

Who here is smarter than Nvidia and can price them wrong? I see a lot of trash talk, but no numbers to back it up.

NVFP4 is as good as Q8. Prove me wrong! Feel free to downvote if you cant prove otherwise.

💬 23 (+8) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/Old-Sherbert-4495 · 7h ago
Qwen3.8: Flash Next iq3 xxs is dumber than 27B iq3 xxs?

it consistently could not pass this. had to follow up couple of times and by the time it has already exhausted 120k context limit.

"build me a 3d simulation of the solar system in a standalone html file. no three.js, No web GPU, no any extra libs, just plain html, css, js, and svg."

On the flip side 27B, one shots them like perfectly.

One thing i noted was flash next version had way more features and a pretty complex UI. 27B some what simpler and gets the thing working.

I run this to get a vibe of the model. Then ran some agentic coding tasks, it sure does think wide but when it comes to the implementation it always fails. and needs few follow up steps. overall loads of tokens used.

Has this been the same experience you had?

💬 33 (+14) open on reddit ↗
▲
4
 
2👁
r/LocalLLaMA · u/leo-k7v · 7h ago
Supertonic 3 TTS dissolution and liquidation?

I use their model in my up with hand-rolled inference in C.
HF page says:
"This project's sample code is released under the MIT License."
and "OpenRAIL-M License" for the model.

Can I still use their model after liquidation?

💬 6 (+2) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/paq85 · 8h ago
How much does your local LLM server really cost? (power draw + tok/s, self-hosted vs cloud API, free in-browser) :)

I've been tuning local LLM setup for many months, and the number people actually want is the electricity bill. This calculator takes your GPU's power draw (W) and token speed (tok/s for prompt and decode), plugs in your energy price per kWh, and shows cost per hour, per request, monthly, and cost per 1M input tokens — with and without KV cache.

There are 3 GPU presets (RTX 4070 Ti Super, RTX 5090, RTX 5090 eco) if you want to start from realistic numbers and adjust from there. The math treats cached tokens at 1/100 the time of uncached, so the realistic scenario is fixed at 60k input / 50k cached / 2k output / 90% uptime.

The cloud API comparison compares each self-hosted profile against a reference API ($0.25/1M input, $1.20/1M output) at the same scenario, so you can see whether self-hosting actually saves money or costs more.

Everything runs in your browser, nothing is uploaded and there is no sign-up. From my experience, the per-request number is the one worth comparing with the cloud — the monthly bill is just the per-request cost times your actual request rate.

https://appdoesit.com/apps/llm-cost-calculator — it's one of the 139 free tools in the catalog.

Try it by yourself :) — how does your server compare?

▲
13
 
1👁
▲
4
 
1👁
r/LocalLLaMA · u/mrgreatheart · 3h ago
Epyc inference rigs

I posted on here a while back asking for advice on upgrading to an Epyc system. Thanks to all who commented.

I’ve taken the plunge and ordered:

\- Epyc 7443 CPU (24 cores, 48 threads, 128 lanes)
\- H12SSL-NT-B motherboard (the variant with two 10Gb Ethernet ports)
\- 256GB 2666MHz 2Rx4 RAM

To this I will be adding 72GB of VRAM across four GPUs:

\- 3090 (24GB)
\- 5070 Ti (16GB)
\- 2 X 5060 Ti (16GB)

Unfortunately I have to wait 2 weeks for the motherboard, so naturally I’m wondering what difference it will make.

The upgrade should:

\- double the PCIe bandwidth of all four GPUs (x4 to x8 for the 5060s and x8 to x16 for the other two). They will also all be on proper CPU lanes instead of chipset and dodgy m.2 adapters.
\- increase the RAM memory bandwidth by at least 50% (96GB to 170GB/s theoretical, perhaps 140 in practice)
\- increase total RAM by 192GB (64 to 256)

And it will open up the possibility for more GPUs later.

Obviously this won’t make much difference to dense models although I’m hoping for a nice bump in prompt processing due to the increased link bandwidth.

But I’m excited about bringing my total memory pool up to 328GB and opening up access to things like GLM5.3-flash and larger quants of Qwen3.8-flash-next.

Has anyone here built something similar?

How does your system handle large MoE models that fit fully in RAM & VRAM?

▲
11
 
1👁
r/LocalLLaMA · u/ResearchCrafty1804 · 3h ago
Qwen-Image-2.1-Turbo released!

Qwen-Image-2.1-Turbo, create and edit images in just 8 denoising steps! Open weights now available!

Built on Qwen-Image-2.1, Turbo is an accelerated checkpoint on the same 7B visual generation architecture.

Fewer steps does not mean lower quality: it still generates strong 2K images from text, and supports continued creation through natural-language edits, from adding accessories to changing a scene.

Start directly with Diffusers: load QwenImage21Pipeline and the checkpoint’s recommended 8-step sampling schedule is ready to go.

Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo

▲
3
 
1👁
r/LocalLLaMA · u/OkFly3388 · 5h ago
rtx 4090 + huawei atlas duo for qwen flash next ?

I have rtx 4090 and 64 gb of ddr4 ram. This is enough to fit smaller quants of qfn, but I found them not great and switched back to qwen3.8. However, there are option that bother me a lot, buy huawei atlas duo, thats another 96 gb of ram thats twice faster than system ram and another AI accelerator, thats faster than cpu.

Anybody tried that ?

▲
24
 
1👁
r/LocalLLaMA · u/jacek2023 · 5h ago
tencent/Youtu-Parsing-Omni · Hugging Face

Youtu-Parsing-Omni is a compact (5B) omni-modal parsing model. Given a single input — a document page, a natural image, a chart / flowchart, a geometry figure, an audio clip or an audio-visual video — it produces one structured JSON envelope that covers both perception (layout elements, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events, camera motion) and cognition (captions, narratives, reports). The output family is selected by the task prompt (--task in the examples, keys of prompts/youtu_parsing_omni.json).

|Input|--task|modality / subtype|Key contents|
|:-|:-|:-|:-|
|Document page|document|image / document|layout elements with bbox, text / LaTeX / OTSL tables / Markdown charts / Mermaid flowcharts, reading order|
|Natural image|natural_image|image / natural_image|entities and text with bbox, tags, captions, global description|
|Chart|graphics_chart|image / document|one chart element: Markdown table, notes, caption|
|Flowchart|graphics_flowchart|image / document|one flowchart element: Mermaid, caption|
|Geometry figure|graphics_geometric|image / document|one geometric element: points, lines, arcs, shapes, geometric relations and measurements|
|Audio|audio|audio / –|vocal / non-vocal segments with timestamps, speakers, ASR, timbre / scene captions, acoustic events|
|Natural video|natural_video|video / natural_video|temporal segments with visual elements, actions, interactions, camera motion, audio track|
|Text-rich video|textrich_video|video / text_rich_video|segments with OCR + ASR and a Markdown structured_report of the whole video|

Highlights (see the technical report for details):

  • Unified schema – one JSON envelope for seven parsing families, driven by the task prompt.
  • Omni encoder – image, audio and interleaved audio-visual video inputs (frames + audio track) in a single model.
  • Strong results at a small size – state-of-the-art on OmniDocBench v1.6 (96.96 Overall), best open-weight model on OmniParsingBench (75.08 Avg., second only to Gemini-3-Pro), and competitive with specialized models on chemical-structure (ChemOCR) and music-score (PDMX-Synth) recognition (results).
  • Easy to serve – a vLLM plugin, pinned serving settings, task prompts and inference examples are included.
▲
21
 
1👁
▲
10
 
1👁
r/LocalLLaMA · u/NoobSolid26 · 5h ago
I made a portable AI memory format that transfers both context and skill

Last month I created a portable memory format as a .txt file as a small project. I named it as: Memory Journal, in a shape of a skill card. It's essentially a large prompt wrapped in a text file.

Initial testing went better than I have expected, then did more tests and improved it in a course of a month, Because it was interesting to work on. It's meant to be general purpose, it can handle both casual and technical sessions, even though it can bend a bit. I found the it kinda cool so I decided to share it here. It's not perfect by any means but it works.

The philosopy behind was simple "If an AI knows the context, it can describe how it created something too". In addition to context it has ability to be reproducible, it can transfer skill, transfer mistakes. along with proven, unproven and disproven facts. Journal mostly gives free reign to the writer (AI) and stays kinda neutral. Journal has been split into A/B/C sections, that serve differen purposes.

\[Capabilities\]:

  • Right now it can allow you to resume your work with journal + project files.
  • AI can write a journal about EGO engine XML file conversion, and readers can reproduce the converter from journal alone. (Deepseek 4.1 (deep think mode) and Claude Sonnet 5 (high effort) recreated the converter in 1st try.)
  • Journal can also cause drifts in behavior and vocabulary if you use a philosophically charged journal

\[Usage\]:

  • You give this "card" to an AI then they will create a text file as a memory journal.
  • You can supply additional commands by yourself, like: "add completion status to goals" or "don't add x part to journal" and such.

https://preview.redd.it/19xcinownfuh1.png?width=795&format=png&auto=w…

\[Note\]:
It's best to use this journal with a capable AI model preferrably with an effort slider. Also yes skill card larp is actually part of the design, I know it looks eccentric but, isn't it cool?

\[Memory Journal format itself and outputs of it is down below\]:

Memory journal format itself: Memory Journal S4 - Pastebin.com

XML converter reproduction journal: XML Converter S3 - Pastebin.com
TRD2 Modding journal (4th writer on the line): TRD2 Modding Final Journal - Pastebin.com
Weird casual chat journal: Psychic pepperoni journal - Pastebin.com
Joke casual chat journal: Pun journal - Pastebin.com

▲
0
-1
8👁
r/LocalLLaMA · u/ZenZombie117 · 21h ago
I was doing some testing on Strata. vs llama vs. runner and was missing something

The MTP head for the file seems to be accountable for some of the speed of strata (maybe not news to anyone but me but i'll digress). To be able to do the comparison I produced A head for ISTA-DASlab's GGUF that strata uses. while on it I also went ahead and produced a "head" for one of my own quants and llama seemed to have a big gain from it.

https://huggingface.co/Joakimpalm-Zen/Qwen3.8-27B-GSQ-RCO-IQ3\_S-recovered-GGUF

Runner still has a long way to go, i need to do some architectural improvements... But llama had a big gain so there's that.

and the bigger one:

ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF (header is here: https://huggingface.co/Joakimpalm-Zen/Qwen3.8-Flash-Next-MTP-GGUF )

Llama sees some improvements with the MTP header, Strata is still WAYYY faster though.

Thought they might be useful for someone else so thought i'd just share, now back to runner!

💬 19 (+19) open on reddit ↗
▲
0
-1
3👁
r/LocalLLaMA · u/harderisbetter · 22h ago
How to get Qwen 3.8 to properly use skills.md?

Noob here, I'm broke so I only use Qwen 3.8 27B through the Qwen chat website via my potato pc. I tried to copy paste the skill.md in the customization option in my profile, I also tried to attach it as part of the prompt, nothing works.

I understand that there is no free API for the 3.8 model, and I don't want to use those sketchy temporary affiliate links that will overcharge my credit card after the trial.

Is there a way to properly use Claude skills (downloaded as zip folders from github) with Qwen for free? I only have free Claude desktop.

💬 9 (+4) open on reddit ↗
▲
0
-1
2👁
r/LocalLLaMA · u/NoahPersaud · 23h ago
Tokenizers and HuggingFace ONNX Model Pipeline (UE5)

I created a tokenizers plugin and an HuggingFace ONNX model pipeline plugin for UE5.

The tokenizers repo is fairly complete for Windows, but does not currently support other platforms.

The pipelines repo only supports text embeddings, text classification, image classification and object detection for now. I plan to add a lot more in the future.

The plugins are open source. Claude was used to build both, but I started Tokenizers myself years ago.

Contributions are welcome.

▲
0
-1
7👁
r/LocalLLaMA · u/Aggravating-Push-207 · 20h ago
guys is it a dumb idea to use one of the smaller jev knockoffs to decide which speculative draft is better

as in like

generated so far: A B C
draft 1: D E F
draft 2: G H I
draft 3: J K L

then some small model that runs locally really fast decides which speculative draft is best

but only when the token entropy is high

this would be in the generation loop itself

💬 10 (+3) open on reddit ↗
▲
0
-1
4👁
r/LocalLLaMA · u/MasterSama · 13h ago
How is it possible to enable tool calling like file read/write/execute in Strata (Qwen3.8-next-flash) web ui?

I've just installed Strada and Qwen3.8-next-flash on my PC and I was amazed by it.

The problem is, I cant seem to find a way to enable or add tool callings, so it can read/write/execute files on my PC. How do you do that in Strata?

Thank you very much in advance

Update:

Per suggestions, I went on to use a harness and ended up liking the deepseek harness a lot. its working great. Thank you very much everyone.

💬 14 (+14) open on reddit ↗
▲
0
-1
2👁
r/LocalLLaMA · u/Loose_Doubt367 · 10h ago
What features do i unlock by switching rx6700xt 12gb (AMD) to 3060 12gb (NVIDIA) in llama.cpp?

is expecting higher tps the only feature that i'll be receiving?

💬 15 (+2) open on reddit ↗
▲
1
-2
8👁
r/LocalLLaMA · u/Bulky-Priority6824 · 22h ago
What are you using for NVFP4 and Do you like it?

What are people using to run nvfp4 on multi-gpu?

the only thing i can get to run is unsloth and vllm is too slow and takes FOREVER to fucking load. TensorRT-LLM has too many issues, so what are people using?

And do you like nvfp4 vs q4 qguf for qwen 3.8? apples to oranges is nvfp4 closer to Q6 gguf than q4 gguf is?

well i tired the model here https://huggingface.co/neroued/Qwen3.8-27B-nvfp4-NInfer

which works with https://github.com/Neroued/ninfer/tree/master

and initial testing has not been great for code but vision and tool calling is very impressive. Brief testing on complex scenes showed slightly better than what I've seen on q6 gguf

but im going to revisit surely im missing something, i had to spend a lot of time wiring ninfer into my frontend so ill look at it again with fresh eyes.

The speed is fantastic on 2x5060ti with 197k ctx and model loading in 4-6 seconds is wild

https://imgur.com/a/ahDnoAZ

💬 41 (+26) open on reddit ↗
▲
0
-2
8👁
r/LocalLLaMA · u/rodrigodevbits · 21h ago
Local hardware vs Cloud APIs: Is it actually worth buying a Mac Studio or 2x DGX Sparks for real agentic coding?

Hey r/LocalLLaMA,

I’m trying to figure out if I should drop serious cash on a local setup for heavy agentic coding (letting agents read whole codebases, refactor multi-file repos, run terminal loops) or if I should just keep paying for Claude Code and ChatGPT.

Right now, cloud APIs are driving me crazy. If you do any serious agentic coding, you easily blow past $500+ a month in API bills. And even if you have the money, you hit a hard rate limit after 2 to 4 days of heavy work and have to wait for a reset. It completely kills my momentum.

The global memory shortage has messed up hardware prices, but going local is looking pretty tempting just to escape these cloud limits. I have a budget of around $14k–$15k max. Here are the two routes I'm looking at and the headaches I'm trying to weigh out.

1x Mac Studio M5 Ultra (512GB RAM)

  • The Cost: Around $13,500 - $14,000 USD because Apple charges an absolute fortune to max out the unified memory.
  • The Good: You get 512GB of VRAM on a single machine. You can easily fit huge models (like DeepSeek V4.1 Flash or GLM-5.3 Flash) and give them huge 128k+ context windows without the system crashing.
  • The Catch: Time-to-First-Token (TTFT) is going to be slow. When the agent reads a 60,000-token codebase all at once, the Mac is going to sit there and "think" for like 2 to 3 seconds before it starts typing. Once it actually gets going, generation is about 35+ tok/s, which is fine, but that initial pause might get annoying.

2x Nvidia DGX Spark Units (Linked directly)

  • The Cost: Right around $14,000 USD (Nvidia jacked the price of the 128GB version to $6,950 due to the component shortage, so two nodes plus cables puts you right there).
  • The Good: Prefill is blazing fast because of the Blackwell cores. It will ingest thousands of lines of code almost instantly. No waiting around for the first token.
  • The Catch: Stacking two nodes only gives you 256GB VRAM total. This means you are seriously restricted on what models you can run. You can't run the massive 300B+ giants unless you use super compressed low-bit quants (like IQ3 or IQ2) just to fit the model and a decent context window without hitting Out-Of-Memory (OOM) errors. If your codebase is too big and the KV cache overflows that 256GB limit, your speed drops to zero.

How the math looks to me

If I take that $14,000 and look at it compared to what I'm spending on APIs:

  • At $500 a month, $14k pays for about 2 to 2.5 years of cloud access.
  • But again, cloud means hitting limits every few days and sitting around waiting for a reset. Local means I can run it 24/7 with zero downtime.

The main reasons I want to buy hardware:

  • No limits: No quotas, no rate limits, no waiting for a reset. I can run infinite loops, try weird models, tweak my tools, and never see a "Quota Exceeded" message.
  • Privacy: My code and data never leave my room. No corporate data center is logging my repo.

The big downsides I'm worried about:

  • Depreciation: The moment I buy a $14k cluster, it starts getting old. In two years, cloud models will be way smarter, but I'll still be stuck with the same physical VRAM limits.
  • Friction: Local agents love to break. I feel like I'm going to spend hours messing with vLLM, debugging tool-calling errors, and dealing with quantization loss instead of actually getting work done.

What do you guys think?

I'm really trying to figure out if anyone here has built a mini-cluster specifically to escape the $500/month cloud tax and quota lockouts.

Did it actually replace your Claude subscription for real development work, or did it just end up being an expensive toy? How bad is the TTFT on the Mac when loading huge repos, or are you constantly hitting OOM errors on a 256GB Nvidia setup?

Would love to hear some real-world experiences before I burn a hole in my wallet.

💬 76 (+35) open on reddit ↗