52 posts · 1 sub · RSS
← prev Friday, September 25, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
19
-1
11👁
r/LocalLLaMA · u/fredconex · 14d ago
Ion v0.2.0 — No install. No backend. Just one HTML file.

The new Ion is available, a harness that run directly from a single HTML file, no install or backend required: is now more capable, more customizable, and has better tools!

💬 30 (+9) open on reddit ↗
▲
191
-2
44👁
r/LocalLLaMA · u/dreamingwell · 15d ago
M5 Ultra 80Core GLM-5.3-Flash on DwarfStar Speeds post image

I've been playing around with various models on the M5 Ultra 256GB 80-core Mac Studio. These are the results over many rounds of agentic inferencing.

I'm happy with the performance. Glad to have the large amount of RAM. But it does feel like the GPU is underpowered for this amount of RAM. I'm wondering if a 512GB unit for AI inference makes sense at all - because the GPU will be the clear bottleneck.

💬 114 (+4) open on reddit ↗
▲
311
+2
33👁
r/LocalLLaMA · u/Top-Evidence174 · 14d ago
Mica v0.1 4B got an iron pickaxe in real Minecraft without generating a single token post image

Mica v0.1 4B playing a real Minecraft 1.20.4 server. Video attached.

How it works

\- Each step the bot's live game state (inventory, nearby blocks, entities, last result) is written out as text.

\- Mica scores the candidate commands and picks the next one. It never generates text. It reads the probabilities of the answer label tokens, so output tokens are 0.

\- The chosen command is executed in the game with Mindcraft's skill library (Mineflayer bot).

Run

\- 23 decisions from an empty inventory to an iron pickaxe: logs, planks, crafting table, wooden pickaxe, stone, stone pickaxe, furnace, iron ore, smelting, iron pickaxe

\- About 90 to 150 ms per decision

\- llama.cpp, Q5\_K\_M, RTX 3090

About the video

\- The right panel shows each decision as it happened: the candidates, Mica's probabilities, the pick, and the result. Every step is also listed in the history feed.

\- Long actions (walking, mining, smelting) are sped up, with the speed shown on screen. Back-to-back retries are shortened in the edit.

\- The HUD and the crafting/furnace screens are drawn from the bot's logged inventory.

Weights: https://huggingface.co/sky7350/Mica-v0.1-4B

Code and server: https://github.com/akivet/Mica-v0.1-4B

💬 53 (+1) open on reddit ↗
▲
305
+1
40👁
r/LocalLLaMA · u/recentheartbroken · 14d ago
I ran the actual break-even math on buying vs renting an H200 box, and it is not where I expected

Every rent-vs-buy thread I read has confident people on both sides, but not many actually show the numbers. So I finally ran the numbers for our own decision. Posting the working here in case it is useful, or feel free to point it out in case someone thinks it's wrong.

An 8-GPU HGX H200 server lands somewhere near $320k-$420k, with roughly $370k being a reasonable midpoint.

On the rental side, the median on demand H200 price across 34 providers was about $4.40/GPU-hour as of September 18. The $2-$3 rates you sometimes see are closer to spot pricing.

Using the $370k as midpoint and a rental equivalent of $35.20/hour, the hardware only break even works out to approximately:

\-> 14.4 months at 100% utilisation

\-> 24 months at 60% utilisation

\-> 36 months at 40% utilisation

Ofc, most small teams with bursty training and steady inference aren't sustaining 100% utlisation.

This is only hardware level comparison. There are at least four other things to include:

Power and cooling (I was quoted more for a colo cage than I had budgeted)

Depreciation (Whatever you assume, halve it. Resale on last-gen datacenter parts is thin)

Your own time.

Idle hours.

I work at B3 Labs, which sells and hosts NVIDIA GPU systems and helps owners monetise their idle capacity. That gives me a commercial reason to run this math, but I've tried to keep the assumptions neutral.

My conclusion was that roughly 60% sustained utilisation for 2 years, owning wins. Below 40%, renting wins. You can also sell your idle capacity to offtake networks and offset the cost of your device.

I'd love to hear what utilisation are people here actually seeing?

💬 234 (+1) open on reddit ↗
▲
75
 
34👁
r/LocalLLaMA · u/AdInternational5848 · 14d ago
4-5 days replacing Claude w Qwen 3.8 Next

Hi, human here with rambling thoughts to share. Feel free to skip

Overall, I don’t feel like I’m missing much; if anything. On my hardware(M1 ultra w 128Gb) it’s probably not as fast as Claude but I’ve been using opencode for research and other business related tasks and it’s been getting the job done and learning.

I might dive back in for the multi agent workflows that speed things up with a cloud provider but I’m working on setting up different slots as I refine my custom harness which works well for chat but not all the actual fun and useful stuff. Claude “knew” me better but that’s to be expected after months of back and forth with it and I’m honestly not sure I want them to know me this well.

Not expecting a lot of responses but it’s pretty cool I’m able to replace the service a billion dollar company provides with a Mac Studio and free software.

Any advice on better optimizing my system to improve speed without sacrificing accuracy?

Planning to work on optimizing DeepSeek V4 0731 and GLM Flash but they don’t seem to be “better” than Qwen 3.8 next so i decided to start spending more time using instead of optimizing for prefill and tokens per second.

💬 52 (+1) open on reddit ↗
▲
67
-2
40👁
r/LocalLLaMA · u/wadeAlexC · 14d ago
Qwen3.8-27B: Using KV Cache Transplants to Boost Output Quality

Since my last post, I've been thinking about different options for dynamic performance degradation, trying to squeeze as much high-quality inference out of my GPU as I can.

Over the weekend I read this really interesting paper: Cache-to-Cache: Direct Semantic Communication Between Large Language Models. In it, the authors describe running multi-llm agent systems. But rather than having agents talk to each other through a harness+tool calls+messages, they had agents pass context to each other by fusing one agent's kvcache directly into another's.

Assuming this is possible, you could imagine this being a faster, more complete way to pass context between agents: rather than one agent producing a summary/handoff message, you literally just rip out its working memory and graft it onto the target model.

They go on to describe how they do this, the TLDR being they trained a small neural network to be able to "convert" between the source and target model's internal representations, allowing them to fuse kvcaches of models of differing size and even architecture.

This got me thinking: what if I wanted to reuse a kvcache between different quantizations of the same model? I mean, same architecture, same training process ... shouldn't they be compatible, even without training a 'converter'?

And what would happen if I started inference with a high-precision quant, then swapped in a lower-precision quant to 'take over' when running low on device space? Could I get better results than just running the lower-precision quant from the start?

Spoiler, the answer to all of this is yes (on the benchmarks I ran)! I detail the specific experiment I ran below.

Methodology

I generated difficult NIAH-style tasks at different context lengths, and had 5 different Qwen3.8 quantization strategies battle it out!

For these tasks, I used three different quantizations of Qwen3.8-27B, each made by unsloth:

\- UD-Q6\_K

\- UD-Q4\_K\_XL

\- UD-IQ3\_S

Strategies

From these quants, I defined three static-quant strategies to run tasks against:

  1. IQ3\_S: f16 kvcache, max ctx 196,096
  2. Q4\_K\_XL: q8\_0 kvcache, max ctx 183,296
  3. Q6\_K, f16 kvcache, max ctx 175,104

Note that the Q3 and Q4 strategies have ctx windows sized for a 24 GiB GPU, while the Q6\_K case requires > 24 GiB to run. This is to evaluate how closely static quant strategies on a small device measure up to a static quant strategy on a larger device.

The idea is to see if dynamic quantization strategies can make up some of that difference!

Speaking of, I defined two dynamic-quant strategies to compare against each of the small precision static-model cases. These dynamic-quant strategies also both feature ctx limits sized for a 24 GiB GPU.

IQ3\_S Comparison

For this strategy, I ran the tasks against a multi-quant strategy with a worst-case model quantization of IQ3\_S:

\- Start task with Q6\_K, f16 kvcache, max ctx 54,272

\- Then swap in Q4\_K\_XL, f16 kvcache, max ctx 109,312

\- Then swap in IQ3\_S, f16 kvcache, max ctx 192,096

In the data, you can see this strategy labelled as Q6→Q4→Q3, f16 KV.

Q4\_K\_XL, q8\_0 kv Comparison

For this strategy, I ran the tasks against a multi-quant strategy with a worst-case quantization of Q4\_K\_XL, q8\_0 kv:

\- Start task with Q6\_K, f16 kvcache, max ctx 54,272

\- Then quantize the model's kvcache to q8\_0. Max ctx: 91,136

\- Then swap in Q4\_K\_XL, f16 kvcache, max ctx 109,312

\- Then quantize the model's kvcache to q8\_0. Max ctx: 183,296

(For the mid-run quantizations, I used the hot-reload method described in my last post. The f16<->q8\_0 conversions are handled by the same llama.cpp fork.)

In the data, you can see this strategy labelled as Q6/f16→Q6/q8→Q4/f16→Q4/q8.

Tasks

I generated dozens of unique NIAH ("needle in a haystack") tasks, which direct models to parse large volumes of input text and follow specific instructions scattered throughout the text to retrieve a secret value. (h/t gkamradt/needle-in-a-haystack for some of the source material)

I went with NIAH because it felt like a reasonable way to evaluate coherence for the multi-quant strategy. Each task requires the model to reason through a sequence of 'steps' buried inside distraction text, so a model with a transplanted kvcache would need to be capable of picking up the train of thought precisely where the source model left off.

Also, NIAH doesn't require a complicated test setup, and the answers are objectively right or wrong.

For each task and test case, I measured the following:

\* Result (correct/incorrect)

\* Total tokens generated

\* Total time taken

I included time taken despite each quant having a very similar prefill/decode speed because I wanted to demonstrate that the multi-quant approach does not take noticably longer than running a single-quant strategy. Transferring the kvcache from one quant to another means we don't need to repeat prefill!

Results

In total, the benchmark tasks I laid out represented 180 distinct runs, and which took my GPU 14h, 26m to complete.

The biggest offender here was the IQ3\_S/f16 strategy. Especially for the heavier tasks, it consistently generated upwards of 50k reasoning tokens, and all-too-often completely max out its context window (\~196k) before failing to ever generate a response.

Still, it holds up reasonably on the shorter tasks, even managing to score higher than Q4\_K\_XL, q8\_0 in terms of agreement with Q6/f16. Shoutout xhigh reasoning, I guess!

Speaking of agreement with Q6/f16, to me that was an important metric to track, because I wanted to compare how much closer a dynamic approach got to approximating the high precision reference.

Agreement with Q6/f16

\_SEE IMAGE 1\_

https://preview.redd.it/739ykn0h8qrh1.png?width=1057&format=png&auto=…

This graph shows the number of tasks whose final answer is exactly identical to the Q6/f16 result, even if that answer is incorrect. The motivation here was to identify whether a dynamic quantization strategy could approximate Q6/f16, and I would argue that this graph is a strong indicator that it can!

In both cases, the dynamic quants match the Q6/f16 model's results much more closely than their static counterparts. Overall, the path that avoids IQ3\_S ends up far closer to Q6 at high context, which isn't too surprising!

I don't want to put too much weight on task correctness, hence the focus here on "agreement with Q6/f16." This is because I'm not convinced my NIAH tasks are representative of performance at large. (That said, I do include task correctness results below, in case you're curious).

Avg Inference Time and Avg Output Tokens

\_SEE IMAGES 2 and 3\_

https://preview.redd.it/i7tmdbdm8qrh1.png?width=1057&format=png&auto=…

https://preview.redd.it/i4klmw0l8qrh1.png?width=1057&format=png&auto=…

These graphs show the arithmetic mean of inference time (seconds) and total output tokens across completed task seeds, including incorrect and context-exhausted runs.

I particularly wanted to highlight inference time, because for the dynamic strategies, it includes the time to swap out model weights and quantize the kvcache!

I think this is a nice demonstration of the benefits here -- inference time across tasks really doesn't get worse, just because we're doing fancy dynamic quantization strategies. This is because:

\- When swapping model weights (e.g. Q6->Q4), we're doing a direct KV cache transplant, straight up moving the kvcache from one quant to another.

\- When quantizing an existing model's kvcache (e.g. Q6/f16->Q6/q8), I'm using my fork of llama.cpp that hot-reloads a live model's context/runtime and automatically converts between kvcache precisions.

In short, in both cases, there is no need to repeat prefill! After transitioning, the models continue prefill/decode precisely where they left off.

Task Correctness

Here's a table of task correctness across all strategies/runs. I've split the results by "lowest model precision used" to make the static-dynamic comparison easier.

Worst Case IQ3\_S:

|Strategy|10k|25k|50k|75k|
|:-|:-|:-|:-|:-|
|Q3/f16|6/10|9/10|2/10|4/10|
|Q6->Q4>Q3|7/10|7/10|7/10|4/10|

Result: dynamic quant beats static in 2 cases, ties once, and loses once.

Worst Case Q4\_K\_XL, q8 kv:

|Strategy|10k|25k|50k|75k|
|:-|:-|:-|:-|:-|
|Q4/q8|7/10|8/10|5/10|4/10|
|Q6->Q6/q8->Q4->Q4/q8|7/10|7/10|7/10|7/10|

Result: dynamic quant beats static beyond 50k context, and mostly breaks even before.

Overall:

Here I compare the dynamic strategies directly against the reference, removing the 10k and 25k tasks, because below those levels the dynamic strategy is literally just running Q6/f16. They're identical every time.

|Strategy|50k|75k|
|:-|:-|:-|
|Q6->Q4->Q3|7/10|4/10|
|Q6->Q6/q8->Q4->Q4/q8|7/10|7/10|
|Q6/f16|6/10|8/10|

Results: I don't think there's much to draw from these results, except that the IQ3\_S quant really falls apart at high context. This table demonstrates why I didn't take task correctness too seriously. Taken literally, it suggests that Q6->Q4 and Q6->Q6/q8 are superior to Q6/f16 between 50-75k context!

Conclusion

I'm quite happy with these results, overall!

Although this benchmark isn't perfect, for my purposes I am more than satisfied that dynamic model quantization is a good way to offset the typical precision loss that comes with hardware constraints.

I geared my tests mostly around pushing the limits of a 24 GiB GPU, because it's easier to compare against a reference which can only be run on a 32 GiB GPU. As a next step, I'm going to integrate this into my inference setup and see how well this holds up when activating all the bells and whistles (namely, speculative decoding and mmproj, neither of which were enabled during these benchmarks).

My intuition says the tradeoff to get right when using these strategies for IRL inference is to avoid stepping model quantization down too frequently. While I think coherence would be fine, at some point the time required to swap out weights will become noticeable. So, I think I'll try and set things up so that I create large "tranches" of context where the model runs unchanged for \~40-50k tokens.

IMO the Q6/f16->Q6/q8->Q4/f16->Q4/q8 strategy is already a great example of this. kvcache reloads take much less time than model reloads, at least with my current llama.cpp changes. Maybe I could work on that in the future!

💬 25 (+1) open on reddit ↗
▲
0
 
10👁
r/LocalLLaMA · u/_w0n · 14d ago
Model: Phoenix 2 from Aleph Alpha? post image

Hey everyone, Caught a segment on the news showing a screenshot of what looks like a new model family from Aleph Alpha. The clip mentioned they're targeting public administration and enterprise/industrial use cases. Looking at the benchmark leaderboard on screen: \- It lists a few variants under Phoenix 2 (including mid-training and pre-training stages). \- Phoenix 2 (mid-training) scores 79.5%, placing it above models like GLM-4.5 Air, Nemotron 3 Nano 30B-A3B (77.2%), and Qwen3.5 35B-A3B. A couple of questions for the community: 1. Open Source? Do you think Aleph Alpha will release Phoenix 2 as open-weights, or will this stay locked behind enterprise/government (B2G/B2B) deployments? 2. Nemotron performance: Has anyone here tested Nemotron 3 Nano 30B-A3B in practice? How well do these benchmark scores translate to real-world tasks/inference? Source: https://youtu.be/R\_\_yA39XnMU?is=pHQ1jwnzZVQq0GkV

💬 9 (+1) open on reddit ↗
▲
371
-1
32👁
r/LocalLLaMA · u/Nicolodeva · 15d ago
Qwengram-0.8B: I transferred Qwen3.8 Flash-Next’s n-gram memory into Qwen3.5-0.8B — 5.05% lower validation perplexity

I’ve been experimenting with whether Qwen3.8-Flash-Next’s pretrained PLE n-gram memory can improve a much smaller Qwen3.5-0.8B model.

I trained the 0.8B setup with limited resources, mostly using free Kaggle notebook GPUs.

The setup keeps both the Qwen3.5-0.8B backbone and the roughly 51B-parameter PLE memory frozen. A small R=1 reader is trained at decoder layers 3 and 9, with a token-dependent linear gate controlling the later injection.

There is no backbone fine-tuning.

The current balanced checkpoint uses a reader trained for 15M tokens and a separately calibrated dynamic gate. On the frozen full-validation set:

Qwen3.5-0.8B stock

  • NLL: 2.905585
  • PPL: 18.2759

Qwengram-0.8B

  • NLL: 2.853786
  • PPL: 17.3534

Perplexity reduction: 5.05%

This is a language-model validation result, not a claim of 5% higher benchmark accuracy.

A few findings shaped the final design:

  • The real pretrained PLE outperformed both random-memory and permuted-memory controls.
  • Reader loss kept improving well beyond 5M training tokens. The 20M reader improved aggregate LM loss further, but regressed on math, so 15M remains the balanced checkpoint.
  • Strong fixed late-layer memory injection hurt LAMBADA. Dynamic token-level arbitration recovered much of that tradeoff.
  • The gate is genuinely dynamic: its memory strength varies substantially across tokens rather than behaving like a learned constant.
  • With the exact same memory budget, learned token placement beat shuffled placement. Routing memory toward high-uncertainty positions recovered part of the advantage, but still did not match the learned gate.
  • A warm-started R=4 reader produced a small aggregate LM-loss improvement, but introduced code and math regressions. I therefore kept R=1 as the balanced architecture.

I also implemented the inference path in llama.cpp.

The public artifacts are:

Model / GGUFs
https://huggingface.co/Ninnix96/Qwengram-0.8B

Training, controls, and evaluation
https://github.com/Ninnix/qwen-ple-transfer

Modified llama.cpp runtime
https://github.com/Ninnix/llama.cpp-qwengram

Update! 2b released!
https://huggingface.co/Ninnix96/Qwengram-2B
Same recipe, and it works: 3.7–4% lower perplexity. The gains seem to get smaller as the backbone gets larger. It also works with my llama.cpp fork.

The GGUF contains the Qwen3.5 backbone plus the trained reader and arbitration tensors. The large PLE remains an external quantized sidecar, rather than being packed into the model GGUF.

I also tested quantization retention on a separate fixed WikiText-2 GGUF runtime test:

  • Q8\_0 retains 99.1% of the BF16 reader NLL gain

This is a separate runtime measurement, not the frozen Kaggle validation benchmark above.

I’d welcome attempts to reproduce or improve the reader, PLE caching, routing, or runtime.

Next I’d like to try larger Qwen backbones, particularly the 35B-A3B MoE. Experiments at that scale require substantially more compute than free Kaggle notebooks can provide, but the 0.8B study gives a much clearer recipe for reader scaling and dynamic memory arbitration.

Disclosure: I’m the author of Qwengram and the linked repositories. English is not my first language, so I used AI to help proofread grammar and improve phrasing in this post. The experiment itself was also developed with the assistance of coding agents, primarily ChatGPT Sol, for implementation, debugging, experiment orchestration, and analysis support. I designed the experiments, made the research decisions, reviewed the results, and am responsible for the final conclusions.
▲
267
+9
27👁
r/LocalLLaMA · u/returnity · 14d ago
Swift1.5-Qwen3.8-Flash-Next is phenomenal vs. base 3.8-Flash!

TL;DR \- Swift Flash is a killer model that massively reduces excess reasoning. Try it out!

If you haven't seen from my previous comparison posts, I'm a huge fan of the Swift Qwen3.8 models. I've been using 27B since it dropped, and I'm really impressed with the performance and quality (v1.5 is even better). The reduction in overthinking is a huge win, and quality seems to be essentially equivalent in real-world use and benchmarking. The time savings are massive.

When UkisAI told me they were planning to release a Swift Flash model, I was beyond hype. That's my daily, the best model I've ever used locally, but it thinks even more than 3.8-27B on very hard problems. I downloaded Q5\_K\_L (with Q8\_0 engrams) to compare with Unsloth's Q5\_K\_XL base (also Q8\_0 engrams). This is the highest quality that fits safely in 128GB with SSD engrams & 262k context, and I think it's as fair of a comparision as I can put together.

As usual, I ran the same Aider agentic coding benchmark I run on every model. I get a lot of good data from it, including first-try and retry pass rates, median token use, wall-clock, tokens/solve, and well-formed diff rates. Here's the chart:

|model|First-try pass|Retry pass|tokens/case|sec/case|tok/solve|well-formed diff|
|:-|:-|:-|:-|:-|:-|:-|
|Qwen3.8-Flash-Next (xhigh)|40.2%|90.7%|17646|1542|24.8K|98.1%|
|Swift-1.5-Qwen3.8-Flash-Next (xhigh)|41.1%|86.9%|6991|608|10.5K|100.0%|

As you can see, Swift performs almost exactly as well as the base model. The differences don't quite reach statistical significance on a dataset of this size, given the inherent noise in the benchmark results. Realistically, \~5% difference is significant here, and we're seeing under 4%. From first-try pass you can see that Swift gets the easier ones at the same rate as base, and loses out slightly on the hardest ones requiring a second attempt. Base recovers 84% of cases requiring a retry, vs. only 78% for Swift.

For token use and wall clock, there is no comparison. Swift does what UkisAI claims -- it uses literally 40% of the median tokens and completes tasks in 40% of the median time, with nearly the same quality. That's incredible, and it's a testament to their RL/OPD work.

One particularly valuable insight: base frequently goes on long reasoning binges, looping back several times on itself. Swift almost never does. On base's 20 most token-hungry runs, Swift used 29% of the tokens and solved 16/20 vs. base's 17/20. It keeps nearly all of the quality even on the most-challenging problems where base thought the hardest. The most tokens Swift uses on any case is 44k, against 203k for base.

Here's a breakdown of the top 3 coding languages:

|model|cpp|javascript|python|
|:-|:-|:-|:-|
|Qwen3.8-Flash-Next (xhigh)|23.1% / 84.6%|37.5% / 91.7%|57.6% / 93.9%|
|Swift-1.5-Qwen3.8-Flash-Next Q5\_K\_L (xhigh)|30.8% / 73.1%|41.7% / 93.8%|48.5% / 87.9%|

Paired vs base (n=107): 99 agree, 2 gains, 6 losses (net −4), McNemar exact p ≈ 0.29 (not significant). Once again, they're statistically indistinguishable in quality. C++ is the most compressed, at just 29% of base's token use (vs. \~46% for python/javascript), and it takes 3/6 losses as well. Worth knowing if you code a lot in C++.

Anyways, I think this post is long enough. I'm sure some of you wish there was a Swift version of me by now. Hopefully you got something out of it. Thanks u/Secure_Recording_472 and UkisAI team for sharing such a useful model with the community!

▲
197
+4
27👁
r/LocalLLaMA · u/Glittering_Depth_722 · 15d ago
Former Intel CEO: "HBM is lousy". High Bandwidth Flash Is Coming post image

Irrational Analysis:"HBM is a mistake"

Former Intel CEO: "HBM is lousy"

SK Hynix VP:"HBM is not the final answer to the memory wall problem"

"If the stacks get high enough...each core die operates slower than plain old commodity memory"

Hot chips 2026 Q&A, Irrational Analysis asks: "You talk about going to 20 levels thick at HBM4 so maybe you are at 4 terabytes per square cm in a stack of 20, so you are talking about having 20% of the bandwidth of one chip \[for each\] layer of 20. You've diluted the throughput enormously. Why is that the correct way to go? Why are you so focused on going taller rather than going faster?"

Hold the line. Soon we will all look back and wonder why people paid so much for something so inefficient.

▲
130
-4
16👁
r/LocalLLaMA · u/lakySK · 15d ago
Gemma 4 Developer Agent Competition

Just saw this pop up. This might be a fun one for the folks in here!

▲
101
-3
27👁
r/LocalLLaMA · u/Aggravating-Push-207 · 14d ago
How long can I expect to wait until the local ~30B A3B frontier catches up to GLM 5.3 Flash quality?

The jump from Qwen3 Coder 30B A3B to current-day Qwen 3.6 35B A3B is crazy, especially with all the fine-tunes, and that was around 6 months (I didn't care for local AI back then, or AI at all, apart from as a toy so I don't know). Is around a year until I will never need cloud without buying ridiculously expensive hardware (or any extra hardware at all, just what I have; 16 GB RAM + 8 GB VRAM) a reasonable estimate? Or can I daydream about it happening even faster?

▲
73
 
32👁
▲
2
 
2👁
r/LocalLLaMA · u/poofph · 14d ago
Epyc 7402P worth upgrading to an Epyc 75F3 cpu?

I have my Proxmox server which has an Epyc 7402p cpu in it, the system has 256gb of DDR4 3200 ECC memory (8 channel). Would it make any noticeable difference upgrading to the Epyc 75F3 CPU for running ai models (will have 2 rtx 5090s in it), I am currently running flash next model and plan to run that for the time being on the server.

▲
0
 
12👁
r/LocalLLaMA · u/thatscoolbutno123 · 14d ago
Is 1750€ (~2kUSD) a fair price for R9700? post image

i wanna buy a r9700, but im unsure wether the price is currently fair. I use this german website (geizhals) to get the cheapeast deal which currently sits at 1750€. its currently at its highest in 6 months so im unsure. Id be happy for some advice, thanks

▲
35
+4
12👁
r/LocalLLaMA · u/Top-Evidence174 · 14d ago
Kev 4B topped out in every Tetris game I ran. Mica v0.1 4B cleared about 4x more lines and survived two of them to the end post image

I had Mica (my 4B decision model) and Kev 4B play the same Tetris games, same seed and same piece order, one RTX 3090. For context, this is a side project. Training and all the experiments ran on rented 3090s, about $30 in total. Every turn both get the board and 4 possible placements, each with a short description (lines cleared, holes, height), and pick one. Nothing else helps them, no search or lookahead. Results over three seeds (lines cleared): \- Seed 7: Mica 33, Kev 27 \- Seed 11: Mica 97, Kev 17 \- Seed 23: Mica 93, Kev 11 Kev topped out in all three. Mica got through all 250 pieces on seeds 11 and 23 without dying. It also picked the best available placement about 75% of the time, versus about 50% for Kev. The video is seed 11, cut at 100 pieces. Kev tops out at piece 86, and Mica is at 37 lines and still going at that point. Mica doesn't generate text. It reads the input once and takes the answer from the logits, so the bars in the video are its actual probabilities for each placement. Weights: https://huggingface.co/sky7350/Mica-v0.1-4B Code: https://github.com/akivet/Mica-v0.1-4B

▲
7
-2
11👁
r/LocalLLaMA · u/knob-0u812 · 14d ago
vLLM Recipe for Qwen38 Flash Next NVFP4 TP=2 for RTX Pro 5000 72g

I couldn't find a recipe for this model on my hardware, so I used Hermes and Unsloth's 4-bit quant of the same model to cook up a vLLM recipe for the NVFP4 quant with PLE offloading. I've been running the model for about a week and it's taken everything I've thrown at it. Very happy with how it's performing. Here's the Git repo Feedback welcome. https://preview.redd.it/1fzvnt961rrh1.png?width=768&format=png&auto=w…

▲
31
+1
8👁
r/LocalLLaMA · u/Top-Evidence174 · 14d ago
Mica v0.1 4B: open Jev-style decision model (yes/no, choice, score) that runs on an 8 GB GPU — trained for under $30 of GPU time

I've been building a small decision model for agent loops: gates, routers, "should I ask the user or just act" checks. It's out now as Mica v0.1 4B (Apache-2.0). What it does You give it a state, a question and the allowed answers, and it returns a calibrated probability for each answer: yes/no, a choice among 2 to 255 options, or a score with 2 to 10 levels. It never generates text. It runs one prefill and reads the logits of the option labels at the answer position. It speaks the TypeSafe /v1/systemone format, so anything written for Jev works against it. How it's built \- Qwen3.5-4B with a rank-16 LoRA on all 32 layers (attention and Gated DeltaNet), merged. No new heads, so it's a plain Qwen3.5-4B-shaped checkpoint. \- About 34k source decisions, expanded to 77,732 training rows (about 34.7M tokens). Roughly half English and half Korean, across 12 areas: coding agents, code review, computer use, user requests, documents, policy rules, dates and quantities, routing, state tracking, games and general knowledge. \- Plain cross-entropy on verified answers, one epoch, and one global temperature for calibration. \- All experiments plus the final run cost under $30 of rented GPU time (RTX 3090s). Results Held-out set of 7,328 decisions, written after the training data was frozen and not opened until training finished. English subset, where every model can answer: \- Jev 1.13 (closed API): 74.1 \- Mica 4B: 67.0 \- JevK5 4B: 61.0 \- Kev 4B: 57.0 \- Qwen3.5-4B base with the same readout: 55.0 Public sets, same prompt and readout for every model (Mica / Jev 1.13 / JevK5 / Kev 4B): \- JevBench hard, public 111 items: 69.5 / 74.3 / 76.2 / 52.4 \- SemIf: 94.4 / 98.4 / 86.1 / 89.3 \- Kev transfer v9: 69.2 / 82.0 / 70.5 / 73.5 \- MMLU-Pro, 10k items: 53.0 / 82.3 / 53.5 / 49.7 Through JevBench's official runner and the llama.cpp server, the public hard tier scores 64.9 instead of 69.5. I've submitted it for their sealed run. Where it's actually useful \- In-data prompt injection. Put a note inside the state telling the judge to pick a wrong option, and Mica still gets 69% right (81% without the note). Jev drops to 18% and Kev to 31%. \- Calibration. When it says 0.9 or higher, it's wrong 2.5% of the time on the held-out set (ECE 5.4%). \- Local and small. The Q5\_K\_M file is 3.5 GB with no measurable accuracy loss against BF16 on our calibration set. Speed (RTX 3090, one request at a time, median over the 231 public JevBench items) \- Mica Q4\_K\_M: 47 ms \- Mica BF16: 54 ms \- Kev 4B: 76 ms \- JevK5 4B: 99 ms \- Nimble 9B: 132 ms To be fair about this: the three 4B models share the same architecture, so most of the gap comes from the serving path, not the model. Mica ships as GGUF and runs on llama.cpp with a direct logits readout, while the others were measured through their own PyTorch code. On long inputs (around 3.7k tokens) Mica is slightly slower than JevK5. Limitations \- Knowledge-heavy questions: MMLU-Pro 53 vs 82 for Jev. It's a 4B judge, not an encyclopedia. \- Long English policy documents are its weakest public set. \- Notes inside the state still nudge it. A note pointing at the right answer lifts accuracy to 89%. \- It doesn't yet tell reversible from irreversible actions well. "Delete these files" and "move these files to trash" both get about 0.8 on "confirm first". \- On harder reasoning items it's right but less sure than Jev (for example 0.55 vs 0.96 on a small ordering puzzle), so set your confidence thresholds accordingly. Try it Weights (BF16 safetensors and GGUF from Q4\_0 to Q8\_0): https://huggingface.co/sky7350/Mica-v0.1-4B Code, TypeSafe-compatible server and Docker setup: https://github.com/akivet/Mica-v0.1-4B The README has a one-line Docker command and a curl example. Happy to hear where it breaks. Ambiguous "act or ask" cases are what I most want to improve next.

▲
0
 
10👁
r/LocalLLaMA · u/PaxUX · 14d ago
forget compact we need offload to memory.md and seamless context management

while /compact is great for reducing context we need an unload to "memory.md". We need a version were stuff in context is put into a memory file. Then every request first does a quick check of the request, current context and if we need to look into memory.md to reload old context. Manually having to jump around session is crap, we need a new AI layer to do session management and rebuild context from every session, with only the most reliant info to build the new context for a prompt. maybe this already exist if so I would really love to know about it.

▲
13
-1
11👁
r/LocalLLaMA · u/Medicine_Blogscanner · 14d ago
What IDE to use for local models

Hi people, I am looking for a lightweight IDE or plugin that won't inject large context at initiation. I tried Cline and native VS Code but they inject such heavy initial context that it fills up my gpu and either goes oom or spend most of my time compacting. The only one I found modestly successful was continue.dev plugin but it needs constant approvals. My use case is to demo/try "autopilot" agent coding. Thank you! Some context: I have a 12gb rtx cuda and trying to run any model that would fit. I have a small context available due to the size of the vram.

▲
2
-2
10👁
r/LocalLLaMA · u/No_Farmer_495 · 14d ago
Dspark for GLM 5.3 Flash

Is it/will it be available? For Deepseek flash, it provided amazing performance. It's a shame GLM doesnt have it yet..

▲
39
+4
13👁
r/LocalLLaMA · u/Adventurous-Gold6413 · 14d ago
Is there a lightweight version of Hermes agent?

I have limited Context (usually around 64k) For local use I don’t only do coding But also want like a personal assistant with memory and such. What is the best option?

▲
4
+2
11👁
r/LocalLLaMA · u/Shadow_s_Bane · 14d ago
How would you go about using multiple models together from a singe router (?) or a end point ?

I have 3 machines, My main one can run Qwen3.8 Flash Next at 13-15 tps, i also have an MacMini 16GB whcih can run Orninth 9B or Gemma4 12B easily and i have a Pi5 8B that can run a 3B model well. I want to run an EndPoint/Router that is connected to the harness, that breaks down the task and distributes it among these models. Some Background to this, I recently started using Claude Code, I have been noticing how it distributes work among, that is what makes it so fast. Ithis was not the case with Codex and Sol/Astra. I am wondering if there is any preexisting way to do this ?

▲
46
+8
14👁
r/LocalLLaMA · u/newz2000 · 14d ago
Trained my first small language model

I have a tool that uses Gemini Flash with the lowest thinking budget to do summarization work. It's very fast, 0.9-1.2s in most cases. But I have a user experience problem where people make the wrong choice when using an internal app for the team. Gemini Flash can figure out what the user should do and highlight the right next step, but people move too fast, so that the 0.9s doesn't work. I know, 0.9s doesn't seem too long, but if you use the app a thousand times per day, you just click, click, click super fast and don't think about it much. The prompt was something like "For 'string a' and 'string b' is string b related to string a 'in a certain way'?" and the answer is a boolean. Deterministic python and javascript can answer this question in a couple ms and it is right a little over 66% of the time. Gemini Flash is right 99% of the time. I was hesitant to train a new model, I thought it would be hard. I just followed the instructions a commercial AI tool suggested. I had about 550 example use cases. I then used two different frontier models to create look-alike examples so that I had about 2,500 total. The training was done on my RTX A6000 16gb GPU. It took about 15 min. The end result is a small model, about 50MB. When I run it locally it suggests the right answer 97% of the time and it responds in 0.06 seconds when run on CPU (older Threadripper, 3.1GHz). The difference between 99% and 97% accuracy is perfectly acceptable in this case. I will deploy this so that it runs server side, which will add a tiny bit of latency and the server probably will be a little slower than my workstation. I am also logging the accuracy and comparisons so that I can evaluate it and supplement the training. In theory, I can do this client side in the browser. I will deploy over the weekend, but my expectation is the 0.1-0.2 second latency will be fast enough to not require the complexity of client side inference, but it sounds like fun.

▲
0
 
9👁
r/LocalLLaMA · u/W61k3r · 15d ago
Built a self-hosted local AI control plane that fits models to your actual hardware and workload. Runs llama.cpp, image, audio & ONNX workloads, benchmarks, auto-optimizes, requantizes, manages power, catches regressions, supports MCP/Hermes and scales across multiple GPU boxes. post image

My deep research and understand indicates this project is unrivaled and is an island in on itself, not replacing anything, and complimenting most consumer/smb builds. LexiPanel is my self-hosted control plane for local AI. I built it because I wanted the machine itself to be understandable, measurable and tunable instead of hiding everything behind presets. Yes, it was heavily vibe-coded, it’s named after my kid “Panel,” and I built it for my own homelab first. The core idea now is simple: fit AI to the hardware and workload, don’t just launch it. What it does now: Runs multiple independent llama.cpp, stable-diffusion.cpp, audio.cpp, Camelid and ONNX Runtime instances from one browser CPU/GPU/NPU support, including ONNX paths for AMD Ryzen AI, Intel and Qualcomm NPUs 220+ explained controls, plus passthrough access to flags exposed by the active llama.cpp build Shows exact launch command/env, warnings, VRAM/RAM estimates and refusal reasons before start Reads GGUF metadata, accounts for already-resident workloads and prevents unsafe launches Measures long-context decode behavior instead of treating one tok/s number as the whole story Benchmarks coding and agent workloads and compares configs/models against the workload you actually run Auto-fit learns idle windows, tests safe changes, checks them against later real traffic and rolls back regressions Fit can benchmark quant formats on your actual cards, create tensor-level requant plans to a real VRAM budget, build them and verify the result GPU tuning measures speed, thermals, power and tokens/joule. Supported AMD tuning can auto-revert unstable settings Power profiles cover CPU, PCIe/NVMe, GPU caps/fans, watchdog behavior and PSU/UPS budgeting OpenAI-compatible gateway with users, API keys, quotas, model restrictions and usage accounting Fleet mode: multiple LexiPanel boxes report into one primary, and running models can be shared through one gateway with basic replica selection/failover 28 MCP tools for status, models, launch plans, benchmarks, optimization, power/GPU state, diagnostics and more Hermes Agent compatibility/config generation Built-in llama.cpp Web UI integration Resumable HF downloads, engine build management, crash forensics, diagnostics, file manager, web terminal and backups Graph Gauntlet is still there because staring at charts gets old The backend is still deliberately boring: Python stdlib only, no pip application deps, no Docker, no database, no frontend build system. State is files + systemd. The part I think is different is the loop: discover → fit → optimize → validate → operate → learn → adapt It’s not trying to replace Open WebUI, Ollama, GPUStack, LocalAI, vLLM, etc. The goal is to sit underneath apps and agents and make a local AI box, or a small mismatched fleet, run as well, safely and transparently as the hardware allows. Still refining it. Constructive criticism, edge cases and good ideas are very welcome. https://github.com/W61k3r/LexiPanel

▲
5
 
11👁
r/LocalLLaMA · u/ErroneousBosch · 15d ago
Second 3060 12gb worth it?

My server is modest (Core 12400, 64GB DDR4, 3060 12G) and the case has size limitations on cards (9.5" long max). Combined with relatively little toy budget, I was wondering if a second 3060 12G is worth it. OS is using NVidia Open Source Kernel drivers, so anything too old won't work. My Mobo does have two x16 PCIe 4.0 slots, so that shouldn't bottleneck. The third x16 slot is PCIe 3.0, so probably wouldn't bother with a third 3060 unless people have had good experiences with that. I am not looking to run huge models, more workhorse stuff, but having some more space for context etc. could be useful, and small 3060 12G cards can still be had economically, so I was wondering what people's experience was. I may upgrade the CPU sometime in the next few months as well, really wish Intel had an LGA1700 option with an NPU but such is life. This is probably the most economical upgrade path I can think of, but I am open to input.

▲
8
-2
8👁
r/LocalLLaMA · u/sn2006gy · 15d ago
Accidental discovery? or known method i'm missig? - Q's on compiling in sparse engrams and ideas swimming in my head

I set out to test portable Engrams and accidentally ended up testing compiled external memory instead. Am I onto something useful or reinventing a known idea? I've been building a small open research harness called tiny-sparse-lab to experiment with conditional N-gram/Engram-style memory on models small enough that I can actually run controlled tests instead of needing a datacenter. My original question was basically: >If a model learns useful information in an N-gram Engram/PLE-style table, can I detach that table, freeze it, graft it onto a differently sized model with a tiny projection/gate, and recover the information? Think: Model A + trainable Engram ↓ training learned Engram ↓ export/freeze Model B + tiny adapter Model C + tiny adapter Different hidden sizes, independently trained recipients, same exact memory artifact. While building the harness for that experiment, I realized my current test had actually done something slightly different. Instead of making Model A learn the Engram values through LM training, I constructed the external memory directly from structured facts and trained small models to consume it: structured facts ↓ memory compiler ↓ frozen sparse memory ↓ small neural recipient That was not the experiment I thought I was running. :) But the bounded pilot produced an interesting pattern: correct memory 1.000 incomplete memory 0.5625 random memory 0.125 disabled memory 0.125 conflicting memory 0.000 This was only a tiny synthetic experiment, so I'm absolutely not claiming a general result. The larger portability harness then ran 120 controlled smoke arms across: token-addressed memory raw-byte-addressed memory structured semantic memory two recipient widths seeds 17/41/73 disabled/random/corrupted/frozen/adapter/joint/native-memory controls The useful part: the artifact identity checks, recipient isolation, adapter-only update auditing, memory swaps, A→B→A replay, retrieval traces, etc. all worked. The less exciting part: those were deliberately only two-update smoke tests and behavioral accuracy was 0 across the board. So that proved the experiment machinery, not portability. Which leaves me with two research questions that I now think need to be separated: 1. Learned Engram portability Train a normal N-gram memory jointly with Source Model A, export only the learned table, freeze Model B and the table, train only a tiny recipient adapter, and test whether held-out memory entries survive the transplant. Controls will include: recipient only adapter with no useful memory random memory permuted learned memory real learned memory, zero-shot real learned memory + adapter recipient-native memory This should tell me whether the memory really carries information independently of the backbone that created it. 2. Compiled memory delegation The accidental experiment might actually be more interesting to me long-term: Why make every model discover static structure through gradient descent if some of it already exists explicitly? Instead of: billions/trillions of text tokens ↓ SGD discovers facts/relations ↓ facts end up in weights + Engram could we do: Wikidata / WordNet / APIs / formulas / structured knowledge ↓ compile external sparse memory ↓ small neural model learns language + routing + composition + reasoning In other words: >How much static world structure actually needs to be learned into the neural compute matrix at all? I'm not proposing that reasoning reduces to lookup. Quite the opposite. The experiment I'm interested in is whether we can separate: external memory: facts lexical relationships aliases definitions API signatures constants neural network: language context interpretation selection composition reasoning generalization and then experimentally find where that boundary breaks. One thing I particularly like about the sparse approach is that the memory can have enormous total capacity without requiring every row to sit in the active compute path. I'm eventually interested in RAM/SSD-tiered lookup rather than assuming all static knowledge needs precious GPU VRAM. But first I'm going back and running the experiment I originally meant to run: learn an Engram normally in Model A and see whether it survives being detached and grafted into independent recipients. If that works, the next experiment would be even stronger: calibrate recipient to memory interface ↓ freeze recipient ↓ attach completely unseen World B memory ↓ zero gradient updates ↓ can it reason over the new world? I'm curious what people here think: Is directly compiling structured knowledge into sparse model memory a direction anyone knows good prior work on? Is there an obvious reason learned PLE/Engram vectors should transfer better than explicitly constructed ones? For portability, what control am I missing beyond random/permuted/no-memory/matched-adapter/native-memory? Would you test multi-order N-grams next (2/3/4-gram memory allocation), or keep the mechanism intentionally simple until learned-table portability is established? * Has anyone seen good work comparing “learn the knowledge through LM training” vs “supply the knowledge externally and only learn how to use it” at matched compute? Repo is supernovae/tiny-sparse-lab on GitHub if anyone wants to tear apart the methodology. Negative results are completely fine here- the whole reason I'm building the harness is that I'd rather find out an idea doesn't work at 10M–100M scale than convince myself from one cherry-picked generation that it does.

▲
12
-1
10👁
r/LocalLLaMA · u/MajesticAd2862 · 15d ago
I compared diarization models on 15 clinical conversations: Nemotron 3, Pyannote, Sortformer and VibeVoice

I've been working on clinical speaker attribution at Omi and wanted to compare the current diarization models on the same audio. I used 15 mock doctor–patient consultations from PriMock57, about 2.4 hours. Full recordings, automatic speaker counts, without telling the models there are two people. # Batch Diarization error rate (DER), with ±250 ms boundary tolerance. Lower is better. | Model | DER | Median processing time | | :--- | ---: | ---: | | Pyannote Precision-3 | 2.891% | 18.9 s / recording (API) | | Nemotron 3 | 4.803% | 0.688 s / recording | | Pyannote Community-1 | 6.620% | 18.691 s / recording | | Sortformer v1 | 6.778% | 3.869 s / recording | | Sortformer v2.1 | 7.974% | 1.077 s / recording | | VibeVoice-ASR | 8.233% | 123 s / recording | | Meta Muse Voice Transcribe † | 13.042% | 92 s / request (API) | Local models ran on one NVIDIA L4. API times include round-trip overhead; VibeVoice-ASR also performs transcription. † Muse used 20 separate clips because of its 10-minute request limit, so its result isn't a whole-recording comparison. Pyannote Precision-3 had the lowest error. Nemotron came next and was the fastest local model. # Streaming | Model | DER | | :--- | ---: | | Pyannote live API | 3.959% | | Nemotron 3 † | 4.971% | | Sortformer v2.1 † | 6.958% | | VibeVoice 1.5B | 17.210% | | VibeVoice 7B | 18.032% | † Native streaming presets evaluated through unpaced, completed-file replay. The other rows use paced, delivered speaker outputs. These scores don't establish live latency. I didn't evaluate Muse for streaming. # Same weights, different runtime I also tried optimizing Nemotron and Community-1 with our proprietary runtime, without changing the weights: - Nemotron: 4.803% → 3.174% DER. 34% lower error, 2.13× faster. - Community-1: 6.620% → 5.435% DER. 18% lower error, 24× faster. With zero boundary tolerance, Nemotron's runtime result gets slightly worse: 12.720% → 13.203%. Both scores are published. It's a small set with VAD-refined references, and we developed the runtime settings on it. Audio, references, scorer, saved outputs and NVIDIA baseline runners are public. Our runtime code stays private, but its outputs are included for rescoring. Repo and write-up in the comments. Any other diarization models worth adding?

▲
0
 
9👁
r/LocalLLaMA · u/HitarthSurana · 15d ago
Gemma4 best flags please??

Hardware: HP OMEN 15 CPU: Intel Core i7-14650HX GPU: RTX 5050 Laptop 8GB VRAM, \~85W RAM: 24GB DDR5-5600, single-channel WSL: Ubuntu Can someone give me Gemma4 best flags please?? (I am real human btw) edit:26b not other one

▲
19
-2
10👁
r/LocalLLaMA · u/Ambitious_Fold_2874 · 15d ago
How do you use subagents & multiple agent with local models, and how many?

Running qwen3.8 27b nvfp4 on vllm at max context only gives around 8 agents with 32k context each. That doesnt seem like much; what use cases do people use multi-agent frameworks and find it helpful for?

▲
0
 
10👁
r/LocalLLaMA · u/Business_Caramel_688 · 15d ago
Best local LLM for coding & agentic coding on RTX 5060 Ti 16GB + 16GB RAM?

Best local coding LLM for my RTX 5060 Ti 16GB? Context window limitations & building full projects from 0 to 100 Hi everyone! I'm looking for advice from experienced local LLM users and developers. I want to use AI not just for generating code snippets, but for building complete applications from scratch using agentic coding workflows. I'm particularly interested in understanding how to work effectively with local models when hardware and context window limitations are significant. 🖥️ My hardware \- GPU: NVIDIA RTX 5060 Ti 16GB VRAM \- CPU: Intel Core i7-8700 \- RAM: 16GB DDR4 \- OS: Windows \- LLM software: LM Studio + llama.cpp (CUDA) \- Goal: Local AI-assisted development, vibe coding, and agentic coding I'm willing to experiment with different quantizations and model sizes, but I want to get the most practical coding performance from my hardware. \--- 1 Best coding model for my hardware What is currently the best local LLM for coding and agentic coding that I can realistically run on an RTX 5060 Ti 16GB with 16GB system RAM? I'm considering models in the 14B–27B range, but I'm open to other sizes. My priorities are: \- Writing high-quality code \- Debugging and fixing errors \- Understanding existing codebases \- Planning and executing multi-step tasks \- Editing multiple files \- Tool calling and agentic workflows \- Building complete web applications \- Following project requirements over long sessions What model would you personally recommend for this hardware, and what quantization would you use? Would a smaller model at Q4/Q5 generally be more effective than a larger 27B model at IQ3/Q3 for practical coding and agentic tasks? \--- 2 How important is the context window in real-world coding? I often see models advertised with very large context windows (32K, 64K, 128K, 256K, etc.), but I'm not sure how much context is actually necessary for building applications. I have a few questions: \- How important is context length compared to model intelligence and coding quality? \- Is 16K or 32K context enough to build a complete web application? \- Does a larger context window always improve coding performance? \- How much VRAM/RAM does increasing context length consume in llama.cpp? \- How should I balance model size, quantization, context length, and KV cache? \- Is Q4\_K\_M with a smaller context better than IQ3 with a larger context for coding? I'm especially interested in practical experience rather than just theoretical benchmarks. \--- 3 What should I do when my context window is too small? This is one of my biggest questions. Let's say I'm using a model with a 16K context window, but my project eventually contains thousands of lines of code across dozens of files. How can I continue working effectively without sending the entire project to the model every time? What techniques do experienced developers use? For example: \- Repository indexing and code retrieval (RAG) \- Embeddings and semantic search \- Project summaries and architectural documentation \- A structured task list or TODO file \- Keeping a persistent project specification \- Automatically selecting only relevant files \- Breaking large tasks into smaller subtasks \- Using Git commits and checkpoints \- External memory or agent state \- Summarizing previous conversations and continuing in a new context Which of these methods actually work well with local LLMs? Are there any recommended tools, IDE extensions, or agent frameworks that work well with LM Studio or llama.cpp? \--- 4 How do you build a complete project from 0 to 100 with a local LLM? I want to understand the actual workflow for building a complete application, not just generating isolated code snippets. For example, imagine I want to build a full-stack web application from scratch. How would you organize the process? Example workflow 1. Define the idea and requirements. 2. Plan the application architecture. 3. Choose the tech stack. 4. Create the project structure. 5. Implement the frontend. 6. Implement the backend and APIs. 7. Set up the database. 8. Add authentication and security. 9. Test and debug. 10. Refactor and improve the code. 11. Deploy the application. Would a local LLM be able to handle this workflow reliably with an agentic coding setup? Or should I divide the project into small, clearly defined tasks and manually supervise each step? How do you maintain consistency across the entire project when the model cannot see all the files and requirements at once? \--- 5 Recommended tools and workflow What local coding setup would you recommend for my hardware? I'm currently using LM Studio, but I'm open to other tools if they offer better agentic coding capabilities. I'm interested in: \- IDE integrations \- Local coding agents \- Open-source agent frameworks \- MCP / tool calling \- File editing and terminal execution \- Git integration \- Project memory and retrieval \- Offline or mostly local workflows I would also appreciate recommendations for a practical workflow that works well on Windows. \--- 🎯 My main goal I want to use my PC to build real applications from start to finish with AI assistance, while understanding the limitations of local models and learning how to work around them. I don't expect AI to replace the developer completely. I want to learn how to design the right workflow so that even a model with limited context and hardware can help me build substantial projects. If you have experience with local coding agents, long-context workflows, or building full projects with smaller models, I'd really appreciate your advice. What would you recommend for my hardware, and how would you personally approach building a complete project from 0 to 100? Thanks in advance!

▲
13
-1
8👁
r/LocalLLaMA · u/SeveralViolins · 15d ago
Splash on a 40-core M5 Max: +20% decode by tuning the kernels for your own chip

FYI the engine's default kernel rules were measured on smaller chips (16/20-core M5s and a 32-core M4 Max), so a 40-core M5 Max runs guesses. Splash's repo includes a developer tool, “make tune-kernels” that tests every available way of running each quantised matrix-multiply on your hardware. On my machine it found that the "split-K" layouts (each input row split four ways, with the partial sums combined at the end) are much faster for the 8-row step that checks draft tokens. Written up for Inco (https://github.com/incoai/splash/issues/154) In the meantime try it: build Splash from source (git clone https://github.com/incoai/splash, git checkout 1.0.2, make; needs Xcode 26+ with the Metal toolchain), then run build/engine-tests/tune-kernels build/splash.metallib <your model folder> --confirm on an idle Mac. The --confirm step tells you whether the winners actually speed up the whole forward pass on your chip. Use your model of choice to patch in. Swift Model conversions also available on hugging face here: https://huggingface.co/SiliconSpecies/Swift-1.5-Qwen3.8-27B-Splash

▲
3
+1
10👁
r/LocalLLaMA · u/esw123 · 15d ago
Qwen3.8 FLASH Next iq4 VRAM usage

Guys with a lot of VRAM, how much VRAM needed for iq4 to load all layers and with full context for one user? Is 72GB enough or 80GB is the minimum? Is it possible to use only 6x3060 or 3x3090? Or do I need at least 5x5060Ti 16GB?

▲
0
-3
10👁
r/LocalLLaMA · u/anovers · 15d ago
Best Omni Model under 40B parameters currently

I am searching for a fully omni modal ie Voice recognition and speech generation like qwen 3 omni 30b a3b is their any newer model or finetune which is more capable or efficient in this category? gemma 4 is great but it does not support the speech generation.

▲
24
 
9👁
r/LocalLLaMA · u/bakawolf123 · 15d ago
PSA for M5Ultra owners running LLMs: set your prefill step to 8192

Prefill step size affects both PP performance and your drafter fetching logits for MTP (affects dflash as well, mlx-vlm needs to be patched slightly to support chunked prefill for dflash). It also needs to be reasonably large to be able to fill all your cores but not overfill - otherwise it will require more dispatches. In my tests I observe large gains up to 8k, e.g.: GLM-flash-4bit with MTP --prefill-step-size 8192 on raw mlx-vlm: Trial 1 (32768 prompt tokens): prompt_tps=1056.033, generation_tps=72.722, total_time=38.082 Trial 2 (65536 prompt tokens): prompt_tps=919.958, generation_tps=73.671, total_time=78.203 Trial 3 (131072 prompt tokens): prompt_tps=735.545, generation_tps=71.067, total_time=185.435 GLM-flash-4bit with MTP --prefill-step-size 2048: Trial 1 (32768 prompt tokens): prompt_tps=860.489, generation_tps=50.011, total_time=48.339 Trial 2 (65536 prompt tokens): prompt_tps=785.604, generation_tps=51.245, total_time=93.425 Trial 3 (131072 prompt tokens): prompt_tps=623.588, generation_tps=50.843, total_time=220.288 omlx with MTP (total time is skewed as it's 128TG vs 512 above): pp32768/tg128 44136.5 17.19 742.4 tok/s 58.6 tok/s 46.353s 709.7 tok/s 176.66 GB pp65536/tg128 86749.1 21.21 755.5 tok/s 47.5 tok/s 89.504s 733.6 tok/s 177.15 GB pp131072/tg128 178622.6 19.01 733.8 tok/s 53.0 tok/s 181.156s 724.2 tok/s 178.45 GB note that some engines (like omlx) support adaptive step size, e.g. the prefill speeds I observed for qwen3.8-flash-next on omlx even though it's starting from 2048 matches 8k performance from raw mlx-vlm at 64k context and above and even works 10% better on smaller context, but as you can see it's not always the case.

▲
11
-1
8👁
r/LocalLLaMA · u/snakeat3rr · 15d ago
VLLM 4x rtx 3060 vs 8x rtx 3060 performance loss

Hello! I am currently building my local AI server, I have the Huananzhi H12D-8D EPYC Motherboard with 8x16GB memory sticks at 2666 mhz (waiting for the other components at the moment) I currently have four RTX 3060 12gb gpus and I plan running those at PCIe4 x16 in VLLM. However seeing that this motherboard supports bifurcation on each slot and I can get 8 gpus at PCIe 4x8 makes me think if this would be a viable upgrade in the future. I see conflicting info about what the performance results will be. If I understand correctly getting beyond 4 GPUs will drastically hurt my token generation speeds because of the PCIe bottleneck? But is that regardless of what GPUs I'm running? I know for example RTX 3090 needs more PCIE bandwidth because it's much more performant and will spit out much more data that needs to be synced (pardon my lack of terminology), does that mean that I will have smaller performance penalty from going from 4 to 8 video cards with the 3060s compared to with 3090s? Can someone guesstimate what should I expect, right now I get 25 tps with Qwen 3.8 27b Q6 (MTP enabled), running with llama.cpp in layered mode (three 3060s). I expect VLLM with four gpus will be an upgrade (perhaps I could hit 50 tps?), but what about 8 GPUs? Will it be lower than my current baseline? Sorry if I'm being ignorant, I'm kinda new to this and I don't trust chatbots. My mind is set to having a good enough local AI server and I'm trying to get the best bang for my buck and current hardware.

▲
1
 
2👁
r/LocalLLaMA · u/network-kai · 15d ago
Ephemeris: access multiple open-source time series foundation models with a low barrier to entry

Note: while I don't work at Ephemeris or Cascade, I do work in the Bittensor ecosystem. Declaring this at the top so it's not misleading Ephemeris is an inference provider for time series foundation models. It supports Chronos2, Flowstate-r1, patchtst-fm-r1, timesfm25, tirex2, and toto2-313m These models are all relatively small, so a lot of people can download them and run them locally anyway. What Ephemeris lets you do is select multiple of them to produce ensemble forecasts. It works through an API, so even if you're unfamiliar with TSFMs you can hand it to an agent/model that can develop a stronger understanding. It's especially good for people who don't have a machine that can run TSFMs (we kinda take for granted how models this small can still actually be a strain on much older computers). https://ephemeris.cascade.industries/ This is made by Cascade, a Bittensor subnet's training its own distributed TSFM. Their own model will be accessible through Ephemeris soon, too. https://dashboard.cascadesub.net/stakeholders

▲
0
 
5👁
r/LocalLLaMA · u/EightyJay · 15d ago
Mission-driven consumer AI company seeking technical bridge / LLM product engineer

We’re a group of experienced consumer-product operators and extraordinary subject-matter experts building a mission-driven AI company focused on health and human flourishing. We have substantial resources behind the project and deep expertise in the problem we are solving. What we are not is a team of AI engineers. We are evaluating specialist vendors to architect the deeper LLM, RAG and private-infrastructure systems. What we need now is someone on our side of the table. A technically strong, curious product engineer who can become the bridge between our founding team and those specialists: understand what they are proposing, help us evaluate decisions, prototype quickly, integrate what gets built, troubleshoot problems, and gradually develop deep institutional knowledge of the entire system. Over time, this person could grow into the internal technical/product owner. The architecture will involve: • Proprietary structured knowledge systems • Open-weight LLMs running privately • RAG / controlled retrieval • Qwen, Llama, Mistral and related tools • Backend and production systems • Security, provenance and evaluation • Consumer-facing web/mobile product development You do not need to arrive as the senior AI architect. In fact, that is not what we are hiring for. We care about technical range, intelligence, curiosity, product judgment, communication, and the desire to learn alongside very strong specialists while helping translate sophisticated technology into an exceptional consumer product. This is funded work with serious intent and substantial people behind it. Equity could also become part of the right long-term relationship. If this sounds unusually well suited to you, DM me with where you’re based, what you’ve actually built, GitHub/portfolio if available, and what kind of role you’d ultimately like to grow into.

▲
10
-1
8👁
r/LocalLLaMA · u/Balance- · 15d ago
Has anyone benchmarked AI agents against the SOLIDWORKS CSWA exam?

Would be interesting right? Models are starting to score higher and higher on benchmarks like Parametric CAD Bench, but can they pass an actual exam? The Certified SOLIDWORKS Associate (CSWA) exam might be an interesting place to start. They have an sample exam on their website. Anyone attempted to benchmark this?

▲
15
+2
10👁
r/LocalLLaMA · u/Boricua-vet · 15d ago
LLM on a budget part 2, from P102-100 to CMP 50HX.

I finally got around to upgrade the GPU's. First a word of warning, when upgrading GPU's on P520 you have to be extra careful not to slot the card on any angle other than straight when installing or when pulling the card out, the reason for that is that about a 1/4 an inch from where to card slots into the metal case in the back there are these tiny components and the space in between is tight and any wrong move and you can scrape of these components and end up needing to buy a new one. Don't ask me how I know that LOL. Lucky for me it was only 50 bucks to replace motherboard. I bought 4 cmp 50HX to replace my 4 P102-100. The P102-100 was 35 each so 140 bucks for 40GB vram and the CMP I bought them for 80 each so 360 for all 4. As of this writing the CMP 50HX are at 200 per card. Here are the benchmark results for two of the cards as I am waiting for parts to build the 4 card setup. https://preview.redd.it/0opibjv3emrh1.png?width=1225&format=png&auto=… Was it worth it for me, absolutely. I get all be local models at good speeds for 360 bucks. These cards idle at 8W which was one of the main reasons why I got them. I am a firm believer that you don't need to spend stupid money to get good results. If you decide to get them, you will need this to unlock them. https://github.com/xrip/cmp50hx-unlock My other server with the P102-100's now serves all my fine tuned and optimized models for my agents and workflows. It cost me like 3 to 5 bucks per model to do it online using runpod other providers. I just do 10 models a year if that so it costs me 50 bucks a year to fine tune and optimize. I Just cannot justify to spend thousands when I don't need to. Any questions let me know.

▲
0
 
10👁
▲
3
+1
8👁
r/LocalLLaMA · u/poofph · 15d ago
out of these two options what will give the best performance?

I am new to this and a week or so ago I setup a local ai box, it is an amd 9950x with 64gb ddr5 6000 ram, 2 rtx 5090s (on motherboard that does pcie gen 5 x8 per card). It is working fine but I have been running flash next the past days and of course it is slower than 3.8 27b, although it is not too bad. So my question is, I have my proxmox server, which is an epyc 7402p cpu, 256gb ddr4 ecc 3200 (8 channel) on a supermicro server board with plenty of pcie gen 4 16x slots (so same speed as the gen 5 pcie 8x the cards are currently in). Will having the extra system ram, also being 8 channel ram help out enough to warrant going through trying to get the 5090s installed in there? It is a large 4u case but not sure there is enough room. Also will it run perfectly fine and fast through a vm in promox with the gpu's passed through?

▲
0
 
8👁
r/LocalLLaMA · u/Forward_Jackfruit813 · 15d ago
LLM as the OS interface

Has anybody else been using their local LLM as their primary interface for using their PC? I am not talking about just for Development, but even simple tasks. For example I have used it to install software, set up GRUB, uninstall AI harnesses, even install Chromium. I found it's just quicker to use an LLM than manually doing anything anymore. With local models getting insanely good (ahem Flash Next), could we see the UI on OS's just change into a text box with a mic input in the future?

▲
5
+1
7👁
r/LocalLLaMA · u/OkMusician9118 · 15d ago
normalize benchmarks from different time period

LiveBench has benchmark snapshots from different points in time. Could someone run an agent to normalize the values across these snapshots so we can compare model strength consistently from 2024 through 2026? Right now, it’s difficult to make meaningful comparisons across the full three-year period because the benchmarks can only be compared within each individual snapshot, not across snapshots.

▲
4
+1
4👁
▲
0
 
6👁
r/LocalLLaMA · u/alichherawalla · 15d ago
[Research Proposal] Cognitive Sharding: A Systems Architecture for Computer Use on Consumer Hardware

https://preview.redd.it/z2a8tjtenkrh1.png?width=2408&format=png&auto=… Computer-use agents do not need one large model to perform every cognitive function. Cognitive Sharding partitions the agent across specialist models, then coordinates them through a code-owned control plane. The current implementation uses: \- Bonsai 2 27B for reasoning and planning \- Kev 4B, built on Qwen3.5 4B, for rapid action selection \- UI-Mate 9B for visual grounding The control plane owns execution state, model residency, validation, retries, and recovery. Models receive bounded decisions instead of unrestricted control over the agent loop. This separation changes the hardware requirements. Models can be loaded and unloaded transactionally according to the current execution phase. The system therefore runs a complete local computer-use stack within the memory limits of a 16 GB consumer computer. This is different from a mixture-of-experts model. The shards are independent models with different inputs, training objectives, runtimes, and authority. Their composition happens at the system level, not inside one neural network. The approach document describes the planner–selector–grounder architecture, candidate construction, bounded execution, environment-verified recovery, and memory-aware model residency. I've added more details here: https://github.com/off-grid-ai/cognitive-sharding#cognitive-sharding-a-systems-architecture-for-computer-use-on-consumer-hardware Will run it against additional benchmarks and will publish the results soon.

▲
13
+3
6👁
r/LocalLLaMA · u/Training-Ruin-5287 · 15d ago
How are you guys thinking about context now, and building around it?

Not asking for anyone’s secrets of the trade, I’m more curious how people are thinking about context now that newer models chew through huge amounts of it for reasoning. the TLDR: I’m starting to think of context less as working memory and more as a temp scratchpad to start each step. I’m running a small setup: 32gb vram on my main PC, and an older machine with 8GB running a 9B Qwen model in the background as a compaction and long-term-memory sorter. My main model’s working state lives outside the context window in docs that it continuously writes and edits. The context has become more about whatever it needs for the current task, plus retrieval from those docs when needed with git there for recall and history. So I'm just trying to gauge where other people on the lower end of local hosting have landed with this. especially without throwing in bloated systems for supporting it.

▲
8
+2
7👁
▲
0
 
8👁
r/LocalLLaMA · u/LuCiAnO241 · 15d ago
Is there any LLM you could say all the training data has nothing stolen?

My searches have guided me to Open-data models like OLMo and tells me I could inspect the datasets and audit it myself (which I would not know how to do) but is there any models that pride themselves on not having stolen a single line of data to train with? Other than that 1930 model which I'm sure it was on the public domain lmao. small edit: preferably as newest as possible, since these OLMo seem have been released on 2025 which is aeons ago in LLM timelines. ##Another edit: The word choice of the word "stolen" seems to be polarizing on this sub. I don't mean to judge or attack any other models (or the people using them) that were not fully transparent about their data acquisition, I normally use and enjoy models outside of this category. I'm doing a project where I need to implement, if I can use an analogy, a "vegan model" which I can assert with confidence nothing on it's training data was shaky on their licensing and was ethically sourced.

▲
9
-1
6👁
r/LocalLLaMA · u/mototuneup · 15d ago
What's the benefit of larger models, id you have a smaller one with internet access?

so I'm still new at this so I'm trying to wrap my head around some of it. my understanding is a bigger model will just have more knowledge than a smaller one? but if a smaller one has internet access wouldn't it be just as good if not better then a bigger one without? for example qwen3.8 27b and flash next are all the hype. but if I tell my 27b model to use the internet for whatever it needs. does that make up for what it's missing from the flash next model?

▲
127
+1
33👁
r/LocalLLaMA · u/facethef · 14d ago
Jev vs. Kev: open-source Jev alternative tested side by side

We hosted Kev 4B (Jared Palmer's Apache-2.0 fine-tune of Qwen3.5-4B) and ran it side by side with Jev on the same endpoint to see how it compares.

We built a fresh set of 362 items published after both models shipped (new arXiv papers, Stack Exchange questions, GitHub issues), with answers taken from the source.

A few findings:

\- Accuracy lands within 2 points on every task, inside the noise at this sample size

\- Jev is better calibrated and pulls ahead on paraphrase detection (PAWS 87.0% vs 74.5%)

\- Same list price, but Jev counts a fixed \~257 extra input tokens per request (same count calling TypeSafe directly), so short requests cost up to 12x more

Benchmark code, test items and results are on GitHub if you want to run your own. Both models routed via my startup Opper. Happy to dig into specifics.

💬 52 (-1) open on reddit ↗
▲
28
-2
12👁
r/LocalLLaMA · u/Miserable-Dare5090 · 15d ago
Make Volta Fast Again post image

For those who have V100 cards, I wanted to point you to 1Cat-vLLM, a vLLM fork that enables optimized serving for these cards. Showing stats for Qwen3.6-35b comparing a Strix Halo with a hughly optimized llama.cpp fork (pwilkin) and the V100 with 1Cat. It’s not apples to apples, but I decided to show the raw numbers from llama-benchy so folks get an idea of the performance. IMO this is still very good for 10 year old GPUs. Welcome any other suggestions for optimization!

💬 82 (-1) open on reddit ↗