50 posts · 1 sub · RSS
← prev past day next →
day hourdayweekmonthyearall
allr/LocalLLaMA
▲
371
+358
3👁
▲
209
+198
4👁
r/LocalLLaMA · u/ResearchCrafty1804 · 10h ago
Qwen-Image-2.1-Turbo released!

Qwen-Image-2.1-Turbo, create and edit images in just 8 denoising steps! Open weights now available!

Built on Qwen-Image-2.1, Turbo is an accelerated checkpoint on the same 7B visual generation architecture.

Fewer steps does not mean lower quality: it still generates strong 2K images from text, and supports continued creation through natural-language edits, from adding accessories to changing a scene.

Start directly with Diffusers: load QwenImage21Pipeline and the checkpoint’s recommended 8-step sampling schedule is ready to go.

Hugging Face: https://huggingface.co/Qwen/Qwen-Image-2.1-Turbo

💬 21 (+21) open on reddit ↗
▲
265
+176
7👁
▲
159
+157
11👁
r/LocalLLaMA · u/TheOriginalG2 · 22h ago
Qwen3.8-27B: 159 tok/s on R9700, 64 tok/s on Strix Halo

LemonSeed Studio is an iPad editor/IDE with on-device inference on an AMD GPU in a Thunderbolt enclosure. It embeds the unmodified upstream Linux amdgpu + amdkfd driver (mac\_linuxgpu) as a PCIDriverKit extension, and runs LemonSeed Engine (LSE) on the GPU. In the photos: iPad Pro + Sapphire Radeon AI PRO R9700, Qwen3.8-27B Q4 with a Q8 DFlash2 draft, 131k context.

LSE is the same engine on every platform: Linux, macOS and iPadOS. It records the model's forward pass as a graph, fuses ops and generates kernels for the GPU it's on. Where a kernel has several layouts, it measures each on the device and keeps the fastest. Speculative decoding (MTP and DFlash2) picks how many draft tokens to verify from measured acceptance and cost.

Decode, code prompt, Qwen3.8-27B Q4 (baseline / MTP=3 / DFlash2 tok/s):
\- R9700 on iPadOS (LemonSeed Studio): 31.2 / 108.7 / 158.5
\- R9700 on macOS: 32.2 / 111.9 / 159.4
\- R9700 on Linux: 32.0 / 111.0 / 144.6
\- Strix Halo (Radeon 8060S) on Linux: 14.0 / 47.8 / 64.4

Prefill at 4K tokens: 1,636 (iPadOS), 1,634 (macOS), 1,418 (Linux R9700), 517 (Strix Halo). It holds up at 32K: 1,395 / 1,401 / 1,226 / 462.

New in 0.5.8:
\- DFlash2 draft trees on by default: one target pass verifies a tree of draft candidates when that's measured to be faster than a chain
\- Strix Halo (gfx1151) prefill and decode work: fused gate/up GEMM, two query tiles per workgroup in prefill attention
\- Fixed a startup GPU fault
\- Linux archive bundles its own HSA runtime, so you only need the amdgpu driver with /dev/kfd access
\- lse-server models / pull: grab Hugging Face models with their MTP or DFlash2 companions

OpenAI-compatible server, CLI, and a C library (libLSE).

Engine: https://github.com/Geramy/LSE
Studio: https://github.com/Geramy/LemonSeed-Studio

Happy to answer questions, especially about getting an eGPU working on iPadOS.

💬 80 (+78) open on reddit ↗
▲
142
+121
8👁
r/LocalLLaMA · u/pmttyji · 17h ago
[Paper] EngramEdit: Decoupled Knowledge Updates in LLMs through Conditional Memory
Conditional memory architectures such as DeepSeek Engram use input n-grams to look up learned embeddings, expanding the capacity of large language models (LLMs) with limited additional computation. Beyond model scaling, this architecture has demonstrated the potential to decouple factual knowledge storage from general-purpose computation, offering a promising route to updating factual knowledge while keeping the Transformer backbone fixed. Realizing this potential is challenging because different expressions of a fact may activate different n-gram embeddings, while updating shared embeddings can unintentionally change the model's predictions about other facts. We propose EngramEdit for decoupled knowledge updates through conditional memory. EngramEdit first computes target memory representations that make the model predict the updated fact across multiple expressions. It then jointly updates the shared n-gram embeddings to match these targets across expressions and edits, penalizing updates to frequently reused embeddings more strongly to preserve unrelated knowledge. Experiments show that EngramEdit enables independent factual knowledge updates through conditional memory, achieving near-perfect editing success. Revised knowledge is usable across unseen expressions and in multi-hop reasoning, with nearly three times the strongest baseline's accuracy under chain-of-thought (CoT) prompting. Unrelated knowledge and general capabilities are largely preserved even as factual updates accumulate. These findings show that EngramEdit turns conditional memory into an editable knowledge interface, extending its role beyond model scaling to support decoupled knowledge updates.
💬 10 (+10) open on reddit ↗
▲
186
+65
2👁
▲
62
+56
5👁
r/LocalLLaMA · u/pmttyji · 16h ago
[Paper] DLoop: Looped Speculative Decoding
Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length methods decide during decoding how many draft tokens precede a verification, but they raise the speedup only for autoregressive draft models. For a parallel draft model, drafting further requires target-model hidden states for draft tokens that have not been verified. We propose DLoop, a looped form of speculative decoding that adaptively performs multiple drafting stages before verification. DLoop continues drafting while the draft model remains confident and verifies all accumulated draft tokens together. Loop-aware training keeps the draft model reliable in the additional drafting stages by exposing it to its own hidden states for unverified draft tokens. By spending additional draft-model forward passes, DLoop reduces the number of target-model forward passes required for verification. Across diverse speculative decoding methods including EAGLE-3, DFlash, Domino, DSpark, and multi-token prediction modules, DLoop improves the wall-clock speedup by 5 to 41 percent while preserving lossless decoding. Code will be available at this https URL.
The code is currently under internal review and will be released soon. Stay tuned!
💬 16 (+16) open on reddit ↗
▲
76
+55
9👁
r/LocalLLaMA · u/Mrinohk · 23h ago
Qwen 3.6 35B appreciation post

Referring to specifically Unsloth's UD\_Q4\_K\_XL quant because that's going to be a question, and is relevant regardless.

It's old now. It's not great at coding medium sized or even small-ish projects. I wouldn't hand my codebase to it by any means. It hallucinates, like any other model. It's not perfect, by any means.

I don't like the fine tunes; they're almost all coding focused, and lose general assistant capability to a strange degree.

What it is good at is general, broad agentic action.

Set the lights in the living room, and bash into this machine to get a movie going.

Send my grocery list to my phone/watch when I get to walmart.

Remind me to clock in at work each day because it's becoming a problem.

Tell my husband to come here because I'm under the car and I can't drop this thing that I finally got in just the right spot but need a third hand to get this wrench in the correct spot.

Tell my dad about the pets or any one of my projects because I'm showing off your memory to him.

This is all shit that it can do, consistently. Sometimes it'll thrash a bit, but it stays on task and is fast enough that little mistakes are a non-issue.

I recently got a couple of Tesla P100s to power the model, and it's made this model, that I already had going at a pretty good clip on limited hardware, to run at speeds that are genuinely conversational. \~120-140 t/s generation, \~1000PP at 0ctx, \~700 by 13k.

It's more than smart enough to know how to do these tasks and, with the right sampler settings, thinks ridiculously efficiently for them. Actually awesome.

Fucker got a minecraft mod pack installed and running on a machine it wasn't even running on through the prism launcher appimage. Don't worry, it only has ssh towards computers on my tailnet. It's probably fine.

It's bad at holding a persona. My TTS model is trained on Paul Bettany's MCU Jarvis. When the model emits the right phrasing, it feels like magic. It doesn't do that very often. Gemma4 is great at that.

I am actually so scared that if they do release a Qwen4 model in this class (30-40B parmeters, 2-5b active) that it's going to be a coding focused, overthinking, genuinely capable but not at all fast, mess. Gemma4 26b feels like it would be so close if it wasn't dumb as rocks. As it stands, I pay anthropic for that shit coding shit. Maybe when I can run Flash next at good speed (current \~30-40 t/s rn with strata, but at q3. It's ok.) I can kick claude to the curb.

I want a model like what 3.6 35b is, but smarter. Just as fast, but thinks of the little things. Memory updates. I ask it to add something to my calendar, but maybe it also sets up a dedicated notification for the specific time separately. That'd be nice. Not more coding focused. There are so many coding assistant bots. They're great. But the general AI assistant isn't solved yet for the normal person. Maybe it could have better general knowledge, but I think an n-gram table would solve that. It said the P100s used ROCm the other night. Silly bot.

💬 39 (+29) open on reddit ↗
▲
108
+52
2👁
r/LocalLLaMA · u/LegacyRemaster · 6h ago
GLM 5.3 Flash opensource @ the top of Artificial Analysis Cyber Index over Claude. post image

It seems Dario's "too powerful for you users" strategy is paying off: we have two open models at the top of the leaderboard, surpassing every single model from Anthropic.

Open source prevails. Even Mistral Large 4 is better!

💬 46 (+13) open on reddit ↗
▲
53
+49
8👁
▲
72
+48
3👁
r/LocalLLaMA · u/jacek2023 · 12h ago
tencent/Youtu-Parsing-Omni · Hugging Face

Youtu-Parsing-Omni is a compact (5B) omni-modal parsing model. Given a single input — a document page, a natural image, a chart / flowchart, a geometry figure, an audio clip or an audio-visual video — it produces one structured JSON envelope that covers both perception (layout elements, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events, camera motion) and cognition (captions, narratives, reports). The output family is selected by the task prompt (--task in the examples, keys of prompts/youtu_parsing_omni.json).

|Input|--task|modality / subtype|Key contents|
|:-|:-|:-|:-|
|Document page|document|image / document|layout elements with bbox, text / LaTeX / OTSL tables / Markdown charts / Mermaid flowcharts, reading order|
|Natural image|natural_image|image / natural_image|entities and text with bbox, tags, captions, global description|
|Chart|graphics_chart|image / document|one chart element: Markdown table, notes, caption|
|Flowchart|graphics_flowchart|image / document|one flowchart element: Mermaid, caption|
|Geometry figure|graphics_geometric|image / document|one geometric element: points, lines, arcs, shapes, geometric relations and measurements|
|Audio|audio|audio / –|vocal / non-vocal segments with timestamps, speakers, ASR, timbre / scene captions, acoustic events|
|Natural video|natural_video|video / natural_video|temporal segments with visual elements, actions, interactions, camera motion, audio track|
|Text-rich video|textrich_video|video / text_rich_video|segments with OCR + ASR and a Markdown structured_report of the whole video|

Highlights (see the technical report for details):

  • Unified schema – one JSON envelope for seven parsing families, driven by the task prompt.
  • Omni encoder – image, audio and interleaved audio-visual video inputs (frames + audio track) in a single model.
  • Strong results at a small size – state-of-the-art on OmniDocBench v1.6 (96.96 Overall), best open-weight model on OmniParsingBench (75.08 Avg., second only to Gemini-3-Pro), and competitive with specialized models on chemical-structure (ChemOCR) and music-score (PDMX-Synth) recognition (results).
  • Easy to serve – a vLLM plugin, pinned serving settings, task prompts and inference examples are included.
💬 9 (+4) open on reddit ↗
▲
41
+31
7👁
▲
50
+29
2👁
r/LocalLLaMA · u/zyxciss · 5h ago
Qwen 3.8 Flash Next-GSQ-RCO-IQ2_XS at ~21 tok/s on just an RTX 3060 12GB + 16GB DDR4 RAM(No gate pruning, 100% bit-exact) post image

(Don't judge by the screenshot, the cache is cold. It hits 24+ tok/s with a warm cache!)

About two months ago, I made a post here asking whether predicting which MoE experts would be used on the next token could actually help speed up CPU/GPU offloading.

Original post: Tried predicting which MoE experts get used next token to speed up CPU/GPU offload

Well, quick confession first. I actually shelved that project shortly after.

The reason? The speeds I was getting back then were kinda fake. My engine was aggressively pruning experts based on their router weights, basically dropping cold experts to get better performance. Sure, the numbers looked great, but doing that on an already quantized model was hurting output quality and coherence.

Didn't really like that tradeoff, so I abandoned it and never released it.

Fast forward to recently, and Qwen 3.8 Flash Next (125B MoE, 512 experts, top-10 routing) drops.

I downloaded the 68GB GSQ-RCO IQ2_XS build, hoping to run it on my daily driver. That's when I decided to revisit the idea, but this time without cutting corners.

My setup

  • GPU: RTX 3060 12GB
  • RAM: 16GB DDR4, single channel (~19 GB/s bandwidth)
  • OS: CachyOS / Arch Linux
  • Storage: Mid-tier NVMe SSD (~2.1 GB/s read)
  • Software: llama.cpp, built with CUDA

And if you've tried running a 68GB MoE on a 16GB RAM machine with stock llama.cpp, you probably know how painful it gets.

I'm talking 1.4–2.1 tok/s, with over 1,500 major page faults per token in some runs. Linux ends up constantly pulling model data from the SSD because there's simply not enough memory to keep the working set around.

Then engines like Strata started showing up with claims of around 40 tok/s on consumer hardware. Pretty impressive, but there's a catch for people with less RAM. Some of these approaches rely on keeping around 24 GiB of experts pinned in memory using mlock. If you've only got 16GB RAM, you're obviously not doing that. Depending on the setup, you either run into OOM issues or end up with terrible performance.

So I went back to my original idea and started implementing it properly as an optional feature inside llama.cpp:

--moe-direct-io

The goal this time was simple. No dropping experts, no sacrificing output quality, and bit-exact output compared to stock.

The numbers

All tests below were run with MemoryMax=6G using a cgroup.

Model: Qwen 3.8 Flash Next IQ2_XS (68GB)

Hardware: RTX 3060 12GB + 16GB DDR4 RAM

| Engine / mode | Decode speed | Major page faults per token | SSD I/O | Output |
| ----------------------------------------- | ----------------: | --------------------------: | -------------------- | ------------------ |
| Stock llama.cpp (mmap) | 1.41–2.12 tok/s | 1,140–1,565 | 208–312 MB/token | Coherent |
| Our engine (blocking, demand-only) | 0.73 tok/s | 0 | ~206 MB/token | Bit-exact |
| Our engine (--moe-direct-io + prefetch) | 20.14–21.13 tok/s | ~0 (+1 across 32 tokens!) | Sequential streaming | Bit-exact to stock |

The blocking version is actually slower than stock, which makes sense. It's basically waiting on disk reads without doing much to hide the latency.

The prefetching version is where things get interesting.

Once the cache warms up, it sustains 20–21 tok/s, with some runs hitting 24+ tok/s. That's roughly a 10–15x speedup over stock llama.cpp on the same machine.

And no, we're not getting those numbers by dropping experts. The output is bit-exact to stock.

That's the part I'm most excited about, honestly. Being able to run a model this large on a 16GB machine without the usual page-fault nightmare is pretty much what I wanted to achieve with the original project.

What's still rough

It's not all perfect yet. There are a few things we're still working on.

1. Cold starts are noticeably slower

Right now, the slots start empty (-1), so the first request on a new topic can start around 3.5–4.5 tok/s before ramping up to 20+ tok/s as the hot working set settles.

We're working on offline hot-profile seeding so it can start with a useful working set instead of learning everything from scratch.

2. Prompt processing is slow

Feeding a prompt of 512+ tokens can touch a huge number of experts in a short period. That puts a lot of pressure on the 72 slots per layer and causes the prefill stage to struggle.

We're working on micro-batching prompt chunks (-ub 32) to help with this.

3. Speculative decoding gets weird with SSD offloading

We found that standard MTP speculation can actually make things slower. Verifying 2–3 tokens can require loading the combined set of experts needed for those tokens from disk, which eats into the gains.

Right now, confidence-gated speculation (min-p 0.8) or suffix prompt lookup seems more promising for this kind of setup.


Anyway, that's where the project is at right now. Still plenty to improve, especially prefill and cold starts, but getting 20+ tok/s out of this setup without pruning experts is a pretty big deal for me.

Happy to answer questions or get into the io_uring and slot-remapping implementation details if anyone's interested.

(The second half of this post was written with some help from Claude.)

💬 11 (+5) open on reddit ↗
▲
44
+23
3👁
r/LocalLLaMA · u/SarcasticBaka · 12h ago
lightonai/LightOnOCR-3-4B · Hugging Face
💬 13 (+10) open on reddit ↗
▲
36
+19
4👁
r/LocalLLaMA · u/pmttyji · 16h ago
[Paper] Stepped MoE: Segment-Level Routing with Configurable Inference Complexity
Training large language models (LLMs) is resource-intensive, and adapting them for diverse deployment scenarios with varying computational constraints remains challenging. While elastic architectures enable flexible model deployment and sparsely activated models allow input-adaptive computation, existing approaches treat these dimensions independently. Moreover, models catered towards on-device edge inference need to conform to the memory and compute limitations of the serving devices. In this paper, we introduce a unified framework that combines elastic structures with sparsely gated architectures to create models that adapt simultaneously to both deployment constraints and task requirements. Our approach employs a model backbone that conditions on both the context and target efficiency specifications, enabling fine-grained control over the accuracy-efficiency trade-off at inference time. The model learns to activate task-relevant parameters within elastically-nested sub-networks, allowing a single model to span multiple capacity points while maintaining input-adaptive routing. Through experiments we demonstrate that we can create a model that allows the flexibility to use 1,2,3,4 billion parameters while being more accurate than their dense counter-parts (2-5\\% on knowledge-intensive benchmarks) and at par with their static versions while delivering similar latency metrics as dense models. Overall, we save on device disk space by sharing the model parameters, allow flexibility of serving based on DRAM and compute available while delivering more accurate results.
💬 3 (+2) open on reddit ↗
▲
86
+16
2👁
r/LocalLLaMA · u/pmttyji · 8h ago
GitHub - google-ai-edge/ml-drift: GPU-Accelerated AI/ML Inference

Blog Post : ML Drift: Next-Gen GPU AI/ML Inference at the Edge - Google Developers Blog

The Google AI Edge Team is excited to announce the open-source release of ML Drift, our high-performance, cross-platform, on-device GPU compute engine specifically built for on-device AI/ML inference, under the Apache 2.0 license. By abstracting hardware and low-level API complexities of on-device GPUs across OpenGL ES, OpenCL, Metal, and WebGPU, ML Drift empowers developers to build real-time, interactive ML experiences from advanced video effects to generative AI across multiple platforms. Serving as the core GPU acceleration engine within LiteRT, ML Drift is also available as a standalone library for custom graphics and inference runtimes providing a unified foundation to deliver peak performance everywhere.
💬 11 (+2) open on reddit ↗
▲
31
+13
4👁
r/LocalLLaMA · u/falconandeagle · 15h ago
Stepping away from Benchmarks and Code, what models are you using for Creating writing projects

All anyone ever talks in this sub is about x new model having x benchmark scores and how well it can code. Or some new tech to increase t/s for qwen.

So I wanted to go back to the roots of this sub and talk about how well local models can write.

I am not going to mention closed source models, this discussion is for open weight models only.

I don't have the best machine, so I am limited in what I can use locally.

So far I have mostly been running Gemma 31b finetunes. I have a finetuned system prompt for creative writing that I have refined over time. This is mostly for long form story writing and not RP. I use it to write Sci Fi and Fantasy novels and short stories and sometimes fanfiction.

I also have created a creative writing harness using Pi. It's mostly for self correcting and getting rid of AI slopisms.

Gemma 31B Mero-mero-v2 has so far been my favourite. It writes well, its does not feel lobotomized and its mostly uncensored. It's instruction following is okay, sometimes it makes mistakes and gets the character traits mixed up but my creative writing harness takes care of this.

Qwen 3.8 27b finetunes, Disappointing, I guess the base model is really fucking bad for creative writing purposes and even with finetuning you can only do so much. It does follow instructions quite well but its writing is just terrible.

Muse Glimmer 30b, this one is an interesting one, sometimes it will give really good prose but sometimes it will respond like a 2b model, the variance seems to be super high for some reason. Maybe something to do with chat template, I am not sure. Anyway I think even when bad it's still better than Qwen.

Gemma 4 26b4a, pretty good but gets beaten by its bigger bother Gemma 4 31b. I do use it when I want faster responses and to check for issues with my initial scene beat prompts.

Gemma 31B Artemis, drummer's finetune, eh, I was a bit disappointed with this. Seems like its more focused on RP rather than long form story writing. Also got refusals which I did not get with the mero finetune.

So that has been my experience writing with some local models. What are you guys using for local creative writing projects? Have you guys tried a creative writing harness?

I just want to mention quickly my writing process before I end this post, I write using an outline and scene beat prompts that are hand written. So far AI has been extremely disappointing in being truly creative and still requires a ton of hand holding to not write the most cliché ridden drivel.

💬 59 (+38) open on reddit ↗
▲
14
+12
6👁
r/LocalLLaMA · u/Distinct-Pie2389 · 18h ago
Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s

HauhauCS ships their uncensored Qwen3.8-27B as GGUF only. NInfer, a C++/CUDA wanted its own format.

Now the same model that ran at 91.6 tok/s / 131K under llama.cpp does:

  • 262K context (the model's full native window)
  • \~130 tok/s decode with MTP3, 70.8% acceptance
  • 3,591 tok/s prefill on a 9K prompt
  • Perplexity within 1.3% of the official artifact, so the conversion is clean
  • Vision and tool calls still work

• E8 4Bit quant

Converter + writeup here: https://github.com/T-Crypt/ninfer-4090/pull/4

Questions welcome.

Repo: ninfer-uncensored

💬 31 (+31) open on reddit ↗
▲
40
+9
2👁
r/LocalLLaMA · u/SoAp9035 · 8h ago
Tested Mellum2.1-12B-A2.5B on PI Coding Agent - surprisingly usable, but not great at one-shot projects

I tested Mellum2.1-12B-A2.5B (Q8) locally using Pi and llama-server. All five tests were one-shot.

Results were pretty mixed:

\- Pelican SVG, Browser OS, Minecraft: Poor results.
\- Bouncing Hexagon: Physics were okay, but surprisingly it made it run in the terminal.
\- Flappy Bird: Completed it, but the visuals were very basic and the game was way too difficult.

Pelican SVG

Browser OS

Minecraft

Flappy Bird

Bouncing Hexagon

Okay, these might not look great, but hear me out. This model is actually pretty good.

its agentic behavior is goood. Tool calling is genuinely good, and it often catches and fixes its own editing mistakes. One time it failed to call the tool and stopped completely.

I also tried it on an existing game project. It explored the codebase and successfully changed the dogs' jump height to 3x.

Another nice surprise: after I recorded the tests, I asked the model to convert the screen recordings into GIFs and organize them into a specific folder. It handled the whole thing without any issues.

It does sometimes struggle with unfamiliar workflows, though.

Running it on a laptop with 8GB VRAM and 32GB RAM at 131K context. I'm getting around 40 t/s.

You might ask why I don't use Qwen3.6 35B-A3B instead. It gets slower with larger contexts, and my laptop gets so hot that it actually burned out the lid sensor. My laptop can barely handle it.

Overall, not a model that blows me away with one-shot creations, but definitely one I can see myself using regularly for smaller development tasks. Thanks for the model jetBrains.

💬 13 (+3) open on reddit ↗
▲
10
+7
8👁
r/LocalLLaMA · u/Felladrin · 23h ago
Decisions in a js13kGame post image

I'm having AI models play Cat Goric, a 2D platformer I made in 2021 during js13kGames: https://js13kgames.com/games/cat-goric-escape-from-the-warp-chamber

The best result so far is 8 of the 14 playable levels cleared in one continuous run. Two models have cleared eight levels so far:
\- Qwen/Qwen3.8-Flash-Next
\- Cloudflare/clef

If you'd like me to try a specific model, please name it in the comments.

You can also try beating the game with a model of your choice. The project with instructions for running the challenge with different models/engines is here:
https://github.com/felladrin/ai-plays-cat-goric (PRs welcome!)

And attached is a 2-minute recording of Clef clearing eight levels (the game pauses while the model thinks, so the video is time-warped).

💬 3 (+3) open on reddit ↗
▲
12
+6
5👁
r/LocalLLaMA · u/woct0rdho · 20h ago
LoRA over GGUF: Train Qwen3.8-Flash-Next in 40G VRAM

https://github.com/woct0rdho/transformers5-qwen3.5-recipe

An update to my LoRA over GGUF series: Now we can train Qwen3.8-Flash-Next (125B-A6B + 51B engram) in 40 GiB VRAM, with no CPU offloading, with engram on disk that does not reduce training speed.

On Strix Halo it trains context chunk size 2048 at 9.5 s/it. That's 200 token/s. There is still room to optimize, compared to > 1600 token/s PP we've achieved, and the common sense that LoRA training takes 2-3x work of PP.

Since transformers 5.18, initial support for modern GGUF has been merged, and we can expect more work in this direction.

Spoiler: In the torch-ggml-ops repo there is something called GGTensile. Basically it's Tensile-like asm-level optimization on MMQ kernels. We already see it's faster than HIP in many cases. I'll make a new post when I have something to show on this.

I guess I'll skip DeepSeek-V4.1, unless someone can quantize or prune it to < 125 GiB.

💬 12 (+3) open on reddit ↗
▲
7
+5
5👁
r/LocalLLaMA · u/jjusko20 · 17h ago
Release: Qwen-2B-RCOL Dynamic Low-Bit Quantization (IQ1_M, IQ2_M, IQ3_M)

Hey guys - primarily a research release with working models,

Not so much a model as a psuedo-new quantization technique. I've been experimenting with a modification of ISTALab's RCO algorithm that can quantize models on a strict VRAM budget.

It's at its core an approximation algorithm that attempts to make up the difference with a few different strategies. I'm quite happy with how well it's working at low bits. I've seen significant KLD improvements, particularly in the 1 bit and 2 bit range. This should be model agnostic and it doesn't require loading the FP16 teacher into VRAM. I've included a little more detail on the model card and will probably publish the full recipes. I plan to apply this method to the larger models in the family next.

Unfortunately Unsloth does not have their KL divergence table available, but I've created a table for you to see the difference between these quants and standard llama.cpp imatrix quants. For accuracies sake, both my quants and the llama.cpp quants are calibrated on the sane wikitext imatrix dataset, and tested on two held out sets.

The KLD table

As you can see, particularly at extremely low BPW, these quants vastly outperform their standard imatrix KLD - particularly drastically cutting IQ1\_M divergence at nearly the same size budget.

These are more proof of concept quants, as 2B is a fairly small model and suffers from such low BPW, but here's a fun example of how these are still relatively functional at extreme compression

4\/5 right on an IQ1\_M version of Qwen 3.5 2B \(green, not purple\) - I think it's kinda crazy it can answer anything

and here's the standardized imatrix IQ1\_M quant (not using dynamic RCOL)

yikes

and the file sizes

GGUFs:

https://huggingface.co/trubisky/Qwen3.5-2B-RCOL

💬 1 (+1) open on reddit ↗
▲
5
+3
6👁
r/LocalLLaMA · u/ailee43 · 23h ago
Best method to train a persona model

Hardware available:

2080ti 11gb

5070ti 16gb

Goal: I've downloaded my whole internet history (reddit posts, emails, chats, etc) and i'd like to train a model to think, act, and talk like me.

I'm currently running a qLORA pass with Qwen3.5-9B as a base, but when i run A/B tests, its unfortunately still pretty clear its not me.

Whats the best method to do this these days? Obviously would like to stay as local as possible since the corpus contains tons of personal info

Arch document to show the path i've already taken: https://markdownpastebin.com/?id=2493bd33b50b48ebbe0a49f437f573a6

💬 13 (+9) open on reddit ↗
▲
7
+3
3👁
r/LocalLLaMA · u/mrgreatheart · 10h ago
Epyc inference rigs

I posted on here a while back asking for advice on upgrading to an Epyc system. Thanks to all who commented.

I’ve taken the plunge and ordered:

\- Epyc 7443 CPU (24 cores, 48 threads, 128 lanes)
\- H12SSL-NT-B motherboard (the variant with two 10Gb Ethernet ports)
\- 256GB 2666MHz 2Rx4 RAM

To this I will be adding 72GB of VRAM across four GPUs:

\- 3090 (24GB)
\- 5070 Ti (16GB)
\- 2 X 5060 Ti (16GB)

Unfortunately I have to wait 2 weeks for the motherboard, so naturally I’m wondering what difference it will make.

The upgrade should:

\- double the PCIe bandwidth of all four GPUs (x4 to x8 for the 5060s and x8 to x16 for the other two). They will also all be on proper CPU lanes instead of chipset and dodgy m.2 adapters.
\- increase the RAM memory bandwidth by at least 50% (96GB to 170GB/s theoretical, perhaps 140 in practice)
\- increase total RAM by 192GB (64 to 256)

And it will open up the possibility for more GPUs later.

Obviously this won’t make much difference to dense models although I’m hoping for a nice bump in prompt processing due to the increased link bandwidth.

But I’m excited about bringing my total memory pool up to 328GB and opening up access to things like GLM5.3-flash and larger quants of Qwen3.8-flash-next.

Has anyone here built something similar?

How does your system handle large MoE models that fit fully in RAM & VRAM?

💬 64 (+64) open on reddit ↗
▲
11
+3
2👁
r/LocalLLaMA · u/EmilPi · 5h ago
Just another purely open-weight models benchmark

Live results, coding benchmarks included, agentic benchmarks included, domains and other classifications filters included. (Spoiler: DeepSeek-V4-Vision-Exp rules, but other models have their rule areas): https://beta.locallm.top

Evaluated by domain experts (my friends mostly; coding part is evaluated solely by me), classified by domains/languages/intents (can be filtered on the home page), agentic coding benchmark included. This is only public part of the data (whoever has private queries and evaluated models on them sees the results differently).

\*\*Some of the plethora of current limitations\*\*: evaluations/questions coverage for the domains/newer models is incomplete and imbalanced, only part of the questions classified, UI/UX under-developed, focus was on small models and lower quants.

\*Also\*: everything runs on local 2xRTX 3090 + 128GB RAM PC, everyone can create queries and evaluate models' answers, new runs (especially agentic) slow to appear (see PC specs).

No LLM-as-judge on purpose (and I sometimes regret it).

Benchmarks currently being extended & evaluated: 1) Agentic coding 2) one-shot coding 3) Agentic retrieval 4) Agentic story-writing.

New models being evaluated: 1) Qwen3.8-27B (Uncensored). New models to be evaluated soon 1) GLM-5.3-Flash-UD-IQ3XSS 2) Mellum2.1-12B-A2.5B

If this looks interesting to you:

CALL FOR HELP

I ask for help from those who also want to build a community benchmark, developers and domain experts alike! Many things are going to be implemented sooner or later, and the agentic benchmarks are just the beginning.

I call for help in these areas: 1) compute resources 2) evaluations 3) development - basically everything. But any proposal and feedback is valuable!

I will add more info and answer questions in the comments section.

P.S. No AI used for writing this post (even though English is actually not my native language /s).

💬 16 (+7) open on reddit ↗
▲
9
+3
2👁
r/LocalLLaMA · u/neph1010 · 6h ago
LlamaTale v0.43.0 - MCP Server for story creation

In the dawn of the LLM craze, when wild Llama2's roamed free and 4k tokens was considered a fairly long context for local models, I forked an interactive fiction and MUD library called Tale to experiment with LLM generated content. The idea being a text based adventure with totally dynamic content that never runs out of context.

That was many generations ago in LLM space, and I haven't really done much with it for a long time either. But with the advent of agents, I started brooding how they could be used to create stories and worlds. I finally got around to implement something together with trusty Qwen3.8 27B. (There were other models involved, but none worth mentioning). So yes, this is vibe coded slop, something you see a lot of here, but it's built upon regular AI-assisted slop (admittedly because vibe coding was not available at the time).

I don't know if anyone is still interested in LlamaTale, but here it is. Why would you use the MCP server? (A more important question is "Why would you use LlamaTale?", but that's a different question).

Maybe you have a regular RP story that you would like to explore in a more structured way. Expand it to a living world of "infinite" size without running out of context (maybe less relevant now, but remember the 4k tokens). The world stays consistent, but characters and mobs may wander (It's tick based).

https://github.com/neph1/LlamaTale

💬 2 (+1) open on reddit ↗
▲
33
+3
2👁
r/LocalLLaMA · u/bodhi371 · 9h ago
Qwen3.8-Flash-Next-GSQ-RCO (IQ3_S): ~20-30 tok/sec decode & 300-90k tok/sec prefill on 12GB VRAM + 32GB RAM + NVME

I got Qwen3.8-Flash-Next-GSQ-RCO-Abliterated running at IQ3\_S with just 12GB VRAM and 32GB system RAM, achieving 20-30 tok/sec decode (Q2 achieves 39-45 tok/s) & 300 to \~90 thousand tok/sec prefill @ 131k context, on a custom fork of Strata.

This fork has tonnes of architectural changes, all are very experimental and will probably break. But the performance makes up for it. This feels like local Opus in some regards, on sub 2k in compute.

https://github.com/bodhi37/Qwen3.8-Flash-12GBRAM-32GBVRAM-SSD-Recipe
https://github.com/bodhi37/strata

💬 20 (+2) open on reddit ↗
▲
4
+2
2👁
r/LocalLLaMA · u/SnooPeripherals5313 · 6h ago
Demo of many knowledge graphs. post image

Over the last year, I've been working a lot on structured legal data. Here are some knowledge graph representations that may be of interest to those of you building your own projects. 5s on each representation.

To pre-empt any questions:

  1. No I do not think KGs are super useful for agents, but I like them as a visualisation.
  1. This repo is not public, but there are lots of good opensource ones like graphiti and it's easy enough to recreate this in 3js
▲
6
+2
2👁
r/LocalLLaMA · u/Historical-Internal3 · 7h ago
DGX Monarch - Dual Spark Rendering w/ ComfyUI

For those of you who have been wanting to utilize your dual Spark cluster for image/video gen with ComfyUI - today is your day.

I started this personal project back in April of this year, and it’s at a point where I’m fine with a public release.

To put this as simply as possible - you can get nearly 2× render speeds with this repo. Results vary by model and settings. BF16/FP8/NVFP4/INT8/etc. supported, depending on the model. No GGUF.

Link to the repo: https://github.com/Deen-Media/dgx-monarch (highly recommend reading the FAQ after the README)

The rest of this will be MUCH more boring, so you can skip if you just want to start using it. You’ve been warned, and I cannot refund any time spent reading what you already have thus far, and especially the rest of this post. More words. Four more pointless words. K just checking.

Before you ask - yes, I heavily used AI to develop this. But most of MY time was spent validating outputs and performance, and actually using this project.

In terms of maintainability, well, I have and will continue to do my best to keep the development workflows I’ve built up to date so that new models can be safely added without breaking anything. I do not intend on sharing these development workflows, as I personally do not see the value in doing so relative to the effort I would need to put into making sure they are even safe for me to share. To be clear, I mean my personal AI-agent development workflows, not the example ComfyUI workflows included in the repo. Workflows like this, in my opinion, should be something you craft and tailor/customize yourself.

I’m fully expecting AI-generated PRs, and I realize at the end of the day it’s my slop vs. yours. All I ask is that you actually sit down and validate/render/test your changes and include those results. I’ll need time to review and test them too, so please be patient. I really enjoy USING this tool and would like to continue to do so. Maintaining comes second to me personally, but ironically, using and maintaining go hand in hand.

It’s easy to vibe-code a small project. It is VERY difficult and costly (money, not just time) to vibe-maintain a large project.

For transparency, a project like this has depleted (and I mean 0%/flatline), on a weekly basis, the following subscriptions of mine since the start:

OpenAI:

\- 1× Astra $200 account\* (now Astra $500 account\*\*)

\- 1× Business Astra seat

\*I took advantage of every single banked/global reset and got real lucky on the global ones.

\*\*Strictly to close extra sloppy PRs hella fast with ultrafast (but really to utilize usage on spontaneous global resets that the community gets an hour or two heads up on). Plus I enjoy paying more than double to have roughly the same usage as I had before.

Anthropic:

\- 1× Fable $200 account

\- 1× Business Fable seat

Cursor:

\- 1× Fable $200 account (Grok did not touch the working code - it had a different use case)

Google:

\- 1× Ultra $250 account (don’t freak out - Gemini did not touch the working code)

(Note: Astra/Fable were not around when this project started. Was fun.)

I will maintain these subscriptions for as long as reasonably possible, as I do have other projects I intend on releasing (more DGX-specific ones) in the coming months here.

However, there WILL be a point where it won’t make sense to maintain all these subscriptions. Whether that’s due to local models finally being sufficient/effective enough to handle my workflows, frontier costs skyrocketing or usage allowances dwindling as the era of subsidization ends, or the projects simply no longer needing a high level of maintenance (unlikely).

That’s all, thanks for reading.

💬 9 (+2) open on reddit ↗
▲
2
+2
2👁
r/LocalLLaMA · u/empiriolabsai · 8h ago
Aplomb 1 is #1 on Typed Decisions

Aplomb 1 is #1 on Typed Decisions, an official Hugging Face benchmark for probabilistic decisions. At 5.3B parameters, its probabilities are the closest to the true answers on the leaderboard, by KL from gold and by Brier score, ahead of a 27B model about five times its size.

Typed Decisions gives a model one piece of unstructured text and five typed questions about it at once. Every answer is a probability distribution, scored against the gold distribution across 400 cases and 2,000 decisions. KL from gold and the Brier score measure how far a model's probabilities sit from the true answers, so lower is better.

Results:

  • KL from gold: 0.123, #1 on the leaderboard and 40% lower than the 27B model in second place
  • Brier score: 0.065, #1 on the leaderboard and a third lower than the 27B model in second place

Aplomb 1 takes text, JSON, images, video and audio in one request and reads up to 1M tokens. On our API it answers a short question in about 15 ms of model time and reads a 1M-token document in about 3 seconds, at $0.02 per 1M input tokens with output free. The weights are open on Hugging Face.

Leaderboard: https://huggingface.co/datasets/LocalLLaMA/typed-decisions?leaderboard\_task\_id=kl\_from\_gold

Weights: https://huggingface.co/empiriolabsai/aplomb-1

Try it: https://platform.empiriolabs.ai/dashboard/playground?model=aplomb-1&utm\_source=linkedin&utm\_medium=social&utm\_campaign=aplomb-1-typed-decisions

▲
5
+1
4👁
r/LocalLLaMA · u/leo-k7v · 15h ago
Supertonic 3 TTS dissolution and liquidation?

I use their model in my up with hand-rolled inference in C.
HF page says:
"This project's sample code is released under the MIT License."
and "OpenRAIL-M License" for the model.

Can I still use their model after liquidation?

💬 6 (+2) open on reddit ↗
▲
4
+1
3👁
r/LocalLLaMA · u/OkFly3388 · 12h ago
rtx 4090 + huawei atlas duo for qwen flash next ?

I have rtx 4090 and 64 gb of ddr4 ram. This is enough to fit smaller quants of qfn, but I found them not great and switched back to qwen3.8. However, there are option that bother me a lot, buy huawei atlas duo, thats another 96 gb of ram thats twice faster than system ram and another AI accelerator, thats faster than cpu.

Anybody tried that ?

💬 9 (+5) open on reddit ↗
▲
1
+1
2👁
r/LocalLLaMA · u/Su1tz · 8h ago
Which Image Gen model is best for infographics?

Hello guys,

I have a summary workflow for my school where I have Gemini summarize the lecture content, then generate infographics on said content. However, gemini's style keeps constantly changing, running out of limits and other stupid bs like that. Its unreliable. What image gen model is best for the task of complex infographic generation? Also, what model could take the input and actually produce it into a infographics suitable prompt? I usually have Qwen 27B (lately qwen flash next) generating text content for me but If I run the image model next to it im afraid ill run out of vram pretty fast.

💬 23 (+4) open on reddit ↗
▲
0
 
1👁
r/LocalLLaMA · u/Loose_Doubt367 · 22h ago
what can i do with my local (qwen3.6 35b) model inside pi harness or any other harness

I've been playing a lot with the models and have settled with the qwen3.6 35b and pi, but im not sure on what to do other than code html games. Any suggestions?

▲
0
 
4👁
r/LocalLLaMA · u/Distinct-Pie2389 · 21h ago
llm-tune: getting local models that actually perform post image

Made a little agent skill called llm-tune to help find the best settings for running local LLMs on your hardware.

Still working on it, but I'm looking to test it across more GPU setups. If you try it out, feedback and benchmark results are welcome.

Supported:

  • Architectures: Dense, MoE, hybrid MoE/Mamba
  • GPUs: NVIDIA, AMD, Intel Arc
  • Apple Silicon: M-series Macs via MLX
  • Engines: llama.cpp, Ollama, vLLM
  • Tuning: Quantization, context/KV cache, GPU offloading, MTP, sampling, reasoning, and agent harness settings
  • Benchmarks: VRAM/RAM usage, tokens/sec, context recall, and output quality

Currently measured on an RTX 4090; other hardware and backends are documented but need more real-world testing.

💬 7 (+3) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/MrCatberry · 17h ago
Reverse Engineering Web Application/Service

Hi Guys!

Is there a known harness/workflow that makes it easier to Reverse Enginerring a Web Application/Service thats behind a payed subscription?

My current problem is that I use a service that costs me quite a lot of money but still does not have all the features or tweaking options I need.

Now I'm asking myself, if it would be best to describe every feature myself or if there is a way to let a LLM "explore" the Web Application/Service by itself and write it's own notes what it needs to code.

Did somebody here do something familiar and has some tips?

Thanks in advance!

💬 10 (+7) open on reddit ↗
▲
0
 
5👁
r/LocalLLaMA · u/General-Spite1222 · 17h ago
NVFP4 is the GOAT, prove me wrong.

It being on par with bf16: https://huggingface.co/nvidia/Qwen3.8-27B-NVFP4#evaluation

It being on par with FP8: https://huggingface.co/nvidia/Qwen3.8-Flash-Next-NVFP4#evaluation

Who here is smarter than Nvidia and can price them wrong? I see a lot of trash talk, but no numbers to back it up.

NVFP4 is as good as Q8. Prove me wrong! Feel free to downvote if you cant prove otherwise.

💬 33 (+18) open on reddit ↗
▲
0
 
4👁
r/LocalLLaMA · u/Old-Sherbert-4495 · 14h ago
Qwen3.8: Flash Next iq3 xxs is dumber than 27B iq3 xxs?

it consistently could not pass this. had to follow up couple of times and by the time it has already exhausted 120k context limit.

"build me a 3d simulation of the solar system in a standalone html file. no three.js, No web GPU, no any extra libs, just plain html, css, js, and svg."

On the flip side 27B, one shots them like perfectly.

One thing i noted was flash next version had way more features and a pretty complex UI. 27B some what simpler and gets the thing working.

I run this to get a vibe of the model. Then ran some agentic coding tasks, it sure does think wide but when it comes to the implementation it always fails. and needs few follow up steps. overall loads of tokens used.

Has this been the same experience you had?

💬 45 (+26) open on reddit ↗
▲
0
 
4👁
r/LocalLLaMA · u/paq85 · 15h ago
How much does your local LLM server really cost? (power draw + tok/s, self-hosted vs cloud API, free in-browser) :)

I've been tuning local LLM setup for many months, and the number people actually want is the electricity bill. This calculator takes your GPU's power draw (W) and token speed (tok/s for prompt and decode), plugs in your energy price per kWh, and shows cost per hour, per request, monthly, and cost per 1M input tokens — with and without KV cache.

There are 3 GPU presets (RTX 4070 Ti Super, RTX 5090, RTX 5090 eco) if you want to start from realistic numbers and adjust from there. The math treats cached tokens at 1/100 the time of uncached, so the realistic scenario is fixed at 60k input / 50k cached / 2k output / 90% uptime.

The cloud API comparison compares each self-hosted profile against a reference API ($0.25/1M input, $1.20/1M output) at the same scenario, so you can see whether self-hosting actually saves money or costs more.

Everything runs in your browser, nothing is uploaded and there is no sign-up. From my experience, the per-request number is the one worth comparing with the cloud — the monthly bill is just the per-request cost times your actual request rate.

https://appdoesit.com/apps/llm-cost-calculator — it's one of the 139 free tools in the catalog.

Try it by yourself :) — how does your server compare?

💬 25 (+5) open on reddit ↗
▲
0
 
2👁
r/LocalLLaMA · u/taylorwilsdon · 9h ago
Started an open source project on Reddit 18 months ago that blew up into one of the most popular MCP servers in the world, just released v2!

This isn't an advertisement, and it's very much local and open - I already don't have enough time to keep up with the existing pull requests and issues... just a fond look back on how much this space has grown and matured in the past year. Shit was the wild west back then. Nowadays I can run qwen3.8 on a mac mini fast enough to drive this at full speed for free using native tool calling all day long. When I first released it, local model tool calling was much more hit or miss.

This week, I shipped v2 to a completely different world - more than 2 million people have downloaded it, 400 people have opened more than 400 issues and 800 PRs, 150 of you wonderful people have contributed code to main. Google Workspace MCP is both an MCP and CLI that can control every part of your google account.

I wrote up a little blog post along with the v2 release notes here and run through the evolution of coding harnesses, model capabilities and community participation over the past year and a half.

The code is MIT licensed, 100% open and yours to use, abuse, steal and sell as you please, the way it should be. As always, feedback, criticism & PRs are welcome!

💬 16 (+1) open on reddit ↗
▲
2
 
1👁
r/LocalLLaMA · u/unraveleverything · 2h ago
Is there a music embedding model?

Has anybody built/released a music embedding model trained on a large library of diverse music?

▲
1
 
1👁
r/LocalLLaMA · u/vexatious-big · 3h ago
The Nvidia RTX Spark laptops with N1X now available to pre-order (UK)

ASUS ProArt P16 OLED 16" Laptop - NVIDIA RTX Spark™ N1X, 128 GB RAM, 2 TB SSD, Nano Black
For the low price of £5999
https://www.currys.co.uk/products/asus-proart-p16-oled-16-laptop-nvidia-rtx-s…
other configurations available with less ram:
https://www.currys.co.uk/deals-on-computing/new-nvidia-laptops

▲
0
 
1👁
r/LocalLLaMA · u/No_Farmer_495 · 3h ago
Strata for GLM 5.3 Flash

Hi. Everyone's focusing on qwen flash on strata engine right now. But what about Glm 5.3 flash? I got a weird system, dual xeon(avx1) ddr3 ram and one 3060,one p100 and one 3050. So far all the strata glm 5.3 flash forks failed me and my system. Could anyone help? Or would it be possible for more people to request GLM 5.3 flash to be officially supported by strata? I'm reffering to the GSQ-CRO quant and unsloth Q4.

▲
2
 
1👁
r/LocalLLaMA · u/SirLordBoss · 3h ago
Best value upgrade for my setup

I'm a GPU poor, unlike the vast majority of you. I have finally secured funding, and would like to upgrade my rig.

At the moment, it consists of:

\- RTX 5060 Ti, 16 GB VRAM

\- 32 GB RAM

\- 2 TB NVME

Given the skyrocketing prices of everything, I'd rather not wait for Black Friday to come around. Last year, it brought the RAMpocalypse, and we've not yet recovered.

So, given this scenario, what would be the best value upgrade for this setup? I'd set the budget between $1000-3000. Make that euros, am in Europe.

I've considered:

\- 1/2 used 3090s

\- swapping the 5060 for:

\-- an Intel Arc B70

\-- a 7900 XTX

\-- a 4090 (5090's are now absurd)

▲
20
 
1👁
r/LocalLLaMA · u/pseudotensor1234 · 3h ago
H2O-Lightning-4B: Apache-2.0 4B Decision model, official #1 open model on JevBench (above Jev)

Disclosure: I work at H2O.ai.
  
We released H2O-Lightning-4B, an open-weight (Apache-2.0) model for the "decisions API" style of inference that Jev made popular: you send a state plus typed questions (pick one / yes-no / score), and get calibrated probabilities back from a single forward pass. No generated tokens, so it's fast and cheap.
  
\*\*Results (JevBench, public leaderboard):\*\*
\- Composite score 72.5, vs Jev 1.13 at 71.5; currently the top open model
\- Leaderboard: https://benchmarkheaven.com/jev-models
  
\*\*Running it:\*\*
\- Base: Qwen3.5-4B, fine-tuned
\- Stock vLLM plus a small open shim (in the repo); \~30 ms per decision on an H100
\- Your data stays local, no per-call fees
  
\*\*Coming soon:\*\* 12B and 31B versions, which in our internal testing are considerably smarter than
 Jev, still open-weight and still one forward pass per decision.
  
 \*\*Demos\*\* (inbox triage of 1,000 insurance claims, a multi-browser web agent, DOOM on the decision clock): https://youtu.be/2Qp04Wu0A14
  
Weights, model card and serving instructions: https://huggingface.co/h2oai/h2o-lightning-4b
  
Happy to answer questions about the setup and latency.

▲
1
 
1👁
r/LocalLLaMA · u/SnooPeripherals5313 · 3h ago
Codex harness visualisation post image

Little animation of how codex works I made for my own understanding. Not too dissimilar to pi (but much more opinionated)

▲
0
-1
6👁
r/LocalLLaMA · u/MasterSama · 20h ago
How is it possible to enable tool calling like file read/write/execute in Strata (Qwen3.8-next-flash) web ui?

I've just installed Strada and Qwen3.8-next-flash on my PC and I was amazed by it.

The problem is, I cant seem to find a way to enable or add tool callings, so it can read/write/execute files on my PC. How do you do that in Strata?

Thank you very much in advance

Update:

Per suggestions, I went on to use a harness and ended up liking the deepseek harness a lot. its working great. Thank you very much everyone.

💬 16 (+16) open on reddit ↗
▲
0
-1
2👁
r/LocalLLaMA · u/Loose_Doubt367 · 17h ago
What features do i unlock by switching rx6700xt 12gb (AMD) to 3060 12gb (NVIDIA) in llama.cpp?

is expecting higher tps the only feature that i'll be receiving?

💬 15 (+2) open on reddit ↗
▲
9
-1
3👁
r/LocalLLaMA · u/NoobSolid26 · 12h ago
I made a portable AI memory format that transfers both context and skill

Last month I created a portable memory format as a .txt file as a small project. I named it as: Memory Journal, in a shape of a skill card. It's essentially a large prompt wrapped in a text file.

Initial testing went better than I have expected, then did more tests and improved it in a course of a month, Because it was interesting to work on. It's meant to be general purpose, it can handle both casual and technical sessions, even though it can bend a bit. I found the it kinda cool so I decided to share it here. It's not perfect by any means but it works.

The philosopy behind was simple "If an AI knows the context, it can describe how it created something too". In addition to context it has ability to be reproducible, it can transfer skill, transfer mistakes. along with proven, unproven and disproven facts. Journal mostly gives free reign to the writer (AI) and stays kinda neutral. Journal has been split into A/B/C sections, that serve differen purposes.

\[Capabilities\]:

  • Right now it can allow you to resume your work with journal + project files.
  • AI can write a journal about EGO engine XML file conversion, and readers can reproduce the converter from journal alone. (Deepseek 4.1 (deep think mode) and Claude Sonnet 5 (high effort) recreated the converter in 1st try.)
  • Journal can also cause drifts in behavior and vocabulary if you use a philosophically charged journal

\[Usage\]:

  • You give this "card" to an AI then they will create a text file as a memory journal.
  • You can supply additional commands by yourself, like: "add completion status to goals" or "don't add x part to journal" and such.

https://preview.redd.it/19xcinownfuh1.png?width=795&format=png&auto=w…

\[Note\]:
It's best to use this journal with a capable AI model preferrably with an effort slider. Also yes skill card larp is actually part of the design, I know it looks eccentric but, isn't it cool?

\[Memory Journal format itself and outputs of it is down below\]:

Memory journal format itself: Memory Journal S4 - Pastebin.com

XML converter reproduction journal: XML Converter S3 - Pastebin.com
TRD2 Modding journal (4th writer on the line): TRD2 Modding Final Journal - Pastebin.com
Weird casual chat journal: Psychic pepperoni journal - Pastebin.com
Joke casual chat journal: Pun journal - Pastebin.com

💬 17 (+7) open on reddit ↗
▲
14
-2
2👁
r/LocalLLaMA · u/Chromix_ · 8h ago
A 24KB stand-alone HTML-LLM that can generate consistent stories

https://preview.redd.it/7sknv6omnguh1.png?width=1023&format=png&auto=…

Want to try? Just go here and click "New seed" a lot. You will get different stories, it just requires some clicking - lucky RNG draws: https://output.jsbin.com/nikupemuta/1
You should even get above 60 tokens per second with a smartphone.

FAQ:

  • Why?! Because it's possible, and fun. My original idea was to simply mash MacroStory and LittleBit together, but that did not work at all.
  • Is it relevant? Not a tiny bit! Even a 3 MB HTML file with external dependencies would be loaded in a second and could serve a way more capable model with a lot less work. Squeezing the model itself to 20 KB at 0.01 KLD was trivial. Saving another 5 KB while maintaining output quality required a lot of time.
  • How? Local Qwen3.8, some ideas, lots of patience. Adaptive quantization and QAT made it happen, LittleBit, Rotation, etc just made it worse. HTML/JS packing with minify, zopfli, and a bunch of structural script changes helped with the result.
  • Was this all made by you? Not at all. This would not have been possible without the 81 KB (FP32) MacroStory model. I "just" tinkered around with it somewhat.

Details for those who're interested:

  • Size-reduction rules are very different for a tiny model than for a large one.
  • The embeddings are 60% of the model size, while the tensors are the largest part in a normal-sized LLM.
  • The tensors of tiny models are extremely sensitive to quantization, especially with the Ouro looping transformer format of this LLM. The embeddings had some room though.
  • It's mathematically impossible to save space with 1 bit quants or less with the LittleBit approach here, as the model matrices are just too small for that - the gains would get eaten up by the overhead of the correction data.
  • GGUF K quants would also not help for the same reason: the model matrices are simply too small for the introduced overhead.
  • Even if some tensor size could be reduced by LittleBit: the 3 KB of quantization gain would be eaten up by 3 KB of newly required safetensors metadata. Nobody thinks about metadata sizes in normal-sized LLMs.
  • Yet aside from that: A modest 4 bit LittleBit quant would still break the model. The currently chosen approach meanwhile comes with a convenient 0.04 KLD. This already causes a very occasional duplicate sentence or non-matching story start.
  • The MacroStory model has a handful of dead tokens that could be exploited for further size reduction, as they'd never make it through the sampler on their own.

Since you've arrived down here, there's a bonus for you: Local Python inference (YMMV) and the non-packed HTML. Just save this image to disk (important: "Download original image"). I originally wanted to use it here or on imgur, but both wouldn't let me. Open with 7-Zip, WinRAR, etc to unpack it. Or on the console - even on Windows - use either:

tar -xf MS256story.png
7z e MS256story.png
python -m zipfile -e MS256story.png

💬 13 (+1) open on reddit ↗