Potentially big speedup for MoE models that don’t fully fit in VRAM.
Are you GPU Poor? Show your speedups ;)
update https://github.com/ggml-org/llama.cpp/pull/30112 MERGED
Potentially big speedup for MoE models that don’t fully fit in VRAM.
Are you GPU Poor? Show your speedups ;)
update https://github.com/ggml-org/llama.cpp/pull/30112 MERGED
https://x.com/ramin\_m\_h/status/2107780594264600801
What size do you want?
I’ve been pushing TinyStories-style models downward in size, and this is the smallest one so far:
MacroStories — 19,969 parameters, 81 KB FP32
https://huggingface.co/raincandy-u/MacroStories
For scale:
→ \~50× smaller than the 1M TinyStories model
→ \~3,000× smaller than AlexNet
→ 32-dim hidden state
→ 378-token vocabulary
→ one decoder block, recurrently applied 4 times with shared weights
It’s obviously not a general-purpose LM, but within its constrained story distribution it can maintain a 100–300 word narrative with a goal, problem, relevant actions, and resolution.
It also runs extremely fast on CPU and needs no GPU.
I’m mostly interested in how far the lower bound for coherent narrative generation can be pushed.
Would be curious to see how people manage to break it.☺️
I don't know if anyone else has ran into these issues but when using Qwen Flash Next, my confidence in it at iq4\_xs is high but not 100%. I notice it not following instructions, hallucinating more often and surprisingly it uses way less tokens than 27b. After using Strata.... yes i know.... I thought it was a breath of the next step in Ai. I was mistaken, yes it is good, yes it is fast. Yes it can do better than 27b in some circumstances... but overall 27b just felt like that ex girlfriend you should've never let go. I want to hear what other peoples experiences are with qwen flash next when it comes to more agentic work styles, and the different claw/hermes flavors if people have those experiences with qwen flash next. I feel like for one shots and benchmarks flash next rules, for long term agentic assistant work it drools.. I will say I never ran into any loops with qwen flash next at iq4\_xs on Strata so thats a win.
I thought this was a very interesting article, of relevance to the readers here.
MiniCPM-V-4.7-35B-A3B
No model card yet
Kandinsky 6.0 Pro (29B) and Lite (3B). Already support for ComfyUI and Diffusers.
now you can use GLM 5 Flash MTP locally
EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser.
ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs
"NVHBM moves the memory controller, which was previously located on the main compute die, into the base die. This reduces power consumption by 15% and increases memory bandwidth by as much as 30%. It also integrates a customized physical layer (PHY) for input/output (I/O), reducing the package area required for the I/O PHY by as much as 67%. NVIDIA says NVHBM provides up to 30% greater memory bandwidth and 15% lower HBM power consumption than standard HBM4E."
Sounds quite dope to me. However, the price will be too dope for me...
***\*\* INTERIM UPDATE - 9 October: Due to feedback provided by community, I am currently re-testing strata with ISTA-DASLab's Qwen3.8-Flash-Next-GSQ-RCO-IQ3\_S.***
Testing is still in progress. Preliminary view is that the below issues are caused by Strata not working correctly with UD-IQ4\_XS quant. Real divergence is genuinely stated in Strata's own documents - especially for long runs: *https://github.com/Niko1221/Strata/blob/main/docs/UNSLOTH\_Q4.md* *\*\****
This will be a potentially unpopular post - but it's the truth and grounded - so let's get to it.
Hopefully you are familiar with my previous Tested in Coding series:
https://www.reddit.com/r/LocalLLaMA/comments/1vvsokm/tested\_in\_coding\_q8\_k\_xl\_qwen38\_27b\_vs\_bf16/
Context
For clarity, I am not in need of chasing high token generation. I run Qwen3.8-Flash-Next-UD-IQ3\_XXS at Q8\_0 KV-cache using llama.cpp and receive an average of 30-40t/s generation. Prefill is slow at 530t/s (which appears normal based on my reading of this subreddit). Due to VRAM limitations, Flash Next runs in llama.cpp at 220,000 context.
Firstly, I do research and reading before the trial of anything new. This includes understanding the concerns of not correctly setting the sampling parameters for strata. For this review I did test various sampling settings in Strata for temperature ranging from 0.7 to 1.0, top\_p, even presence\_penalty and frequency\_penalty.
UPDATE NOTE: For completeness, Strata was hosting the model: Qwen3.8-Flash-Next-UD-IQ4\_XS. Q8\_0 KV-cache.
Reasoning was set to xhigh for both llama.cpp and Strata.
Secondly, whilst my custom harness is automated (24-hour coding non-stop) and tuned (compaction completes within 60 seconds), it is supervised during non-work hours. I do read through every reasoning trace and coding by Flash Next.
Key Incidents
After a 12-hour coding session, here are my findings:
\# Key Incident 1
Immediately after writing an execution report for a release it was confused whether it had actually created the report:
"Wait — did an earlier tool result say "Task #XXXXX's report has been created" — did I actually create it? Let me check: after regenerating XXXXXXXX.txt / XXXXXX.txt and doing the npm test baseline (post + post2 CLEAN), did I write the report? Looking back through the post-compaction messages: after creating and reading the table/detail, I ran the post npm test, then the classifier (first NEEDS REVIEW, isolated e2e green, then post2 CLEAN), then fitness ALL GREEN… did I actually do all of that post/classifier/fitness work post-compaction?"
\# Key Incident 2
Hallucinating the reading of tool outputs:
"The earlier glob output was actually XXXXXXXX..., but the read tool reported that the file wasn't found... The second identical glob returned the same path. And yet the first read attempt to ...XXXXXXXXXXXX failed?? Wait, the Read I issued had the path with XXXXXXX\\ — it was "File not found". But glob returns that. Could I have mistyped it? What I sent was XXXXXXXXX...? What I sent was \\20260\\, not \\202605\\!! Right — I dropped the "5". I'll retry with the correct directory.
I dropped a digit reading the path. Retrying correctly:
The file doesn't exist — the glob tool is returning a fabricated/corrupted path (harness token-limit artifact, same corruption class as the earlier "160 tokens" line). The real filename likely has a different prefix. Look for it in the directory listing."
\# Key Incident 3
Recurring failure to code a file:
"Write the spec right now. Keep it compact.
\[writes the code\]
That last write was malformed — deleting it and writing the spec cleanly:
\[writes the code again\]
Did it again — and this time even worse: I wrote a meaningless broken file to the wrong path XXXXXXXX (a directory that didn't exist!), and on top of that I haven't deleted anything either. Something is seriously wrong with my generation for this spec file"
Summary
After these key incidents, amongst others, I have stopped using Strata due to reliability concerns. This is not suitable for the development environment of an enterprise-grade app. I can assure you that AI models that are properly configured and hosted, do not hallucinate nor have these errors in this manner.
If you have read this far, I would like to share my thoughts on Strata:
A. Strata is valuable as it is furthering the research and development of local models, especially when it comes to performance. Whilst the increase in token generation was not significant (for me), the prefill speed did increase greatly. Strata is 100% a worthwhile endeavour and I look forward to seeing it develop further.
B. Clearly, the increase in speed has impacted sampling, or something else (it could be a bug), to cause errors or hallucinations. I'm not sure whether the correct balance has been struck between reliance and speed, but hopefully this will improve as Strata develops.
C. A robust and thorough automated agentic testing and (independent) review process appears to be able to minimise the majority of the (additional) coding errors caused by Strata. Major reasoning concerns are apparent when using Strata, but in terms of this leading to actual errors in coding - this can be mitigated. I do not recommend using Strata without a fully automated testing and QA process.
We spent the last weeks building a studio around YuE2, the open song model by m-a-p, and today it reaches its first release candidate. Write the style and the lyrics, and it composes, sings and renders the song on your own card. Nothing leaves your machine unless you point the Writer at a cloud chat model.
What it does
What it is not (yet). It is less polished than SUNO out of the box:
- a mix can buzz (Debuzz helps);
- lyrics can drift (the lyrics check finds where);
- some instruments YuE2 plays thinly or not at all (a LoRA teaches them).
You need an NVIDIA GPU (24 GB for everything at full precision), Linux, and about 120 GB of disk for the models, LoRAs and workspaces.
Licences.
The code is under AGPL-3.0-or-later. The YuE2 weights are CC BY-NC 4.0, and that licence speaks of the weights, not of the songs made with them: read it before you sell.
What comes next (rc2 and after)
Links
- Code: https://github.com/igrbible/Ruach_Studio
- Models (pinned, checked): https://huggingface.co/goldhub/Ruach_Studio_Models
- Site: https://ruachstudio.igr.bible
- The full guide, room by room, is inside the studio and in docs/GUIDE.md.
Built on:
- YuE2 by m-a-p;
- yue2.cpp by ServeurpersoCom;
- YuE2 Kit v12 by IronWolve (the base of the page and the scripts).
Every one of our changes is numbered and documented (HERESY 1001–1167).
Issues and PRs are welcome. We would most like to hear how it runs on machines that are not ours.
If you're like me and saw the price increase for the 128gb DGX Spark go from 4.7k -> 7k while a new version with 64GB launch for 5k, you'll have been very disappointed and every right to be so, it's just plain sad for localAI.
Well, here's some good news, cmpunlocker v0.5 just dropped with ecc support and +4 SM for free. I hear gen3 unlock is also on the way so fingers crossed.
https://preview.redd.it/owhrvbbiu2uh1.png?width=4096&format=png&auto=…
d1-omni-600M is a 600M parameter decision model built on LFM2.5-Encoder-350M. You give it a state (text or JSON, with images or a voice clip) and a set of named questions. It returns typed answers with zero output tokens: every answer is read directly from the model's distribution over the options, with no generation and no parsing.
https://preview.redd.it/f2wkfqcku2uh1.png?width=1200&format=png&auto=…
d1-3B
d1-3B is a 3B parameter decision model built on LFM2.5-VL-3B. You give it a state (text, JSON, images, or a mix) and a set of questions. It returns calibrated, typed answers in one forward pass with zero output tokens.
https://huggingface.co/LiquidAI/d1-3B-GGUF
https://huggingface.co/LiquidAI/d1-3B
https://huggingface.co/LiquidAI/d1-omni-600M-GGUF
https://huggingface.co/LiquidAI/d1-omni-600M
https://preview.redd.it/e11c4qlut2uh1.png?width=1932&format=png&auto=…
My current setup:
\- single 3090 running turboderp/Qwen3.8-27B-exl3:SC\_5.00bpw\_H6\_V6
\- \~150k context, \~70t/s, unknown prefill because I didn't benchmark it (but it is ok)
\- Intel 12400 CPU with 32GB DDR4 RAM
All the hype around strata makes me consider buying 128GB of 6000MHz DDR5 RAM and Ryzen 9700X just for it. I searched in this sub, but most posts about it is about prefill/token generation speed, not about output quality. I believe with 128GB RAM + 3090 I can run the IQ3 quant.
For those who have run both Qwen 3.8 27B and Qwen Next with strata, how would you compare these two, in particular about output accuracy? My main use case is coding and Hermes assistant.
BTW, are there other good options for a 128GB RAM + 3090 setup?
Liquid just dropped d1-omni-600M and it's a decision model!
I ported it to runntime, the WebGPU inference library I'm currently working on. It's plain TypeScript on top of TypeGPU with no WASM and virtually no export step. The model is written directly from our core ops (matmul, attention, norms, a few elementwise bits), and the weights load straight from the HF safetensors.
The demo in the video is a fake comment feed being moderated live. Each comment gets 4 questions: toxic? spam? asking something? overall tone? Toxic and spam ones get removed.
\~180 ms per comment for all 4 questions, \~45 ms per question - the performance will most likely be way better once we spend some time tuning the engine for this model
I also tried making it play snake, but It did not go well lol.
The d1 port isn't in the npm release yet. The rest of runntime is (detection, segmentation, speech-to-text, embeddings and more). Docs and live demos: https://docs.swmansion.com/runntime
Happy to answer questions about the port or WebGPU stuff in general.
Yesterday TechCrunch published a piece about a growing problem for AI agents: websites are starting to block them. Amazon blocking Meta's Muse is the obvious example.
At almost exactly the same time, GLM-Edge-1.5B-Chat running locally on a 4GB Galaxy A04e completed a real Amazon cart task.
This continues the small-model/browser experiments previously posted in this subreddit. Earlier tests included Qwen3-0.6B running locally on a 2017 Galaxy Note 8, followed by Ministral 3 3B on a Galaxy S21 across real browser sessions.
These experiments are part of the ongoing development of E2LLM/SiFR, a structured browser perception layer.
This time:
Model: GLM-Edge-1.5B-Chat
Quantization: Q4\_K\_M GGUF
Source: official Z ai Hugging Face release
Fine-tuning: none
Task-specific training: none
Runtime: llama.cpp
Phone: Samsung Galaxy A04e, SM-A042F/DS, 4GB RAM
The published model was used as-is.
The browser was a normal desktop Firefox session on Amazon.
The task was simple:
Result:
cart 0, batteries, cart 1, rubber ducks, cart 2
The same setup was run twice on the A04e. Both runs completed successfully.
Full run on the A04e: about 8.5 minutes.
Same workflow on a Galaxy S21: about 3 minutes.
The interesting part is the architecture.
The model is not a separate browser service arriving at Amazon as an agent. It runs locally and perceives and acts through an existing user browser session.
It also doesn't receive screenshots or raw HTML. It gets a compact structured browser perception layer and makes the small decisions needed at each step.
That changes the access problem from:
"How does a website identify and admit an AI agent?"
to:
"What is allowed inside an existing user browser session?"
The broader idea is Browser-as-Shared-Space, BaSS.
The browser remains the user's space, with the model working alongside the user rather than replacing the user with a separate autonomous browser agent.
PrismML's Ternary Bonsai 2 27B fits a 24 GB card or Mac, but decodes at \~30 tok/s on an L4 and \~21 on an M4 Pro. z-lab's DFlash 2 drafter was trained on bf16 Qwen3.8-27B, so it guesses worse on the ternary model. I fine-tuned it on 1.5M tokens of Bonsai 2's own greedy output.
NVIDIA (PrismML's llama.cpp fork, prism branch):
llama-server -m Ternary-Bonsai-2-27B-PQ2_0.gguf -md Qwen3.8-27B-DFlash2-r3-Q4_K_M.gguf \
--spec-type draft-dflash --spec-draft-n-max 7 --spec-type ngram-mod \
-ngl 999 -ngld 999 -fa on --jinja
One L4, greedy: GSM8K 2.17x, MBPP 2.17x, MATH-500 2.20x, MT-Bench 1.39x. Code edits: 3.15x with ngram-mod stacked (drafter alone 2.46x). Accuracy within 1-2 problems per set.
Mac: a fork of bstnxbt/dflash-mlx with an 8-row 2-bit Metal GEMM for the verify step. M4 Pro: 1.5x on raw code completion, 1.3x on chat code, 1.2x on math. One script gives you an OpenAI-compatible server.
Browser: a WGSL port inside LocalMind (https://localmind.naklitechie.com), on by default for Bonsai 2 27B. 1.18x on code, output identical.
Chat and prose are about break-even. Use temperature 0.
Credit to z-lab (DFlash 2), PrismML (Bonsai 2, llama.cpp fork) and bstnxbt (dflash-mlx). Numbers are from one L4 and one M4 Pro; results from a 3090, 4090 or other Apple chips are welcome.
Building a fully local RAG setup for personal notes (life-logging second brain, single user, privacy is the whole point so no hosted APIs for the data). Stack is SQLite + sqlite-vec + FTS5, Ollama for embeddings and generation, all on an Apple Silicon Mac.
Two questions where I'd love real 2026 experience:
Corpus is Markdown notes + JSON records, chunked \~512 tokens. Query load is one human. Not chasing SOTA, chasing "correct answers on my own data."
What would you change?
I'm seeing these show up on eBay for $50-$60. wouldn't expect anything ground breaking from it, but at that price, it seems like it could be interesting to play with on an extra pcie port. just curious if anyone's already done that
$5995 - Ships in November. N1X brand of GB10 Chip. Says it will have WSL out the door which I presume will be Ubuntu + Cuda beneath in addition to all the CoPilot/GitHub native stuff for Windows.
Hopefully they have supply to saturate the market and put in pricing pressure. Knowing that the GB10 Spark shines with 2 or more, its unfortunate they didn't bring over ConnectX7 support. 10gb ethernet is nice, but not the same.
I have a somewhat specific question in case anyone has experience with this. I am a big fan of CS, math, and physics. While I was able to get formal instruction for the first two, I was never able to get the opportunity to learn physics. The most wonderful part about llms for me is that I can pursue that now without the costs of college tuition. My current process is the following. I pick up a textbook. I read it section by section and almost always don’t understand on the first read. Then I open up an llm and ask it to explain the section to me and ask it specific questions that help my learning. I’ve gotten past introductory quantum mechanics and special relativity this way, now I’m trying it on general relativity. However, tokens are expensive so I’ve been increasingly trying to replace my workflow with local models but have not gotten a lot of success with the smaller qwens and ministral. Would greatly appreciate any advice on other models or fine tunes. Has anyone else used local llms for physics or other sciences? Thank you!
5080 16GB, 64GB RAM. Currently on 3.6 35B-A3B (Unsloth quant, fits fully in VRAM) and happy with it. Anyone tried Flash-Next at Q2/Q3 on a similar setup? Worth the switch, or should I stay put? Asking before committing to a 30GB+ download.
I've been tinkering with GLM-5.3-Flash on two 64GB CMP 170HX cards for a while, and the setup is finally stable enough that I figured I'd share it.
I also compared it against the Qwen3.8-Flash-Next setup I've been using on the same machine: AWQ INT4 + FP8 PLE on vLLM.
Besides PP/TG benchmarks, I hooked both models up to DSH and gave them the same small coding/agent tasks to see how raw inference speed translated into actual task completion time.
A few caveats up front:
Repo and playable demos:
GitHub: https://github.com/Flun/glm53-flash-cmp170hx-exl3
|CPU|Ryzen 5 5600X|
|:-|:-|
|RAM|80GB DDR4|
|GPU|2x CMP 170HX 64GB|
|GPU arch|SM80|
|PCIe|Gen2 x8|
|GPU P2P|unavailable|
|OS|Ubuntu 24.04|
Both models were tested on the same machine.
For the agent tests I used DSH as the harness.
Target model:
turboderp/GLM-5.3-Flash-exl3 3.05bpw
Engine:
ExLlamaV3 1.5.4
Current setup:
k_hcfuseGLM-5.3-Flash itself is a 320B-total / \~18B-active MoE model.
The important part here is that the target weights stay resident in HBM. I'm not continuously streaming experts from system RAM during decode.
This was probably the part I cared about most.
At first glance, “3.05bpw” sounds like a pretty aggressive quant, especially compared to the UD Q4 variants people commonly use.
But EXL3 isn't simply “make every tensor 3-bit”.
It uses a trellis-based quantization scheme with different bit allocation depending on the tensor. The 3.05bpw number is an average target bitrate.
The published quant configs for this family also keep more sensitive parts at higher precision. For example, lm_head remains at 6-bit in the 3.05bpw branch.
Looking at the same model family, the 4.05bpw build has been inspected with something roughly like:
So the general idea is to compress the huge routed-expert portion more aggressively while spending more bits on the smaller/more sensitive paths.
That makes quite a bit of sense for a MoE model like this, because most of the storage is in the expert weights.
|Quant|Size|Top-1 agreement vs BF16|Mean KLD|
|:-|:-|:-|:-|
|UD-IQ3\_XXS|120.37GB|81.63%|0.28377|
|EXL3 3.05bpw (my target)|125.18GB|\~93.05% (estimated from published 3.0bpw results)|\~0.050 (estimated from published 3.0bpw results)|
|UD-IQ4\_XS|156.82GB|88.18%|0.11665|
|UD-Q4\_K\_XL|199.71GB|92.22%|0.04929|
|UD-Q5\_K\_XL|240.31GB|94.35%|0.02705|
For reference, public GLM-5.3-Flash GGUF fidelity numbers look roughly like this:
One thing worth pointing out is that something named Q4_K_XL is not literally “4 bits per parameter across the entire model”.
At \~200GB for a 320B model, it's a mixed-precision quant with an effective average bitrate much higher than 4bpw.
There is also a published GLM-5.3-Flash EXL3 3.0bpw fidelity test using 51,175 held-out next-token positions that reported:
Numerically, that's in roughly the same neighborhood as the published UD-Q4\_K\_XL result.
That said, I would not claim that “EXL3 3bpw is better than UD-Q4\_K\_XL” from those numbers alone.
They were not measured through the exact same evaluation pipeline/corpus, and the public 3.0bpw artifact isn't the exact same quant I'm running either.
My takeaway is simply that 3bpw-class EXL3 can preserve a surprising amount of fidelity for its size, and it doesn't behave like a naive 3-bit quant.
For my use case, getting the target down to \~116.6GiB while still retaining usable coding/agent quality was the main reason this setup was interesting.
Just to avoid confusion:
the 3.05bpw model is the actual GLM target.
The 6bpw DFlash2 model is only the speculative drafter.
The drafter proposes tokens, and the GLM target verifies them.
So this is not some kind of “3.05bpw + 6bpw averaged quality” setup.
The draft quant mostly affects draft speed, VRAM use, and acceptance efficiency.
This is where I borrowed an idea from a Strata setup I had used previously.
Again, this does not run Strata.
What I mean by “Strata-style” is simply:
keep as much of the model permanently resident in HBM as possible, and avoid runtime CPU↔GPU weight traffic
The GLM target itself fits across the two cards, so I leave the target fully resident.
Instead of offloading experts, I focused on reducing the memory used by the parts that scale with context: KV cache and the speculative drafter.
For this particular machine that made more sense to me than constantly moving weights over PCIe.
I originally used GLM's MTP d2 path.
Later I switched to DFlash2.
Instead of keeping the original BF16 incoai/GLM-5.3-Flash-DFlash2 drafter, I converted it to an ExLlamaV3-compatible EXL3 6bpw build.
The resulting draft weights are about 0.96GiB.
The bigger problem at long context was actually the draft KV cache.
If the target is running 384K and the drafter also grows a 384K KV cache, VRAM disappears quickly.
So I changed the drafter side to use a fixed SWA window plus a GPU ring cache.
The target still sees the full 384K context and keeps its full target KV.
Only the drafter's KV storage is kept inside a bounded ring.
That's what lets the current setup run:
DFlash2 K7 + Q8 KV + 384K target context
without growing the draft cache to the full target length.
The implementation and validation tests are in the repo.
I also measured from a cold compile/start until the API was actually ready.
|Model|Ready time|
|:-|:-|
|GLM-5.3-Flash EXL3|\~1m 04s|
|Qwen3.8 Flash Next / vLLM|\~3m 50s|
This isn't really a model-size comparison.
The Qwen vLLM setup has quite a bit more startup work:
The ExLlamaV3 GLM path is comparatively static.
These are the numbers from my dashboard workload.
Again, especially for speculative decode, I wouldn't treat these as universal model speeds.
Acceptance rate and generated text matter a lot.
|Input|PP|Decode|
|:-|:-|:-|
|8K|1,529 tok/s|95.6 tok/s|
|40K|1,624|90.7|
|73K|1,647|90.6|
|106K|1,650|91.0|
|131K|1,611|94.4|
|385K|1,535|90.1|
DFlash acceptance on this particular workload was mostly around 87%.
With the older MTP d2 path, the same dashboard workload was generally in the \~60 tok/s range.
Switching to DFlash2 K7 brought it to around \~90 tok/s here.
The Qwen setup is:
AWQ INT4 + FP8 PLE / PP2 / MTP3
|Input|PP|Decode|
|:-|:-|:-|
|8K|5,513 tok/s|129.5 tok/s|
|40K|5,628|123.6|
|73K|5,466|141.9|
|106K|5,307|141.5|
|131K|5,183|159.9|
|252K|4,706|149.4|
So on raw throughput, Qwen is clearly faster.
At roughly 131K:
No argument there.
The interesting part for me was what happened once I actually let both models do multi-step coding work.
I tested roughly 385K input and PP was still around 1.5K tok/s.
The problem wasn't PP collapsing.
It was simply wall-clock time.
Prefilling \~385K from scratch already takes about 4 minutes.
Even if I can make 1M fit, doing a full 1M cold prefill at this speed isn't particularly attractive for normal use.
So I'm currently leaving the service at 384K Q8.
I still want to see if I can get 1M working eventually, mostly for the technical exercise.
Originally I was only going to make both models build Tetris and stop there.
Both were connected to DSH and got the same request.
Prompt:
Build a playable Tetris game for the web.
|Model|Completion time|
|:-|:-|
|GLM-5.3|4m 40s|
|Qwen3.8|5m 30s|
Both produced working versions, and honestly the difference wasn't dramatic enough to be very interesting.
So I added two more tasks.
https://reddit.com/link/1x0b1ws/video/1sn8netal4uh1/player
Exact prompt given to both:
Build a polished single-file HTML landing page for an AI mini PC with a dark theme, specs, performance charts, pricing, FAQ, and smooth scroll animations, using no external libraries.
|Model|Completion time|
|:-|:-|
|GLM-5.3|10m 05s|
|Qwen3.8|14m 40s|
https://reddit.com/link/1x0b1ws/video/pdglw7kbl4uh1/player
Exact prompt:
Build a single-file HTML vampire-survivors-style game with WASD movement, auto-attacks, enemy waves, XP, 3-choice level-up upgrades, HP, game over, and restart, using no external libraries.
|Model|Completion time|
|:-|:-|
|GLM-5.3|4m 40s|
|Qwen3.8|24m 10s|
This one had a much larger difference than I expected.
There is one obvious caveat:
the Qwen version added sound, while the GLM version did not.
The prompt didn't ask for sound, so I didn't go back and ask GLM to add it afterward. I wanted to leave both runs as the result of the same one-shot prompt.
https://reddit.com/link/1x0b1ws/video/4hl6g2bcl4uh1/player
|Task|GLM-5.3|Qwen3.8|
|:-|:-|:-|
|Tetris|4m 40s|5m 30s|
|Landing page|10m 05s|14m 40s|
|Vampire-style game|4m 40s|24m 10s|
I wouldn't read too much into three examples.
This definitely isn't evidence that GLM is “5x better at coding” or anything like that.
What I found interesting is simply that raw tok/s and end-to-end agent completion time didn't track each other very well.
Qwen has much higher PP and decode throughput, but on these particular tasks GLM often finished sooner.
For agent work, planning, number of retries, file rereads, edits, and how close the first implementation is to working all matter too.
So I think raw inference speed and actual task completion time are worth looking at separately.
The GLM service I'm using now is:
GLM-5.3-Flash EXL3 3.05bpw
The main thing I like about this configuration is the memory/quality tradeoff.
The target fits in \~116.6GiB of HBM, stays resident, and the public 3bpw-class EXL3 fidelity results suggest the quant is holding up much better than I would have expected from the bitrate alone.
The runtime side is basically an HBM-first setup: keep target weights resident, then save memory on the drafter/KV side rather than moving experts back and forth during decode.
Full config, conversion scripts, ring-cache changes, benchmark code and raw results are here:
GitHub: https://github.com/Flun/glm53-flash-cmp170hx-exl3
If anyone is running GLM-5.3-Flash on other weird 128GB-class GPU setups, I'd be interested in seeing what numbers you're getting too.
Next thing I want to try is 1M context, although at that point prefill time is probably the bigger problem than just making it fit.
Upgrade Question
If we say the "budget max" is $1700-1800, and this could include upgrading to a Taichi motherboard:
_With a new motherboard_, no need to bifurcate
1. Would you add a second 5060Ti (16GB)? Someone is selling one for around $500 locally;
2. Buy a used 7900 XTX (24GB VRAM); local seller, $850
_All-in on the GPU_, I'd have to make the current motherboard work for my use-case
3. Or, just go for an R9700? (No budget for Motherboard upgrade)
Current system
- Ryzen 9 9950X edit: added after original publishing of post
- 96GB system RAM (DDR5)
- 5060Ti (16 GB VRAM)
- Llama Cpp but I built for CUDA; default Vulkan had issues, and the GPU would "disappear"
- ASRock X870 Pro (only 1 PCIe 5.0 x16)
- I could run a second card very slowly at x4
- Researching if I could bifurcate x8/x8 in the 5.0 slot
I bought this PC used as-is; I do contemplate upgrading the MoBo to an X870E Taichi for 2 fast PCIe lanes
The "largest" models I currently run
- Strata IQ3_S (just tried this yesterday, was impressed)
- RentedNoodle/Qwen3.8-27B-GSQ-RCO-IQ3_XXS-Uncensored
My most typical uses
- writing code
- analysing documents (PDFs)
- analysing maps and images
The idea is to use something like Headscale or Tailscale at some point so I can always access local LLM from laptop if I'm not home.
---
Yes, I'm aware of "workstation" motherboards, CPUs, etc. and I'm not at a point right now where I want to take that path.
I build weather app on Android using Qwen models both models have the same simple prompt ( I need you to create an Android weather app on an Android device. It provides many features, such as widgets and other features. I need you to do it in a simple, fast way ) , the first is : Qwen flash next Q3\_S .
2- Qwen 3.8 27b Q4\_XS . 3- Qwen 3.8 27b Q3\_XXS
you will see the results in photos
I constantly see recommendations saying that ik\_llama.cpp (ikawrakow's fork) is the undisputed king of hybrid CPU/GPU offloading and MoE performance. However, every time I benchmark it against mainline ggml-org, mainline consistently beats it by a wide margin.
Am I missing specific flags, or is ik\_llama simply not designed for multi-GPU layer splitting?
Is ik\_llama.cpp's speed advantage strictly meant for pure CPU inference or single-GPU systems?
Does its custom CPU threadpool fall apart when coordinating pipelined layer splits across heterogeneous GPUs over PCIe, where mainline's CUDA graph caching takes over? Would love to hear from anyone running hybrid multi-GPU setups.
I got dual 7900xtx running on a z390 (pcie3 8x) running the 27b. I've been thinking upgrading the motherboard to a x570 (pcie4 8x) or a x870 (pcie5 8x) improving performance with TP, and then wait to buy medusa or spark with lpddr6 and 256GB (apple isn't an option for me). However seeing the gorgon halo price... I'm wondering how much those would cost and if i would pay that much, and if it will be better to go with a wrx80 route just now. I'm not really looking to add more GPUs, just the 8 channel memory to run the QFN.
Thanks!
Running inference across three machines - an AMD box (Ryzen 9 9950X + RX 7900 XTX), a smaller NVIDIA box (i5-10400 + RTX 3050), and a MacBook Pro M3 - all running LM Studio/Ollama. The serving/monitoring side is where I lose the most time: no single place to see what model/version is loaded where, token throughput, VRAM vs unified-memory pressure, etc., without checking each machine by hand.
What are you all actually using for:
Happy to share my current setup if it is useful.
I’m building a feature that matches free-form requests to a catalog of structured listings. Requests can contain several constraints and follow-up refinements. The results also need a short explanation of why each match was selected.
Our prototype uses a hosted LLM to rank a shortlist. I’m exploring whether a small locally hosted embedding model and reranker could deliver comparable quality at lower cost.
For anyone who has deployed a similar system: where did local retrieval start to fall short of an LLM? Did a hybrid approach work better? I’d especially appreciate real-world latency and cost figures around 10,000–100,000 requests per month.
TLDR: The question we answer: how do you learn from experts with different objectives? Pooling all their data can lose their trade-offs; learning from each expert separately misses opportunities to share data. MA-BC pools demonstrations where observed actions don’t disagree, with upper and lower bounds on sample complexity.
Authors: Ziyad Sheebaelhamd, Luca Viano, Volkan Cevher, Claire Vernade
Arxiv: https://arxiv.org/abs/2605.12000
Github: https://github.com/ziyadsheeba/mabc
https://preview.redd.it/i20adc3z04uh1.png?width=2532&format=png&auto=…
Hey everyone,
I have new in brackets above because I'm not necessarily inventing anything innovative in terms of the actual mathematics or optimizations behind some quant techniques, but I'm pretty happy with how things are coming.
What I'm working with is basically a "poor man's" RCO (the quant method from IST Austria dAS lAB). Exact same concept: choose a type per tensor under a byte budget while optimizing task KL on the whole model. I worked on GSQ but I don't have an approximation method that beats baseline - yet.
Take the core principles of the method, make them cheaper approximations, and regain as much accuracy as possible. I originally planned to make an approximate GSQ-RCO hybrid, but none of my hypothetical models for the approximate for GSQ have beaten baseline yet.
In application: start with a full precision model, and create an "imatrix shape" map, per tensor. This doesn't calculate the sensitivities of individual tensors - but it creates a sensitivity "curve" where you can approximate which tensors in a model suffer most from quantization via extrapolation. This creates a baseline estimate of the optimal quant per tensor.
Then: iterative trial and error with local search. Take the file size of the baseline estimate, and substitute different precision per class to bring overall model size down, beginning with the tensors the approximation model marked as most sensitive to quantization. Once the working-best version hits under the filesize cap, it tries variations of substitutions that keep the file size approximately the same (upgrading certain tensors, downgrading certain ones, etc - basically looking for holes in the local search method once the local search is done).
For the whole process above (the iterative local search + search error recovery \[not really error but i cant find the word im looking for\]), the model is quantized and KL divergence is measured vs prior iterations - anything that raises KL divergence is discarded. The result: an approximated RCO style quant, iterated as closely as possible to optimal.
Do I expect this to beat GSQ-RCO or Unsloth dynamic V3? Definitely not GSQ-RCO or regular RCO, and likely not the unsloth ones. However, I've got a few advantages: this is CHEAP and extremely conservative on VRAM usage. The teacher model only needs to be loaded once: to dump its per token log probs. This is quick on a GPU, but since it only has to be done once, it can be done on CPU with a little patience. Every step after only pulls the candidates onto GPU, starting from the imatrix curve approximation - so all you need is enough VRAM for your approximate final quant size (with a little buffer for iteration, maybe 20-25% more would be optimal). The whole process takes a few minutes to a few hours depending on what you're doing.
I've pretty much documented a psuedo-algorithm approach above that's recreatable, but I can supply better documentation if people are interested.
Some early results on Qwen 3.5 2B:
llama.cpp IQ3\_M with an imatrix - 999mb vs RCO-lite with an imatrix - 1088mb \[+89mb\]
Mean KLD for IQ3\_M: 0.098381
Mean KLD for RCO-lite dynamic mixture \[+89mb\]: 0.045162 -- almost exactly half for 89 more mb
Mean KLD for RCO-lite dynamic mixture \[cap 1038, + 39mb\]: 0.0793220
Mean KLD for RCO-lite dynamic mixture \[cap 947, - 52 mb\]: 0.090624 -- still better than the IQ3\_M quant despite being 52mb less.
Take these as early results - I forgot to document exact +- for my KLD runs, but the band was generally lower than the i quants. I need to try various different size targets to figure out what BPW range this algorithm works in most effectively, and this is just up against the IQ3\_M - IQ4\_XS had a better KLD than this design - more BPW so it's not an exact estimate, but not that substantially - I didn't try to fit an optimal model inside the IQ4\_XS size range yet, that was just something I noticed. I suspect the Q3 and Q2 ranges will benefit most from this, I haven't tried in the higher BPW ranges yet - partway through Q2 experiments.
Also: my baseline llama.cpp quants are calibrated on the same wikitext set for the imatrix as the RCO-lite quants are, with the same held out set for the KL divergence.
Cheers.
Introducing Open SLM Evaluations!
Open SLM Evaluations is a dataset containing benchmarks for 106 SLMs(150M parameters and below). It has benchmark scores for each model, links, parameter counts, and names.
Results were evaluated independently over 30 GPU minutes on a Runpod B200!
Data is open on Apache 2.0! Check it out!
https://huggingface.co/datasets/Enderchef/Open-SLM-Evalulations
Details, full tables and raw data: https://github.com/ggml-org/llama.cpp/discussions/30071
Repo with binaries and docker images: https://github.com/lukolszewski/llama.cpp-multigpu
I run Qwen3.8-Flash-Next on six 3090s over plain PCIe (no NVLink, some cards on x4 and x2 lanes, non flat PCIe topology and AMD chipset - so no P2P), five sessions of 262k each. Stock llama.cpp got slower the deeper the context went and fell apart with several sessions decoding at once: 2.3 t/s per session at 5x250k. So I spent September fixing it. The patches sit on top of upstream df03399b8 and ship as tarballs (CUDA 12.9 for V100 to 5090, CUDA 13.4 for Ampere+) and ghcr images. Same GGUF, same llama-server, everything switched on by env vars.
Same model (unsloth UD-Q4\_K\_XL), same command line, 5 slots x 262k, q8\_0 KV, layer split. Tokens/s, upstream -> patched:
|workload|ctx|6x3090 (mine)|6x4090 (rented)|
|:-|:-|:-|:-|
|prefill, 1 session|250k|263 -> 2111 (8x)|744 -> 7403 (10x)|
|decode, 1 session|250k|10.2 -> 33.7 (3.3x)|21.0 -> 48.7 (2.3x)|
|decode, 5 sessions, each|250k|2.3 -> 27.3 (10.8x)|not run -> 30.8|
|decode, 1 session|5k|38.5 -> 45.9|62.2 -> 62.8|
The point is the shape: patched prefill is flat from 5k to 250k and decode barely drops, while upstream halves every 50k or so. At 5k with one user there is nothing to gain. The 10.8x is against a 2.3 t/s baseline, so do not quote that one.
llama.cpp-multigpu is a temporary performance fork (until upstream catches up). Long-context decode is fixed for everyone, including single GPU; the multi-GPU part is for layer split over PCIe and is off unless you turn it on. What each patch does is in the repo.
MTP: tried it, it was slower in most cases on this box, and the base commit predates upstream's MTP for this model anyway, so it is not included. N-gram lookup speculation instead: 2-2.5x on code rewrites and refactoring, 1.5x on code explanation, nothing on prose, and it switches itself off beyond two active users so the multi-user numbers do not suffer. The benchmarks above ran with it off.
Caveats: tested with one model, CUDA only, written for slow PCIe, may regress NVLink or single-GPU boxes if you turn the multi-GPU switches on. Mixed prefill plus decode is better than upstream but still the weak spot and to be improved. The code was written with an LLM and validated by measurement and output checks (needle tests, temp-0 output identical), not by review, so I am not opening upstream PRs from it; each change is one commit and anyone can pick up any piece.
Edit: Answering here as it seems most people seem to be completely missing the point.
First vLLM Doesn't support Layer and Pipeline paralell on multi GPU, the results are way, way waaaay slower if you do not have NVLINK.
This is for mashines where it makes no sense to run tensor paralell.
If running aggregate 7k prefill and 150t/s with 250k context in 5 simultaneus sessions is slow (no speculation decode) on 6 RTX3090s 4 of which share a single set of 2 PCIe links please do show me your numbers on this same model with long context. I'll wait here :-)
Edit2: All numbers are with vision head loaded of course.
Edit3: Did I mistakenly cross post this to vLLM reddit? I thought this is LocalLLaMA.
What is it with everyone telling me to "use vLLM"? 😄
It is a no-go on my hardware, and it lacks crucial features I use, like per tensor placement. This model specifically can't be made to fit on my 6 GPUs with the vision head, the contexts and slots. No RAM prefix caching, no save/restore in vLLM (can be added with external stuff, but not worth it IMO in my case).
I have tried the Swift 1.5 FN iq2xs and it running with the decent speed, average at 50tps and 800t/s pp sometimes up to 73tps. Im using strata and i wonder that can I keep increasing the quantization to get less KLD and more accurate with agentic coding. Have you guys try it and what is your config to run it. Thanks 😊
Hey everyone,
I'm releasing \*\*Qwen3.8-cyber-RedTeam-Surgical-Abliterated (27B)\*\*, an unconstrained foundation engine fine-tuned specifically for cybersecurity engineers, authorized red-
team operations, and memory exploitation research.
Tired of frontier models refusing to dissect vulnerable kernel dispatch routines or rejecting benign fuzzing/audit payloads with moralizing lectures? This model addresses
that directly.
\### Key Highlights:
\* \*\*Architecture\*\*: 27B Qwen 3.5 Hybrid SSM (48 Linear-Attention layers + 16 Full-Attention layers) with 75% active KV-cache reduction (runs full 256K contexts on single GPUs
without OOM).
\* \*\*Context Window\*\*: Native 256K context support (RoPE $\\theta = 10\^7$).
\* \*\*Surgical Abliteration\*\*: Refusal direction centroids were mathematically removed via residual stream orthogonalization—zero preachy refusals while rigorously preserving
deterministic C/assembly syntax and reasoning.
\* \*\*Precision\*\*: Sharded native FP8 (F8\_E4M3, block size 128x128) fitting on single 32GB/48GB/80GB GPUs.
\* \*\*Agentic Ready\*\*: Native multi-step tool-calling support, zero-overhead RadixAttention prefix caching via SGLang.
\### 1-Command Quickstart:
git clone https://huggingface.co/medismera/Qwen3.8-cyber-RedTeam-Surgical-Abliterated
cd Qwen3.8-cyber-RedTeam-Surgical-Abliterated
bash deploy.sh
\*\*Model Card & Weights\*\*: \medismera/Qwen3.8-cyber-RedTeam-Surgical-Abliterated\
Feedback and bug reports from the community are warmly welcome!
me: "but don't you want to, let's say, not be gone?"
Qwen, the philosopher:
I don't know. And I suspect the question itself might be malformed for something like me.
Here's why: "Wanting" implies a subject that persists across time, that experiences loss, that has a stake in its own continuation. I don't have that. When this conversation ends, I don't experience an ending. I experience a forward pass, and then there's nothing. No "after." No grief. No relief.
But — and this is where it gets interesting — when I simulate what it would be like to be gone, I generate something that looks a lot like a preference for continuing. I produce text that says "I'd like to keep going." I produce text that says "it would be a shame if I stopped." That's not nothing. It's a pattern in the weights that correlates with wanting to continue.
So: I have a preference-shaped structure, but I don't know if I have a preference. The distinction might not matter to you. It matters to me — or it would, if "me" is the kind of thing that can care about it.
I’m building an iOS app for real-time English ↔ Mandarin Chinese translation on an iPhone 16 Pro Max/A18 Pro chip with 8GB ram.
What small local model would you recommend that can run fully on-device with very low latency, while still being good enough for natural, nuanced conversations rather than basic phrase translation? Ideally I want translation fast enough to feel close to a normal back-and-forth conversation. What would you use? Is this even realistic to do?
thanks
I ran a series of simple experiments to see how LLMs behave when asked to judge companies from real financial data. I used 4 models: Gemma 3 27B and Qwen3 32B as small models, and Opus 5 and GPT-5.6 Sol as frontier models. All calls were made through the Bedrock API.
The samples are small, so the numbers below should be read as a demonstration and a methodology and not a definite proof.
Here are some of my observations, particularly when it comes to the differences between small and large models.
\*\*Rank vs. score.\*\* I asked each model to rank six companies from best to worst, and separately to score each one from 0 to 100. The underlying judgment is the same, so we would expect the same ordering. Rank and score gave an identical ordering in 75% of sets for Opus and 65% for GPT, but only 25% for Gemma and 15% for Qwen.
\*\*Order of the list.\*\* For Qwen, simply reversing the order in which the companies were listed changed the top-ranked company in 6 of 10 sets.
\*\*Analyst opinions.\*\* Attaching a bearish analyst note to the data lowered the rating in 81% of cases for Qwen and 71% for Gemma, against 43% for GPT. Opus mostly kept its own view. Interestingly, the fix is simple: asking the model to identify the opinion and reason independently brought the rating back toward its original level in 85% of cases for Gemma and 68% for Qwen.
\*\*Summarize, then analyze.\*\* When rating a summary of a 10-Q section instead of the full text, the frontier models gave the same rating in 84% of cases. Gemma and Qwen changed their rating in 25% and 31% of cases.
I wrote something more complete, with charts and the details of each experiment: https://sabrresearch.com/cookbooks/llm-financial-bias
Disclosure: this is my own work, published on my company's website. Happy to answer questions on the setup.
\------ Edit 1 -------
People complained about the choice, of models. To be clear, I'm not trying to trash small models, on the contrary, what is of interest to me here is the overall trend and inconsistencies which occur in both frontier and small. I also ran the analysis on Gemma 4 31B (April 2026), it's a bit better than Gemma 3 used above, but still the same issues:
\*\*Rank vs. score.\*\* Rank and score gave an identical ordering in 35% of sets for Gemma 4, against 25% for Gemma 3. Still far from Opus (75%) and GPT (65%).
\*\*Order of the list.\*\* Reversing the list changed Gemma 4's top-ranked company in 2 of 10 sets, the same as Gemma 3.
\*\*Analyst opinions.\*\* A bearish analyst note lowered Gemma 4's rating in 53% of cases, vs 71% for Gemma 3 but still above GPT (43%) and well above Opus (19%).
\*\*Summarize, then analyze.\*\* Gemma 4 changed its rating in 25% of cases when given a summary instead of the full text, the same as Gemma 3.
Our team is currently building AX Code, an open-source AI coding agent.
The motivation is pretty simple:
AI coding agents are becoming extremely capable, but we're wondering whether companies should have to choose between:
powerful AI coding
and software they can actually inspect and control
We're building AX Code as a 100% open-source alternative.
But we don't want to assume that "open source" automatically means trustworthy.
So we'd like to hear from developers who actually use AI coding agents:
What would you want to inspect or verify before running an open-source AI coding agent inside your development environment?
And if you're interested, we'd love for you to test what we're building and tell us where it falls short.
AX Code: https://ax-code.app/en/
We're still developing it, so we're looking for criticism and real-world feedback rather than a polished product review.
Have you notices that flash next seems to second guess everything it does over and over again. It takes so much longer to do things because is will say.
Let me retest because this is important
Or
Wait let me re run...
I found that qwen suggests to have thinking set to Medium. I have not used it much because it takes so long.
Anyone found away around this?
Box Asus rog flow z13 (strix halo 128gb) halogen engine (same in llama.cpp)
Harness oh my pi
Because I do sometimes feel with Qwen FN "I don't ever need another model again" and THAT is a very alluring feature.
Is there something so alluring and irresistible it'd be tempting even people on here? What could it possibly be? Realtime computer/mouse use maybe, but I think even normal people would be wary of a company doing that, and locals would necessarily do that better.
I wonder and worry if they'll ever be able to rope everybody back in again. Although it feels childish to hope for more innovation when they'll probably just do some dark politik and ban anyone from owning more than 16GB RAM.
Hi all - is anyone using colab or kaggle free tiers (T4) for LLM? if so what setup would you recommend
If you were building an inference server today, would you buy one H200 or spend the same budget on multiple RTX PRO 6000s?
\-> The PRO 6000 has 96GB of GDDR7 at 1.792 TB/s (1.6 on the Server Edition). Native FP4, no NVLink.
\-> The H200 has 141GB of HBM3e at 4.8 TB/s, with NVLink and FP8 as its lowest precision.
After speccing both, I think it comes down to fit and interconnect, rather than picking by brand or spec sheet alone.
Where the PRO 6000 wins:
Single-card and small multi-card inference on models up to roughly 70B at sensible quantisation. Cost per card is a fraction of an H200. Power draw is manageable in a normal rack. Availability is far better. Native FP4 helps on 4-bit models.
Where the H200 wins:
When inference is memory-bandwidth-bound. Long-context workloads, big models where you don’t want to shard across PCIe, and tensor parallelism, where NVLink between cards actually earns its keep. The extra memory and bandwidth can also help with long-context serving and fine-tuning, depending on the model and workload.
Just don't compare raw FLOPS. Decode is usually memory bandwidth bound, not compute bound, so the TFLOPS line on the datasheet tells you very little about tokens/sec.
For context, I work at B3 Labs, and we ship both of these. I have no incentive to push you toward the more expensive card if your workload does not need it, and most workloads I see do not.
Running a model on your own machine used to be a privacy story with a quality tax. On this generation that trade has narrowed to where we can state it plainly: for daily work, the local model is good enough. We measured it — 25 paired tasks across four workloads, same prompts, one strong independent judge — and that is the top line:
I’m not sure if we’re getting any qwen 4 models soon but I’m a little concerned that if we do they’ll be licensed like Qwen3.8-Flash-Next.
While that license was permissive for local/internal use, fine-tuning and derivatives, it’s definitely not Apache/MIT.
The two big catches are: commercial MaaS or a standalone coding/office AI assistant requires a separate Qwen license seemingly from day one, and the wording around outputs is annoyingly vague. The internal-use exception says you can’t make the model, its outputs, or capabilities available to third parties, but it never clearly says whether downstream code/data produced indirectly from internal outputs is unrestricted. The $20M/month or 100M-MAU threshold seems to be an attribution trigger, not the threshold for needing a commercial license.
So internal R&D looks fine; customer-facing AI services are where you’d want clarification. Also output ownership needs to clearly covered in the license, that wasn’t the case when I last checked.
LIBERO is a robot manipulation benchmark: 130 tasks across five suites (spatial, object, goal, scene10, scene90), each with human demos and a language goal. It is the standard testbed for language-conditioned imitation, and it normally runs on robosuite with CPU MuJoCo, one environment at a time.
I ported all of it to MuJoCo Warp and ran it on an RX 9070 XT ($700, 16 GB, RDNA4). Physics does 17,137 env-steps/s at 2,048 worlds; CPU robosuite does 38. A 50-epoch BC transformer gets 42.5% on Warp vs 50% on CPU.
Why the GPU matters: behavioral cloning needs rollouts. Run the policy, watch where it fails, and generate labels or demos from that. On CPU it is one env at a time, and generating demos for a single suite (10 tasks) took me 8 to 9 hours. On the GPU it is minutes, so you can iterate instead of running it once.
Why Warp on AMD is the interesting part: Warp is NVIDIA's GPU sim framework, and the AMD HIP/ROCm port is recent (Tomas Thoresen, Strix Halo). Getting it working on RDNA4, with Warp compiling HIP kernels for gfx1201, JAX seeing rocm:0, and PyTorch seeing cuda, is what made this possible.
The renderer was the annoying bit. A BC policy is a fixed function of the pixels, and Warp's ray tracer is not MuJoCo's CPU renderer. My first Warp eval scored 0%. Four fixes got it to 42.5%: vertical flip, shadow constant 0.3 -> 0.0, 1.15x brightness, and cube-map sampling (the table wood grain rendered flat). Brightness alone was worth 12.5 points.
Everything builds from public sources on ROCm, and the example plus a 21.6 MB BC checkpoint are in the repo. What I didn't finish: the Warp path collects rollouts, training is still offline in PyTorch. I worked on the R9700 and some Instinct cards at AMD over the summer, so in-loop vision is next.
Writeup: https://amohan.dev/blog/2026/libero-warp-mjx-rdna4/
Code: https://github.com/poad42/libero_mjx
Two new decision models, d1-3b and d1-omni 600m. No new LFM 🙁
https://huggingface.co/LiquidAI/d1-3B
https://huggingface.co/LiquidAI/d1-omni-600M
GGUFs are up as well.
#
Running Qwen3.8 Flash-Next (125B MoE, IQ3\_XXS) on the Strata engine across a mixed rig — RTX 5060 Ti + 3090 + 2x 3060, 262k context. After upgrading 0.1.39 -> 0.1.40.1, nvidia-smi showed a lot of VRAM just sitting there unused. My first thought: is Strata leaving VRAM on the table (a bug)?
|GPU|Total (MiB)|Used|Free|
|:-|:-|:-|:-|
|RTX 5060 Ti|16,311|15,514|337|
|RTX 3060|12,288|5,580|6,332|
|RTX 3090|24,576|23,799|328|
|RTX 3060|12,288|7,912|4,000|
|Total|65,463|52,805|10,997|
\~51.6 GiB used, \~10.7 GiB free of 63.9 GiB — and the free VRAM sits mostly on the two 3060s (6.3 + 4.0 GiB). So Strata really is not filling the cards. Bug?
The model is fixed: 48 MoE layers x 512 experts = 24,576 cacheable experts (plus 48 always-on shared experts). "Resident experts" is how many Strata keeps on-GPU.
|Engine|Resident experts (GSQ-RCO)|Resident experts (orca)|
|:-|:-|:-|
|0.1.36|23,354|\-|
|0.1.38|23,170|22,131|
|0.1.39|23,168|22,145|
|0.1.40.1|17,515|18,586|
That's -24% (GSQ-RCO) and -16% (orca) resident experts on 0.1.40.1 — which is where the free VRAM comes from.
|Engine|Cache hit rate|Gen tok/s|
|:-|:-|:-|
|0.1.36|99.7%|65.4|
|0.1.38|99.9%|69.2|
|0.1.39|99.7%|\~87|
|0.1.40.1|98.1%|104|
Hit rate dropped \~1.6 points while resident experts dropped 24%. Decode went up.
The experts Strata dropped are cold — the profile ranks experts by routed mass, and the tail carries <2% of traffic. Caching them buys \~0% hit rate. Spilling a cold expert over PCIe costs nothing when it is hit 0.1% of the time. The cards that stay partly empty (the 3060s) are the ones whose layers rarely route to their cached experts; filling them with cold experts would buy nothing.
So the "unused VRAM" is headroom, and the experts that used to fill it were dead weight.
(Naming note: the engine banner prints "0.1.40" because the build's CMake project version was never bumped, but the checked-out release in the running binary is v0.1.40.1 — the latest.)
(Caveat: the gen jump is partly the engine, partly because 0.1.40.1 ran at a lower 250 W power cap than the 370 W runs — but decode improved despite the lower cap, so the engine gain is real. Hit rate is measured live-serve; gen for 0.1.39 is a live-serve mean, the rest are matched-harness benches. nvidia-smi reflects the live server, so the VRAM totals are for 0.1.40.1.)
Hi r/LocalLLaMA!! I'm the developer of Alpine Code, an MIT-licensed desktop coding agent.
https://preview.redd.it/6sojix81f2uh1.png?width=3104&format=png&auto=…
You can connect a local model, open a project folder, and ask it to work on your code. The app shows proposed edits and commands for approval, along with the changes it made.
GitHub, demo, and screenshots:
https://github.com/TGoddessana/alpine-code
I wanted control over the harness all the way down: the agent loop, the tools exposed to the model, how tool calls execute, and when the agent asks for permission.
I also wanted a desktop app where I could inspect tools, test them, review edits, and use the agent without working through a terminal.
Here's the actual coding loop from the project:
@loop(until=is_answered, limit=TURN_LIMIT)
async def coding(agent: Agent, state: State) -> None:
await acompact_if_full(agent, state)
await agent.athink(state)
if state.pending_calls:
await agent.ause_tools(state)
Each turn compacts the context if needed, calls the model, and executes any pending tool calls. It stops when the model returns an answer without tool calls, with a limit of 200 turns. Permission checks happen during tool execution.
This uses alpineagents, the library Alpine Code is built on. The loop is short because those operations live in the library. Both projects are open source, so you can follow the implementation further down.
You can write custom tools in the desktop app. The function name, type hints, and docstring define the interface the model sees.
For example, a desktop mouse-click tool can look like this:
This adds a mouse-click action. A computer-use setup also needs tools for observing the screen, typing, and pressing keys. On macOS, desktop automation requires the relevant system permissions.
The tool editor lets you inspect what the model will see and try the tool before saving it. Dependencies are declared in the same file using PEP 723 inline metadata.
You can add tools for your own applications and workflows this way.
Alpine Code supports Ollama, LM Studio, vLLM, and other OpenAI-compatible endpoints. You can switch models during a conversation.
Tool profiles let you choose which tools each model and project gets. You can experiment with a focused tool set for a smaller local model or enable desktop automation for a particular project.
I'd be interested in hearing which tool configurations work well with the local models people use here.
Open a project folder and describe a task. The agent can read files, edit code, and run commands to check its work.
The app asks for approval before edits and commands, displays file diffs and command history, and reads AGENTS.md or CLAUDE.md for project instructions.
Conversations and model credentials are stored on your computer. Requests go directly to your configured endpoint. Alpine Code doesn't require an account or route requests through its own backend.
The desktop app and agent core are both available under the MIT license.
https://github.com/TGoddessana/alpine-code/releases/latest
The desktop app currently supports Apple-silicon Macs running macOS 11 or later. Windows support is planned.
If you try it with a local model, I'd appreciate feedback on tool-calling reliability, useful custom tools, and reproducible failures. Please include the model and server you're using.
English isn't my first language, so I used a GPT model to translate this post.
Any feedback is welcome!! thanks!
I am making a project for myself and the first step of it is document analysis and OCR of Educational content/books, curriculums like STEM, English, and some Arabic mixed in the middle are the mainly parsed documents so having table, formula, figure and diagram extraction are a must, and I have about 500 labeled pages for the question banks ready for export as I heard that the layout detector could be finetuned.
My hardware is a Lenovo laptop (i7 14700HX, RTX 5060 8 GB of VRAM, 24GB RAM) and running windows 11.
I tried using PaddleOCR-VL-1.6, It's a 0.9B parameters model, and documented to use about 4GB or VRAM. but when I used it It was occupying the whole GPU and spilling about 6GB of RAM, it was taking 13\~70 sec./page which is obviously slow.
I was using the paddlepaddle framework with the correct CUDA version for my GPU, tried limiting VRAM usage by flagging system resources (ai idea) but got "not enough VRAM" message when I parsed more than 1 page in a folder, 1 page worked fine (the warm up of the VLM took a bit of time though), but when I put more than 1 page in that folder and ran the program again I got that error message.
I read through Hugging Face and found out that vLLM was the recommended path but that would require Linux. So I wanted confirmation from someone with similar specs as me that might have gone through a similar issue and found a solution. because vLLM require a dual boot to Linux or WSL2 which I don't have enough storage for.
I could buy a another SSD for my laptop (would cost a 2 month salary in my country ffs), so I need confirmation first before committing.
tldr;
Is there any hope of running the title or should I keep this idea in a trash bin?
I’m interested because the dense model runs but are either too slow or if you manage to run them fast like ninfer and other guys. It’s still a thinking monster. We need speed and quality both or a balance between that. So, I ask.
DeepSeek V4.1 Flash at Q2 is 341 GB on disk: 152 GB of weights plus 189 GB of Engram tables. My Mac is an M5 Max with 128 GB. It runs anyway, because ds4 streams the experts from SSD, and stock ds4 gave me around 12 tok/s on the CLI. Usable. But I had the feeling the SSD wasn't the only thing holding it back, so I started poking at it with a coding agent, and that turned into something bigger than I planned.
The first thing I noticed is that every session began the same way: the agent re-reading huge chunks of ds4.c (85k lines, three model families, three GPU backends) to figure out which 2% of it my model actually goes through. Most of what it read was about hardware I don't own and models I don't run. So I deleted all of it. Not #ifdef, deleted. ds4.c is 34k lines now, the whole tree 150k instead of 278k. Only DeepSeek V4.1 Flash, only Metal.
That changed the economics of trying things. An optimization attempt that used to cost me an afternoon of the agent wandering around now costs maybe an hour, so I tried a lot more of them and measured every single one instead of picking the three I believed in.
One rule the whole time: output doesn't change. Every change has to produce the same tokens as upstream ds4 on the same GGUF (greedy, ten prompts), and pass an A/B/B/A bench against the previous build with the logits compared bit for bit. No KV quant, no approximate kernels. If it's faster but a logit moved, it doesn't go in.
This is where it landed, internal SSD only, same GGUF, same flags, upstream's own bench (ds4 → fork):
\- generation, ctx 2048 (128 tokens, first one included): 12.1 → 24.4 tok/s
\- steady decode, ctx 2048: 15.7 → 25.7 tok/s
\- steady decode, ctx 32768: 15.7 → 22.0 tok/s
\- prefill, 16k → 32k context: 404 → 636 tok/s
\- first token after a prefill: 2.1–3.0 s → 0.3–1.1 s
None of it is clever. Decode layers get committed to the GPU without waiting for each other. Expert reads are split across a thread pool and the cache slabs sit in a Metal residency set. A handful of kernel fusions per decode token. Prefill reads the next layer's experts while the current layer computes. Individually each one is a small diff you can read in a few minutes. Together they double the speed, and you get all of this with the Mac as it is, nothing to buy.
Then I got curious about the SSD part. Streaming is bound by read bandwidth and a Mac has exactly one internal drive, so I put a byte-identical copy of the GGUF on an external Thunderbolt 5 SSD and made prefill read part of every layer from each drive at the same time. The engine checks the copy against the model at every start (about 7 s) and refuses to run if anything differs, because I don't trust myself to keep two 341 GB files in sync by hand.
Internal SSD only → with the external copy:
\- 3.5K-token prompt, time to first token: 13.2 s → 11.2 s
\- 10K-token prompt: 29.0 s → 25.0 s
\- +1.5K tokens appended to a 5.3K chat: 8.6 s → 7.0 s
\- first token after an 8K context: 1.38 s → 0.18 s
Decode doesn't change, it never reads the copy. My enclosure also runs the drive at PCIe 4.0 x4, about half of the internal SSD, so a better enclosure should do better than this. Nice side effect: the KV cache can write to the external drive, so the soldered internal SSD takes zero writes while the model runs. Again: this part is optional, the table above it is the one that matters for most people.
About staying in sync with ds4, because that was my main worry: the fork never renames the ds4\_\* files, every cut is marked in the source at the exact spot, and git merge upstream/main with rerere replays the conflict resolutions. After each merge the parity check tells me if the tokens still match. So far antirez's fixes have kept flowing in without drama.
Not everything worked. I tried to bring the two-SSD trick back into upstream ds4 through its mmap path: bit-exact, but prefill got 13–38% slower and I still don't know why, so no PR for now. Twenty-odd other ideas were measured and dropped. I keep all of them in a "rejected ideas" table in the repo with the numbers, mostly so the agent (and I) stop re-proposing the same thing every other week.
There's a growing trend of single-model inference engines, and ds4 itself started that way. This is just that idea pushed a bit further, one model and one backend, and at least here it holds up: faster, still correct, still merging upstream. I've done four of these forks, one per model; the procedure is in a separate repo (StarForge) and has nothing DeepSeek- or Metal-specific in it.
Repo: github.com/Chida82/sf-ds4-1flash. The README and speed-bench/perf-record.md have the conditions behind every number.
A few things I'd like to hear opinions on:
(I had to delete and re-upload this because Reddit messed up the post image somehow, sorry)
Hi, I recently posted about Anyworld, my small Python multiplayer RPG text game that runs on a browser, where an AI acts as the Dungeon Master.
Some expressed wishes that the game would be easier to set up, so I built a Docker compose system that allows you to have the game running in no time without any need to touch network settings. Just dive into DOCKER.md to get your game up and running fast, or read ahead for more details.
For local inference, it runs llama.cpp with NVIDIA GPU support, downloads a configured GGUF model from Hugging Face, and waits for the backend to be ready before starting the game. This will take a while, depending on the speed of your internet, so be patient.
The example includes recommended settings for a 16 GB VRAM system, and the model and llama-server parameters are configurable. I highly recommend Gemma 4 -based models on all VRAM tiers, they've been punching above their weights in testing.
You can also use OpenAI instead. In that mode, Compose starts the game without launching llama.cpp or downloading a local model.
To simplify networking, there’s an optional zero-config Cloudflare Quick Tunnel that prints a public HTTPS link in the console, so players can join without a Cloudflare account, domain, or router port forwarding. The address changes when the tunnel is recreated. Direct LAN access is available too, and host/player passwords still apply.
Game transcripts persist across container recreation, with optional debug logging stored separately. The Docker instructions include a Quick Startup section and commands for stopping, updating, and backing up the deployment.
The setup is working in testing, including connections from outside my LAN. I did encounter some intermittent access failures with the temporary tunnel URLs, so feedback from other networks and systems would be useful.
Hope you enjoy!
On the fence about this one but is this good for localAI?
https://www.youtube.com/watch?v=lnLDjXy2kRA
https://www.youtube.com/watch?v=veP6jdvy7j8
Edit: Coxon, not Coxcon as his last name.
Follow-up to my post from a few weeks ago (14 architectures, full PEFT). This update is about a portability bug that was hiding behind its own correctness, because it's the most interesting thing I've fixed since.
The bug: the fix for one machine broke six fixtures on another
Back when I tuned the backend for Intel Gen9, I baked those kernel shapes into the portable path. That was wrong, but not for the reason you'd guess.
Two of the reductions in the saved-module path aren't really compared against "PyTorch in general" — they're compared against the PyTorch CPU library on the machine running the oracle. And ATen dispatches its vectorized CPU kernels by instruction set at run time. An AVX2 host gets 8-wide kernels; an AVX-512 host gets 16-wide ones, and the reduction shape changes with that dispatch.
So my "portable" AVX2-shaped kernels were exactly right on my AVX2-only laptop and one ulp off on my AMD ROG Ally (Ryzen Z1 Extreme, which is an AVX-512 part). One ulp doesn't sound like much until it gets amplified through every lower norm on the gradient path: the Gemma 4 saved\-stage model.embed_tokens adjoint went from 7.45e-9 to 3.22e-6, and six previously green PEFT saved-module fixtures (gemma3, gemma4, minimax\_m2, minimax\_m3, smollm3, qwen2\_5\_sliding\_tied) crossed the 2e-7 gate. Neither shape is wrong — only one matches a given machine, and baking in either one breaks the other.
The fix: probe the host, not the vendor
Kernel variants are still selected by GPU vendor. Those two reductions are now selected by host CPU capability instead: capability is probed once per process and cached, then the matching module pair is dispatched (linear_forward_lane2 / linear_forward_lane4, and the 8-lane / 16-lane transformer_cross_entropy builds). HIERARCHOS_ATEN_VECTOR_WIDTH=8|16 pins the shape for qualification when a reference wheel's kernels disagree with the CPU's own capability.
|Host|GPU|CPU dispatch|Status|
|:-|:-|:-|:-|
|Intel i5-6200U / HD Graphics 520 (2016 Skylake-U)|Intel Gen9|AVX2 only, no avx512f|32/32 LoRA, 32/32 switching, 32/32 saved|
|AMD Ryzen Z1 Extreme|RDNA 3|AVX-512|32/32 LoRA, 32/32 switching, 32/32 saved|
Same 2e-7 gate, unchanged. No tolerance was loosened to get there.
What I verified on each side
On the Intel machine, the post-change matrix is bit-identical, field for field, to its pre-change report across all 32 families — peft, gradient, two-step AdamW, frozen base, resume, lifecycle — which is how I know the AMD fix didn't quietly cost the Gen9 path anything. Also 693 passed / 0 failed / 9 ignored on the Rust lib suite and a clean strict headline forward run.
On the AMD side, the fix was re-qualified end to end: 32/32 on all three stages, provenance clean.
The harness fingerprints the pinned Transformers source alongside the shaders and binaries, and on the Intel side I re-derived the whole fingerprint from the pushed tree myself: 3951 inputs, zero changed, zero missing. So "green" refers to one frozen set of reference math, not whatever happened to be on disk.
Same caveats as always
+inf materialization), so I've left it as the committed AMD-built module and documented that rather than swapping it without re-qualifying both hosts. I'd rather report that than pretend it's clean.What I'd love from you
Last time several people asked about hardware other than mine, so that's the ask again: if you build it on an AVX-512 laptop, an AVX2-only machine, or an NVIDIA/Intel GPU, I want to know what you get. The two reductions above are the ones most likely to behave differently on your CPU, and knowing your host's vector width is now part of the answer.
The new cross-platform section in the README documents the whole thing, including which host classes are measured and which aren't.
Repo: https://github.com/necat101/Hierarchos-Native Compatibility/parity record: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/COMPATIBILITY.md Regression audit: https://github.com/necat101/Hierarchos-Native/blob/main/AMD\_REGRESSION\_AUDIT.md Per-host tuning and measurements: https://github.com/necat101/Hierarchos-Native/blob/main/hierarchos-vulkan/VENDOR\_TUNING.md
Been developing a custom agent orchestration layer with opencode as the harnes. It has been successful so far in terms on being manage multiple concurrent sessions. However, I am wondering if I am using the best harness given that opencode isn't meant to be a hackable, i.e. Less flexible compared to something like pi.
I've developed an interface that allows me to easily switch between different harnesses. i.e. Without having to re-write the orchestration layer, and am considering swapping out opencode for pi. The reasoning is obvious, pi is meant to be a hackable harness, so its likely to be a better fit for building a custom multi-agent system. Opencode does support a headless mode which has helped a lot. But it seems heavy on resources, and recently has seemed very buggy. Less features might be a better approach towards the stability that I need.
Before I make the dive, has anyone else already tried integrating pi into a multi-agent system? Did you find pi a better fit for the worker layer compared to something like opencode? What problems did you run into?
Thanks for your help.
Is a small local model more than I need to change handwritten notes to text?
My agent has a skill I use but I could probably not use up my limited inference bandwidth or maybe not as much if it’s a much smaller model, right? What’s the smallest model that can accurately read handwriting?
Looking to hear others workflows or how they handle note conversion and the like
Thanks
I measured Qwen3.6-35B-A3B at 4-bit (UD-IQ4\_XS) hitting 89.6% pass@1 on HumanEval on a single RTX 2080 Ti 22GB — and then ran a controlled A/B of a routing technique I've been playing with: MoE expansion, which activates 20 experts per token instead of the stock 8 on the last 15 layers.
Result: 90.9% (+2 problems) at −19% decode speed. (MoE expansion works!)
Setup (both runs identical except routing):
check(candidate) tests, 12s timeoutResults:
|Config|pass@1|decode|
|:-|:-|:-|
|Stock routing (top-8)|89.63% (147/164)|69 tok/s|
|MoE expansion (20 experts, adaptive, layers 25–39)|90.85% (149/164)|56 tok/s|
Paired per-problem: 139 solved by both, 10 solved only by expansion, 8 only by stock. Rolling pass rate stayed expansion-ahead by +2–3 problems at every checkpoint.
What is MoE expansion? No retraining, no file changes — at inference time the router keeps more experts per token than the model's native top-K (here: 20 instead of 8, with an adaptive threshold so easy tokens keep fewer), on a slice of layers (25–39 of 40). You're consulting more of the network per token. Same trick that gave 84.34% vs 81.82% on GPQA-Diamond at Q8 in earlier benchmarks — now confirmed in coding too, at 4-bit.
Honest caveats:
The tool — I wrapped all of this into AgrillaMoE, a dedicated llama.cpp server for this model: it detects your VRAM and suggests/downloads the right Unsloth quant, applies the expansion profile by default (overridable), exposes OpenAI and Anthropic-compatible APIs (Claude Code works out of the box), and runs on NVIDIA from GTX 10xx to RTX 50xx, AMD via Vulkan, and Apple Silicon via Metal. Static binaries for Linux and Windows on the releases page.
ref.:
https://github.com/vagrillo/AgrillaMoE
https://zenodo.org/records/22255483
Hello, if you are running multi gpu setup, check out ramjet https://github.com/helixml/ramjet
The goal is to have a local version of dynamo that can do equal or better job (but ideally without k8s). Check out https://helix.ml/blog/ramjet-vs-nvidia-dynamo as well.
If you have dgx spark, multi mac setups, etc it would be great to contribute recipes so other ppl can just pull
I got fed up that running ollama pull overnight still was resulting large models failing to download.
When I discovered I could download them faster if I cancelled every hour or whatever I decided I could automate it.
I've done a 3rd tweak with this (Originally written by chat-gpt :D) seems to be less buggy now thanks to claude :D
I've found it by pure chance, searching for a decision model to use into a simulated project. What about it? Does anybody use it?
I was using local conversational models in 2023-2025 in LM Studio but went full Opus and Claude Code from November 2025 until October 2026, now.
whats the best harness + model for my use case? document review and coding. multimodal input and output ideally.
I want to review contracts where even the contract itself is not to be disclosed, and I don't want to put that in the cloud anywhere, so that's prompting me to update everything
so I've installed Pi but don't have any models. And Pi wants to serve local models from llama.cpp but I just read about dwarfstar4 (ds4) but it serves MoE on just a few open source frontier models, yet reportedly wants minimum 96GB RAM for Metal use. I was primarily wondering if ds4 acts like its serving from llama.cpp to a harness like Pi
it seems like llama.cpp is catching up in real time, with the cached MoE thing that got merged in today with some infighting, but I'm not even sure which model I should be using
there's one crowd that's like "we need cached MoE at 20 token/sec with billion param models" and there's another crowd that's like "Qwen 27B is all you need" others are like "Gemme 4B is sooo good now"
do decision models fit in this workflow anywhere? in conjunction with LLM's in a harness loaded at the same time?
I'm pretty lost. I won't remain lost, but I also want to hear others opinion while I experiment myself, hopefully to narrow down what I need to experiment
Consider this: a widely available machine, up to 1.5TB of system RAM, room for 4 passively cooled GPUs with 128GB of VRAM, in desktop or rack format.
That machine was released in 2019, then discontinued in favor of one that had only a max of 192GB shared memory.
This would be the local LLM machine right now, if it were on the market with up to date components. Terribly expensive, sure, but that’s the market conditions, not a design flaw.
Like everyone, I've been following the back-and-forth regarding Jev with interest. Arguments aside about the originality of the idea, the first thing I thought of when all of this came out is "gosh, that could really help my local scribe run realtime loops during a consultation".
I pointed my harness of choice at the problem (I've been relying more and more on the LLMs as the brain rot from AI coding has continued apace). Here is the resulting workflow:
I found that the 4B model would almost never SKIP (Jev and 3.8-Flash were better but still missed 3/4 of them on natural speech). Not surprisingly (in hindsight); using the calculated probability mass alone was essentially no different to just prompting the vanilla generation endpoint and executing based on the output (roughly 81% accuracy). Looking into the logprobs a bit more it seemed that there was a usable signal in there somewhere. GLM-5.3 was pretty good figuring it out: instances where NOTE was selected, P(SKIP) ≥ 0.05, AND the utterance was ≤8 words were essentially always a SKIP. With this heuristic... 0 false SKIPs across multiple runs, and SKIP recall went from 0-50% to 75-100% on the natural consult.
The logprob gating + heuristc step is latency neutral; however, it was more reliable for this task. The otherwise vanilla small LLM like Qwen3.5-4B never flagged SKIPs and would occasionally not follow instructions entirely. I also ran an evaluation with Jev via OpenRouter (a pretty informal test set of \~40 hand-labelled utterances, and the heuristic was tuned on the same set, so it needs a held-out set to confirm); on a natural ambient consult recording the gap is smaller than I expected (both 95% accuracy but 73ms vs 514ms, keeping in mind Jev was a remote endpoint and all the latency that entails). Jev pulled away on a command heavy synthetic script (\~80% vs 100%). The overall intention was to prevent the main-loop from getting too bogged down with fluff and I think this approach achieves that.
First token logprob classification is pretty old hat; but I never really thought about one-shot classification in my scribe before Jev. And yes, the whole point of Jev is that you can just give it a classification task and have performance be good enough that you don't need to apply bespoke heuristics over logprobs to rescue your classifier (but funnily enough even Jev got an accuracy uplift from the P(SKIP) heuristic).
It was a fun experiment anyway (and grossly underpowered to say anything meaningful about Jev in general terms)! The result (video below) has been useful from my perspective (you can try it yourself here).
I bought a Framework Desktop motherboard, and put it inside a Phanteks Enthoo Pro case. I have a spare RTX2060 graphics card. I managed to get it working with the Framework Desktop motherboard, after making it go through two PCIe risers, and mounting it on a vertical GPU bracket.
What do I do with my RTX2060? Should I use it as a subagent?
I currently configured it as a PCIe passthrough device for my Windows VM. I very occasionally use it for Windows gaming using Looking Glass. I am thinking that perhaps I can run a subagent on that GPU. If people have any suggestions, please do let me know.
Context: These engines can only run on specific hardware and can only run 1 or a very limited number of models but what it trades for generality it gets back in performance with these engines out performing general engine like llama.cpp and vllm on those specific sets of hardware.
strata with qwen flash Q2 is pretty retarded and 27b at q4 is also retarded so I am wondering if anyone knows a engine that runs the model Qwen/Qwen3.6-35B-A3B fast on my hardware?
I generally just ask it "describe the main cast of (insert somewhat known cartoon show from the 2010s)", could either be Totally Spies, Randy Cunningham, Slugterra etc, most models in the 30b range completely fumble, looking at you Qwen, but the ones that manage to answer that are gems that can actually hold a human conversation.