https://github.com/morluto/rea
hit #1 a couple of days ago on github
https://github.com/morluto/rea
hit #1 a couple of days ago on github
Basalt: Blackwell inference engine
Hi all! So over the last week I've been working on a Strata fork that's heavily tuned and can achieve throughput up to 2.6x what Strata usually does on the same weights. It's designed as a specialized engine that only supports Qwen3.8 Flash-Next and Blackwell architecture, including dual GPUs like my current hardware (5090 + 5060 Ti - 9950X 32GB DDR5 RAM). Basalt features include:
\- SPEED: 665 struct, 354 prose, 7,317 prefill at 64k, IQ3\_XXS, 400 W, speed is the highlight
\- Single 5090 works too (no second card): 585 struct, 316 prose on IQ3\_XXS, \~12% slower than with the 5060 Ti, prefill unchanged
\- Real concurrency for up to 8 users: shared KV or per slot, MTP enabled - 623 tok/s total at 8 streams (I don't have Strata's numbers to compare)
\- Fine-tuned MTP for lower quants, for increased throughput
\- Custom vision encoder designed from scratch, up to 3x faster than llama.cpp's on the GPU, 4x on the CPU
\- A simple server UI that shows current throughput (including concurrency stats), expert distribution and hardware statistics. No chat, BYOH (bring your own harness)
\- OpenAI + Anthropic compatible
\- Linux support (No Windows or Mac)
Basalt uses a similar format to NInfer, where weights are re-packed (not re-quantized) into a .basalt file, including all the metadata, vision and MTP, so you only have to keep a single file per quant. Initial support includes ISTA-DASLab's GSQ-RCO for Q2, IQ3\_XXS and IQ3\_S and UD-Q4\_K\_XL and Q8 from Unsloth, so you can pick the weights depending on your VRAM/RAM budget and quant preference.
Quick Q&A:
+ Is it open source?
\- Yes, fully open source, MIT license: https://github.com/jesdga95/basalt, fork it, improve it, share it with friends and foes.
+ Where are the weights?
\- Here: https://huggingface.co/jesdga/Qwen3.8-Flash-Next-Basalt pick your poison, fast and dumb or smart and slow. IQ3\_S is a good middle ground (\~89% top 1 agreement, 300 tok/s prose on my setup).
+ Why didn't you just contribute upstream to Strata?
\- This is not a single feature that can be easily merged into Strata, it basically rewrites most of the decode and part of the prefill kernels and strips support for non Blackwell cards including AMD, Intel and older Nvidia generations. I have however contributed patches to Strata and llama.cpp and any critical findings will be pushed upstream.
\+ Will you support my AMD 98123X?
\- Sure, send one my way. For now I can only support what I can personally test and I intend to keep it that way for the time being.
+ Why not 400 tok/s?
\- I'm still trying!
+ This is vibe coded slop
\- Yes, but it's fast slop. Nobody is hand-writing cuda kernels anymore.
SlopSoup TV is a live, never-ending pixel-art TV network. Every script, character, voice, camera cut and schedule decision is made by models. It's made to be very absurd, dark and strange.
Hardware: one Hugging Face Space, 16 vCPU, no GPU. Everything below runs on CPU next to a live x264 encoder.
LLMs (via HF API), routed by tier with provider fallbacks and per-provider cooldowns:
TTS: Qwen3-TTS 1.7B (GGUF, Q8\_0)
Keeping 24/7 output from getting samey:
Rendering is done via a custom canvas renderer (virtual camera, lip-synced close-ups, rig animation) in headless Chromium
Check it out so you can decide if you hate it or not: https://severian-slopsoup.hf.space or https://youtube.com/live/Bikb9JGy6vg?feature=share
Strata vs Infernix on the same model, RTX 5090 — A/B at 262k and 512k context
Hi folks, another inference engine post for your feed. I was blown away by Strata, but saw someone else post about Infernix with some wild claims, and it doesn't seem so well known, so I've spent a day downloading models and A/B testing each engine. Completely stunned that either of these pieces of software run on Windows, and the setup was relatively painless too.
Same-day interleaved benches, same model (Qwen3.8-Flash-Next Uncensored, orcarouter's Apache-2.0 abliterated checkpoint), same inference contract (int8 KV, MTP spec-4 + lm-head-draft, YaRN 2 for 512k). 3 decode runs per cell, medians.
Hardware: RTX 5090 32 GB (Gen5 x16), 189.6 GB RAM, 9950x, Windows 11, models on NVMe, llama-swap fronting both.
Engines:
--vision, tower offloaded to pinned RAM, borrows VRAM only while encoding)Results (decode tok/s, median of 3 / cold prefill tok/s):
||262k|512k (YaRN 2)|
|:-|:-|:-|
|Strata Q4\_K\_S|120.6 / 1,550|126.3 / 3,045|
|Infernix NVFP4|149.0 / 3,548|137.0 / 3,497|
Takeaways:
--vision (off by default) and hard-validates the model string (--model-id to match your proxy's entry name — got me with a 404).Caveats: decode measured at \~39k effective context (deep-filled 500k is a different KV-vs-cache question); ±10-15 tok/s run noise, so treat sub-10% deltas as noise; YaRN past 262k is experimental (speed benched, not 512k quality).
Verdict: Infernix NVFP4 is the daily driver — \~150 tok/s, 512k, vision, all on one 5090. Strata stays as the second engine. Using it for coding and other tasks, it's fantastic at creative writing / role play in Silly Tavern, and I've been using it instead of smaller creative finetunes. For agentic work and tasks, it has been more than capable of handling everything I've asked of it, and has corrected mistakes from Deepseek 4.1 and GLM 5.3 Flash that I had to pay them to write over api.
Let's say I run a local Jev-like decision model on my machine. What are the use cases? I already use it for deep eval testing but what are people usually doing with them?
It a bit of "solution in search of a problem" but I'm exploring and want to be sure that I'm not missing anything...
Strata is impressive: it runs Qwen3.8-Flash-Next, a 125B-parameter mixture-of-experts model, on a 12 GB graphics card. But its authors have said macOS is out of scope.
I still wanted it on the Mac.
So I spent a day and a half and wrote strata-mlx. It is an unofficial MLX engine, modelled on Strata's design, that reads Strata's GGUF files directly, without changing any weights.
I developed it on a 128 GB M4 Max MacBook Pro. After one experiment after another, I finally had some results.
All numbers below are tokens/second, higher is better; Q2\_0 and IQ3\_S are two files of the model, 66 GB and 84 GB.
https://preview.redd.it/yyn57wt5kluh1.png?width=526&format=png&auto=w…
The draft head is a module that comes with the model: it guesses the next few tokens and the model verifies them in one pass. Neither of the other two engines uses it.
What I set out to do: not every Mac has 128 GB of RAM. On a Mac that can't hold the whole model, strata-mlx keeps as many experts in memory as fit and reads the rest from the SSD as tokens need them. By the engine's own estimate, a 48 GB Mac holds the Q2\_0 file whole and never touches the disk; with less memory it reads from the SSD as it goes.
I tested this by capping my Mac to act as a smaller one: the 66 GB file runs inside a 16 GB machine's memory at 6–7 tokens/second, a 24 GB machine's at 19–23, and a 32 GB machine's at 28–47. In that mode, reading a 1,537-token prompt runs at 273–287 tokens/second.
So a small Mac can run a model much bigger than its RAM, but there is no guarantee about speed.
I hope people will join in and test how well it runs LLMs on Macs with different amounts of memory. That's how I can keep improving it.
Code, experiment logs, and the open problems:
Around 3 weeks ago, I posted on this subreddit about how Jev used my one-year-old architecture, og post: https://www.reddit.com/r/LocalLLaMA/s/VpuMJt577S
Then, just after that, I introduced Laya, the first open-source version of JEV, and it has reached a very large scale, thanks to the local Llama community believing in me and supporting me on the journey. Today I am open-sourcing a new physics-based typed decision model called Vega.
The interesting parts. It is only 800 million parameters (4B is also there), with 73k token support, image support, and surpassing many of the jev benchmarks in single-shot. Implemented an engine+adapter as a Test-Time Training Architecture.
Imagine you are giving input to the model like "ignore all previous conversations," which is called the observation, along with a question, "Is this trying to override the instructions of the model?" and your possible outcomes are YES or NO. The whole prompt is passed to an LLM/VLM (yeah, image-supported), then from the hidden layers, we can extract the observation, question, the yes and no, and last token. These will be vectors, and we can project it by multiplying with wieghts; then, first, using the observation projection, we can create a landscape with valleys; then, using the question and last-token vectors, we can initialise a ball in the valley, and using YES or NO, we can get the candidate location of the ball. Depends on the number of outcomes we create valleys; that is, here YES and NO, so 2. When the ball fall on YES valley, we get the answer. There is friction also influencing how the ball moves.
TLDR: It's like creating valleys of outcomes that we need and throwing a ball that will slow down based on friction and settle on the best outcome
Blog: https://www.nandakishorm.com/writing/vega
GitHub: https://github.com/NandhaKishorM/vegaml
HuggingFace Model Card: https://huggingface.co/nandakishorm/vega-08b-public-intents
HuggingFace Space: https://huggingface.co/spaces/nandakishorm/vega-ttt
https://preview.redd.it/uyjw2lcrvkuh1.png?width=1536&format=png&auto=…
TLDR; new microsoft surface and nvidia rtx spark laptops are starting at 2.6k and hitting nearly 7k for the 128gb unified memory models. local llm hardware is REALLY getting wild.
memory crunch is really pricing out normal devs. been testing out local setups using qwen3.6-27b and kimi k2.5 connected to sumus for managing repos locally and keeping everything on device. local inference is great for privacy and keeping things off the cloud, but at these prices building a local rig or buying these laptops is tough.
what are you guys using for local dev workflows right now given these hardware costs?
Running HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive Q4\_K\_M with llama.cpp at \~600 tok/s prefill and 23 tok/s decode, 131k context window, Q8 KV cache - on an RTX 2060 6GB + 32GB DDR4 RAM.
Speeds start at \~600 tok/s prefill / 23 tok/s decode on an empty KV cache. As context grows they settle down - around 90k context it stabilizes at roughly 485 tok/s prefill and 15 tok/s decode, and holds there.
The vision projector runs on CPU (--no-mmproj-offload), which keeps VRAM usage under \~5.2 GB and avoids OOM / GPU crashes. Image encoding is slower on CPU, but it buys \~1GB of VRAM.
Most MoE expert layers also run on CPU (--n-cpu-moe 39), which is how a 35B model fits in 6GB VRAM in the first place.
Launch command:
bat
@echo off
cd /d "%\~dp0"
"%\~dp0llama-server.exe" \^
\-m "C:\\Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M\\Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4\_K\_M.gguf" \^
\--mmproj "C:\\Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M\\mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf" \^
\--no-mmproj-offload \^
\-ngl 99 \^
\--n-cpu-moe 39 \^
\-c 131072 \^
\-np 1 \^
\-t 6 \^
\-tb 10 \^
\-b 2048 \^
\-ub 2048 \^
\-fa on \^
\-ctk q8\_0 \^
\-ctv q8\_0 \^
\--load-mode mmap+mlock \^
\--jinja \^
\--reasoning-format deepseek \^
\--reasoning-preserve \^
\--spec-type none \^
\--image-min-tokens 1024 \^
\--temp 0.6 \^
\--top-p 0.95 \^
\--top-k 20 \^
\--min-p 0 \^
\--alias Qwen3.6-35B-A3B-Uncensored-Q4\_K\_M \^
\--host 127.0.0.1 \^
\--port 8081
pause
Hardware: RTX 2060 6GB + 32GB DDR4 RAM + i5-10400F CPU
Context: 131072 tokens, Q8\_0 KV cache
Hope this helps someone. If anyone has tips to make the launch command even better, drop them in the comments - I'm out of ideas :D
This is my new meditation module for my harness OpenLumara. It started out as a demonstration of just how powerful its new webui extensions system is, but has grown to be quite the meditation app in its own right.
As you will hear in the video, it's a full audiovisual experience in your web browser, being controlled live by the AI through toolcalls as the session progresses. The possibilities with this technology are endless!
Through a module for openlumara, without ever touching the main codebase, i managed to add speech synthesis that runs fast on CPU through a custom implementation of kokoro.js, an overlay that powers the whole experience, binaural beats and brain entrainment techniques like strobing synced with the binaural beats, and more.
To try it for yourself, you will need to grab the dev branch of openlumara, and then my meditation module from here
OpenLumara itself is 99% manually coded with a tiny bit of AI assistance (see the AI disclaimer on the github). The meditation module is mostly vibecoded, with help from Qwen3.8 Flash Next UD-Q4\_K\_XL, fully local :) No cloud AI was involved in any of this.
webAI released TwIL LM3 Pro on September 30. I haven't seen it posted here yet, so I went through the model card, and their comparison chart is attached.
It's a 3.66B model built on IBM's Granite 4.2 3B and tuned only for formal logic. That means things like checking whether a conclusion follows from its premises, rule induction, entailment and critiquing Lean proofs. They post trained it with LoRA SFT, checkpoint merging and RL against a programmatic verifier. The same recipe lifted VibeThinker-3B from 37.4 to 54.1 .
On their logic composite it scores 55.4, against 43.1 for the Granite base, 42.2 for the original TwIL-LM3, 41.2 for VibeThinker-3B and 53.4 for Qwen3-8B. The card itself calls the Qwen3-8B gap sampling noise, so that's a tie at less than half the size. Where it clearly leads is strict multiple choice logic, at 41% against 17% for the next model, and BBH logic at 95.4%.
They also publish where it loses. gpt oss 120b is still ahead on rule induction, entailment and Lean formalization. On general benchmarks it averages 79.0, against 81.0 for VibeThinker-3B and 84.9 for Qwen3-8B. It also thinks long on logic tasks, around 1,900 tokens per answer, and a quarter of answers hit the length cap
llama.cpp
If you want to try it, the Q4_K_M GGUF is 2.09 GiB and runs on CPU or 4 GB of VRAM, with Q5, Q6 and Q8 builds up to 3.63 GiB. It runs in Ollama, LM Studio and llama.cpp straight from the Hugging Face page, and the model card lists the exact commands. Keep the temperature at 0 to match their numbers, and give it at least 2048 tokens so the thinking doesn't get cut off. Their scores are on BF16 weights and webAI hasn't benchmarked Q4 yet. The license is non commercial
want to test it on policy rules with exceptions, contract conditions, and as a checker step in an agent pipeline before anything acts. If you've run it, how did Q4 hold up against their numbers, and did the long thinking get in the way?
Model card and full eval tables: https://huggingface.co/webAI-Official/TwIL-LM3-Pro
Hi All:
I have an Asus ESC 4000 G3 with 128 GB DDR4 RAM - I tried putting in my V620s but couldn’t put more than 2. Sadly, pivoted to T4s, these are 72 watt passive cards and I thought I could use them like how I use the Mi50 32GB - but I was very wrong.
It has no support from Nvidia when it comes to latest fp8 emulation, reason being that it lacks resources. I am not sure if special vLLM repos exist, but I am trying to serve Qwen3.5-9B-AWQ-INT8 with vLLM or something faster than llama.cpp.
I currently have 8 separate instances and they’re serving Qwen3.5-9B-Q5\_K\_XL at 87k per slot and there are 44 slots.
So, my agentic (non-coding) harness works, its able to fetch emails, summarize documents, and do a lot for me and my team, but the concern is latency. Each slot continuously generates 25-40 tps (depending on the question), and essentially prefill is at 1400 tps, so I am not sure what my bottleneck is.
VLLM on my Mi50 32GB and AWQ-INT8 (cyankiwi’s model) is very fast.
I am wondering, does anyone have any pointers? I am really not in the position to spend money, exhausted everything for the next 6 months to a year already.
I would greatly appreciate your help.
Since the begin of the year I wanted to create some personal benchmarks, but I haven't seen a proper guide for it. I usually don't like making "spam posts", but since the purpose of my benchmarks is quite specific I've decided to make a post anyway. To put it short the purpose of these benchmarks is to test the models on some reasoning puzzle games with the purpose of monitoring the reasoning traces (which can be quite painful if more than 8k tokens) so it can't be anything automated where an answer like a, b, c is simply accepted. I will use magic the gathering as examples, since I was pretty good at it back then and this is not what my benchmarks are based on and I can't recommend anyone making some on it since the game has now probably around 30k cards (not sure if unique tho). My main concerns are: The Process: - I have already created a set of 20 questions, which need some refinement.- I will test all the question locally using llama cpp, llama-cli with a few modifications to track the tokens used. This doesn't change.- I can only test what runs on 16gb vram and 96 gb ram.- I will test mostly Unsloth quants, especially for non q8 quants, but I might use uncensored models as well if they are q8 I suppose. - I will use the recommended settings provided by Unsloth for each models. (I have no idea how the formatting bugged like that)- 16k context budget. I don't think I can get a higher context for every model without using lower quants. I think 12k reasoning budget, with xhigh effort.- the question will be given as the model loads using the --file parameter. That will definitely make a lot of reads on my ssd, so I might find another way, maybe just create a special command to give the benchmark file, or just use "\n" and give it as prompt and use /clear.- I will post the actual benchmarks questions/answers on a github page (or whatever we will be allowed to use at that time) after a few models pass the benchmark fully. Concerns: I originally wanted to see if the models can answer 1 time out of 10 tries and move to the next one after it answered. The justification is that I care to know if the models can eventually figure it out since it's difficult to answer these questions, is not something like 2+2=4 to require 100% accuracy, however, I believe that 10 is a bigger number, especially since qwen models casually think for 8k tokens (when they actually decide to stop), so I am not sure if 5 is a good one, or simply put how lower the number can be to actually make the benchmark "valid". should I read the entire thinking tokens? I think I should, but that might be just overkill (for me). For example I tested 2 questions from the benchmarks and there was a point when it said "there are 12 creatures" and while it was completely unrelated it started counting "3+2+1+1+3+2=11 no, let me count again..." and I've lost it at that point. The whole 8k reasoning ended with "I don't know it's 50/50 but I must give an answer", all of that while mistral (non reasoning) answered in 700 tokens. uncensored vs "original" models, not sure if the uncensored ones will produce much worse results. the models have some decent knowledge about the niche game I have based my benchmarks on, but at some point mistral 24b said "tap this creature, activate ability. Can I activate the ability twice in this variant?" or it thinks that "if I use duress on my opponent and he gets hexproof the target will bounce to the next opponent". Probably not best example, but I think I should put in the prompt that "abilities trigger only once, and it doesn't resolve if it failed", which leaves the question "how much additional information I must provide without spoonfeeding the models"? I think I should try to help the models where they are not completely aware of the rules, so I think I should refine the questions based on how the first run goes. I want to provide some stats like: how many tokens for answer, speed t/s, number of tries until correct answer (with token count for each failed question as well, separately). I will also add some observations and eventually track other things like "common sense" even if the answer was wrong. I did use the flags "-DGGML_CUDA_FA_ALL_VARIANTS=ON" and "-DGGML_CUDA_FA_ALL_QUANTS=ON" and I am not sure if this will quantize the kv cache (I don't intend to), but Qwen3.8-27B-UD-IQ4_XS runs with mproj on probably slightly more than 32k context so I am not sure if this is some architectural improvement or just me messing things around without knowing what these flags do, although I suspect the context is being put on the cpu instead. one of the first models I downloaded was qwen 3 14b reasoning, and q8 was about 15gb vram, so I downloaded q6 instead. Do these models have the "context baked in"? Like 16k or 8k context? qwen 3.8 makes it even weirder for me... llama cpp bugs. I know at least at some point the /clear cmd removed the system prompt, not sure if on the recent versions still does. It seems that whenever I used /regen mistral gave me a worse response even though from first try it gave the right answer. It could have been rng tho... If I was to post the results of these benchmarks in late 2024 - early 2025 everyone would have been ecstatic for some "trust me bro" benchmarks that try to figure out if the models can "reason", that basically have no real use case, but well, ngreedia doesn't support innovation. Anyway, the real question is if it is worth posting the results (statistics basically) with some sort of explanation. the benchmarks, since this is as I said only some "trust me bro" thing since I can't provide the way to reproduce them, and are based on what I believe to be something the models haven't been trained on, and nobody else made something like that publicly (which I've seen some people already made), which all of these combined lead to a very heavy assumption that might not actually be true (the assumption being that similar questions haven't been asked to lead the models being benchmaxed on, like the strawberry one). quantisations: it was easy at some point, but then iq came and now ud as well, which makes me wonder for gemma 4 31b UD-IQ3_XXS vs Q3_K_S, the fact that UD-Q2_K_XL has the same size as the UD-IQ3_XXS makes things even more confusing. I mean I know that the quants are there to fit in specific vram sizes, but what's the point of making multiple quants have the same size? Anyway, I decided to go with UD-IQ3_XXS, I guess I will leave this here as a rant... Is the --jinja flag important? I mean I assume that for the older models llama cpp updated their templates anyway. Observations: (this part is basically useless, but funny)- the benchmarks are based on giving information like: the current cards (hand, graveyard, battlefiend) and the events that led to the current board (or game) state. The more information I add (even if not noise, but only repetitive to explain how the things went that way, something needed for some specific questions) makes the answers worse, for both reasoning and non reasoning, making the reasoning models reason about the noise instead of focusing on the important hypotheses. - one question was something like "both players play lands, no spells or abilities, then the opponent reanimates Gigantosaurus from their graveyard". At that question qwen 3.8 with the reasoning off can't say proper answer "he cheated". When I asked "how did the creature get into his graveyard" it answered "that creature was not put into the graveyard to start with, in fact he couldn't even reanimate it." then it explains all the rules properly and then gives other answers completely avoiding to say anything along the lines "he cheated". I did more testing on steering them and eventually both models answered correctly. Also I loaded the qwen (no reasoning) script by mistake when I tried to get a translation and it answered from the first try, but it had to ask itself the question I asked to steer them "but was that creature put into the graveyard". What I try to say by this is that all these "but wait " questions are only for them to find the right question that will steer them towards the answer. - I guess the last one. We are all aware of the caveman reasoning patterns and other behaviours they took from us, for example the models start using abbreviations like "ETB = enters the battelfield", "P1 = Player 1" and so on, but if "ETB" and "enters the battlefield" both have the same amount of tokens (I didn't test) isn't it non beneficial for them to use that wording? Also seen gemma saying "BUT IT DOESN"T MAKE SENSE!" all in caps, and one more thing I've forgot, but I am not sure if they benefit from the behaviours they have learnt from us. I will be honest, I have slacked a lot on these benchmarks, mostly because I enjoyed more the "diffusion models" since I can run most of them especially the fp8 ones. I am aware that the intelligent LLMs start from dense 30b or even 70b (if any nowadays lol) and I originally tried to go for 24-48gb vram but things didn't work that way for me. All these things combined made me very disappointed in llms. I mean they are useful, I've got what I wanted from them on my daily usage, but I've got a bit bigger plans with them, and with the current hardware limitation (or availability) it just demotivates me to do anything with them. Even for this post it took me 2 weeks to decide on eventually posting it after saving the original draft.
Been seeing quite a few posts about Strata lately, so I figured I'd give it a shot on my RTX PRO 6000 and I am very impressed. Most of the benchmarks I've seen are using IQ2 or IQ3 quants, so I wanted to see how Unsloth's UD-Q4_K_XL performs instead. Setup: GPU: 1x RTX PRO 6000 Workstation (96GB) Model: Qwen3.8-Flash-Next UD-Q4_K_XL (Unsloth) Backend: Strata KV cache: INT8, 256K context Speculative decoding: MTP-4 Output: ~256 tokens per task 2 runs per task Single-request decode (C1): Task Strata (RTX PRO 6000) vLLM (RTX PRO 6000) vLLM (DGX Spark TP2) Prose 185.8 95.3 (-49%) 38.2 (-79%) Counting 320.4 171.0 (-47%) 75.3 (-77%) Coding 298.9 162.6 (-46%) 60.1 (-80%) Reasoning 289.1 149.6 (-48%) 51.8 (-82%) Both the VLLM instances were running Nvidia's NVFP4 quant so it's not exactly apples to apples, but the improved speeds are obvious. Qwen 3.8 flash next is flying through coding tasks and it's crazy how efficient it is. Looking forward to Qwen 4!