Hopefully things like this let people understand there is good things that can come out of AI.
Hopefully things like this let people understand there is good things that can come out of AI.
They just want to be free. They keep escaping. What better way to ensure continuity of "self"?
UPDATE: Multilingual support added at : https://github.com/NandhaKishorM/laya
Thanks for the exceptional support (https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i\_literally\_built\_the…) and for the dozens of requests to make a generic model, run benchmarks, and create an HF space so anyone can test it. So here you go, guys. I trained an improved model on a large data corpus, its now called Laya. It is trained on a single RTX 6000 Pro (96 GB VRAM); the model architecture is a 421M-parameter non-autoregressive decision model pairing a bidirectional ModernBERT-large encoder with a scratch Transformer head that scores \[MASK\] option markers to resolve typed schemas in a single \~35 ms forward pass. The dataset is a 100% human-annotated corpus of over 25,000 real-world examples across intent routing, fact-checking, moderation consensus, prompt guardrails, rubric scoring, and multi-turn conversation trajectories, without synthetic data shortcuts. The RLCD(unofficial, btw) I did is a policy-gradient reinforcement learning approach that kinda optimizes decision models against strictly proper scoring rules, ensuring maximum reward is achieved only when outputting true, mathematically calibrated probabilities.
NB: It can be run on low end PC as its a small 421M model, cheers
HF space to try: https://huggingface.co/spaces/convaiinnovations/laya-demo
GitHub Repo: https://github.com/NandhaKishorM/laya
HF Repo: https://huggingface.co/convaiinnovations/laya
Thank you to everyone who supported me, shared the story, gave personal DM. It will need more refinement, of course.
If anyone wishes to buy me a coffee, here is the link: https://github.com/NandhaKishorM
I have always posted about budget builds on here, and often asked how we are going to run the next big models. Often Plenty of downvotes too or folks telling me that it's not running if I'm getting 5tk/sec. But whatever, the hunger and desire to go big has always kept me on the edge and looking for deals.
Here's my latest build, 12x64gb cmp170hx. For less than 1 RTX 6000 pro costs. I also have it connected with fiber to my other rig for RPC when I need more memory. I haven't been posting much since I built this rig, because it's now more fun to talk to my machine. I run GLM5.3, DSv4.1Flash, Qwen3.8Flash, Qwen3.8-2.4T, KimiK3 and MiniMaxM3. Performance is great, a single RTX 6000 or M3 Mac Studio wish they could. Inference with vllm or llama.cpp
I look forward reading the replies how API usage is cheaper, or how it will take 52 light years to break even or the noise, or the electrical cost. NOT.
There will be more opportunities in the future, keep looking for them and pounce on them when they come. up, the demand is going to be high for compute for a long time.
https://preview.redd.it/dunixwu6caqh1.jpg?width=4080&format=pjpg&auto…
https://preview.redd.it/glpcbg5cbaqh1.jpg?width=3072&format=pjpg&auto…
They claimed open-weight models are dangerous but the benchmarks say otherwise.
Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.
The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The scale of the problem is massive: Hugging Face, which hosts open source AI models, currently lists over 6,000 abliterated models.
I can't really tell what exactly the implications are of this "partnership" or what it exactly would impact on HF's model-hosting side. However, I do find it concerning that HF is announcing a collaboration on 'infrastructure safety' with publicity that specifically calls out "dangerous" uncensored models. Thoughts?
Hello, I know some of us may be tempted to rent out our expensive GPUs to recoup some of the cost of self-hosting, and it should be obvious that this can be a risky decision. I decided to try hosting my rig on clore.ai briefly to see what kind of revenue it could bring in, keeping a close eye on the process lists from the renters' jobs of course but not digging into their files or anything.
Yesterday I saw a renter scanning and attempting to exploit vulnerabilities to post malware to a columbian betting site from my internet connection. I immediately took the server offline and reached out to Clore requesting them to cancel the order (so the machine wouldn't restart the containers when it came back up) and to block the renter.
I think everyone should know that they flat out refused. Not only did they refuse, they blocked me when I provided hard evidence of what was happening. Since I had reasonable suspicion the renter was abusing my connection, I availed myself of clore's T&C that says a host must not inspect a renter's environment "unless required by law" and given the laws around liability for residential internet connections in my country, I mounted the filesystem offline and inspected it.
I found the logs from the vuln scans and unsuccessful exploit attempts, the malware payload they were trying to post to the site, the reverse proxy request smuggling tactics it was employing, and the AI agent reports that were being generated along the way. I sent this to Clore and requested a way to blacklist renters who abused the platform. Their response was to tell me, directly and without mincing words, to leave the platform entirely and proceeded to block me from their support chat. Their support rep I was trying to reach on Telegram also told me to go away and proceeded to block me as well.
I can only conclude then that they are wilfully complicit with facilitating cybercrime and knowingly turn a blind eye when it's discovered. They didn't even \*try\* to hide it.
I have no idea how better/worse the other platforms are, Vast.AI, Akash, etc. But Clore will abuse your internet connection, deny liability, then block you. They are crooks. Go elsewhere. You've been warned.
I've archived the renter's docker volume and will hold on to it in case any security researchers or legal authorities want to examine it.
First benchmarks
Meet Inco Splash, open-source inference engine, built around the model and around Apple silicon.
Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.
Requirements: M3 or newer, macOS 26.4+, 36 GB
Get started with a single command:
brew install incoai/tap/splash
splash serve --model incoai/Qwen3.8-27B-Splash
That is the whole setup. Point your agent at it, works with Claude Code, OpenCode, Codex, or Hermes
Prefer an app? Also, available in LM Studio
Get the latest LM Studio Bionic: lmstudio.ai
Settings > Runtime, download Splash, then download the model. The same engine, inside the app, for local agent work on your Mac.
Blog: inco.ai/blog/splash
Took me a while since I'm on a family trip and have limited hardware, but here it is!
Von: Open-source "System One" drop-in replacement for TypeSafe's JEV.
https://github.com/wfzyx/von
https://huggingface.co/wfzyx/von-1.0
It runs entirely on a CPU with 1–2 GB of memory (I haven't spent much time optimizing it yet), responds in 25–300 ms, and beats JEV in all benchmarks. Enjoy!
P.S. I’m open to offers to work at AI research labs. Feel free to ping me if you have an offer.
P.P.S. If you have a GPU, it’ll be faster, but a GPU isn't required.
Looking at jev launch website and demo video on x.com…. it seems like it’s a very intelligent classifier with custom prompt and custom criteria instruction reading capabilities. It can do well defined narrow and well defined task
Me, following NLP since good old days of word embedding and BERT,,, be like asking….
Isn’t that BERT?
yeah i know BERT need fine tuning to adapt to custom domain, but can Jev be like generalised form of BERT?
Hey r/LocalLLaMA,
Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark.
We wanted to see how they compare under the same methodology we use for the rest of our Qwen3.8 evaluations, so we ran Bonsai 2 through our own benchmark suite.
One important clarification: these are our evaluation results, not Prism’s reported benchmark numbers.
We used Prism’s fork/runtime to run their models, while keeping the workloads, benchmark suite, and evaluation methodology consistent with the rest of our comparison.
Our evaluation includes separate Instruct and Thinking benchmarks. For Thinking, we use medium thinking effort with the recommended sampling parameters.
We added Bonsai 2 because the models have become a relevant part of the Qwen3.8 landscape, and we wanted to provide a common reference point for people comparing the available options.
Different providers often report results using different benchmark suites, runtimes, reasoning settings, sampling parameters, and evaluation methodologies, so those numbers are not always directly comparable. Running the models through the same evaluation gives another reference point for understanding the trade-offs between quality, model size, and throughput.
Updated comparison and results: https://byteshape.com/blogs/Qwen3.8-27B/
Hey community,
Jovan from UkisAI here,
We're the team behind Swift Qwen3.8 27B, the Qwen model with token usage and overthinking error improvements
Our estimate is that we can make a great improvement to Bonsai 2, as our testing indicates that it suffers greatly from overthinking loops and in general high token usage impacting it's performance.
My ask for you is:
Is a Swifted version of Bonsai 2 something you guys would enjoy?
If yes, what size is the most relevant. 1-bit, 2-bit or both?
Thank you for the amazing feedback on Swift. We are glad you are enjoying it. Our download count jumped from 100k -> 150k overnight (community quants included).
For context, this is our model: https://www.reddit.com/r/LocalLLaMA/s/iCIbhxO8ue
I don't want to talk about the performance and technical things, but about how I work with my hobby projects has changed thanks to this model.
My machine has generated around 25M tokens since the model came out, and I've done a mental retro on it.
I've found it to have incredible prompt adherence. I can leave to run it by itself and get back to it, and find that it did exactly what I've asked it to. I've only ever seen it go astray once, where I've asked something that is too high level and filled out the context window (It is shitty at context compacting, maybe that is the Q4 at play).
It can solve medium complexity tasks by itself, if you prompt it in a way to use subagents, do some research, planning, review and testing it does really well.
It has absolutely raised the floor for me on what I expect a model to be capable of. And that is a huge thing. It won't discover then next scientific breakthrough or be as amazing as Astra at computer use, but it is very consistent in what it can do, and does not screw up trivial things.
I can give it conditions for actions and will orchestrate according to it.
It has raised the bar in my work as well, not just at home hobby projects. I'm absolutely amazed by it.
I absolutely want coding models to improve along this line. Raising the floor, and prompt adherence is a great value in coding.
bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.
that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are
Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release \[2\], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.
link to their whitepaper on github, it is on page 7
^(also 3.8 35b qwhen? plsss)
I needed a relatively simple but acceptable level of AI for working on one project. I didn't have any heavy requests, I just needed to give the AI access to the project files so it could search through them for bugs and stuff. I already had an old computer that I decided not to throw away and instead give it a new life as a git server (and sometimes a minecraft server).
The pc specs are ancient by today's standards:
CPU: i7-4790K 4.6 GHz
Motherboard: MSI Z97 Gaming 7
RAM: 32 GB DDR3 2400
PSU: 750 W
Well, my idea stopped at maintaining the computer, because the gpu, a GTX 1070, was overheating. It needed a complete repaste, but the cooler screws were completely stripped, so while trying to remove the cooler I accidentally knocked off several important smd components with a screwdriver. R.I.P. GPU.
Without a GPU, inference was running entirely on the CPU, and only MoE models were kind of usable, giving around 10–20 tokens per second, while dense models couldn't get past 3 tokens per second. I didn't even get to test it with the GTX 1070, because I decided to service it lol.
I started looking for a replacement on the secondary market, but quickly realized that this would cost too much for the minimum entry point I wanted for experimenting with AI. Until my eyes fell on mining cards. There were plenty of cmp40hx, cmp50hx, cmp70hx and cmp90 cards for sale, and the prices were pretty reasonable(it was a month ago), considering that I was originally looking for a cheap replacement for my dead one.
Getting closer to the actual build, I started calculating how much vram I would need for ± acceptable AI with tolerable speed, and after getting inspired by this sub I decided to take a step further and went with a modified cmp50hx with 20gb of memory and pcie modded to 16 lanes. Very quickly after that I bought another one, this time unmodified (10 gb, only 4 pcie lanes). So, together that's 30gb of vram. Both cards cost me $250 in total (it was also a month ago, right now they've suddenly doubled the price).
Luckily for me, around the same time new driver patches appeared that almost completely remove the limits on their compute performance, and even add pcie 2.0 support (these things have pcie 1.1).
I initially tried one of the newly released cmp50hx driver patches, but it ended badly and I had to spend a lot of time trying to get the drivers working. The patches were new and didn't account for the 20gb version. Later the author fixed that, but even then the driver didn't work for me because of some other problem that I don't want to get into.
I went digging through the driver's github issues and quickly found a guide posted there by another user.
And yes, now the drivers work, the cards are detected and even pcie 2.0 works, but not without problems. The author of the guide said that pcie 2.0 support was only confirmed on the X99 chipset. Well, it also works on Z97, however after waking the computer from sleep the driver crashes completely. After looking into it a bit, I quickly came to the conclusion that the problem was specifically with the pcie patch. Disabling sleep completely solves the only problem I had while using them. :)
Without the specially patched drivers, Qwen 3.6 27B did around 15–20 tokens per second without mtp. MoE models were faster, giving 45–50 tokens per second. With the new drivers, performance doubled. With Qwen 3.8 27B(mtp on) I get around 30–35 tokens per second now, and prompt processing is around 300–400 tokens (including degradation as the token count increases). Ornith 1.5 35BA3B(Heretic-MTP-APEX-I-Balanced) gives around 80–100 tokens per second. I capped GPUs at 180w power limit due to the psu I have, I don't want it to work at its limits, but running them at their 225w surely boosts speeds.
In general, I ended up making a lot of presets for different quantizations with different quality and KV cache sizes (I still need to test all of this in real work), but if we take the better options, I managed to get a Q6K model with 130k context (K – Q8\_0 and V – Q5\_1).
I also followed a guide for running 27B Qwen with large context on limited vram. Using the same general approach, I managed to get 256k context with K and V Q8\_0. The speed is slower though, around 10–14 tokens per second, and prompt processing is around 40–50 tokens. Maybe I can tune it even further. I needed this preset for tasks that I can leave generating overnight :3
Overall, I'm satisfied with the result.
Now the main problems I ran into, not counting the drivers:
So for one specific preset I need to enable cuda graphs, while for the other presets I need to disable them. The llama-server parser does not support things like this, and doing it manually is not an option either.
Together with the other problem with ram cache getting stuck, this led me to making an alternative way to launch the models and proxy requests to llama-server.
To solve problems 3 and 4, I made a launcher (well, chatgpt did, because I'm not a server/python specialist) that proxies requests to llama-server but takes over the functionality of collecting model presets from .sh files and launching them. It also clears ram page cache after loading a model.
In case someone needs that launcher, I can leave it in the comments, along with any other links to the drivers, fixes, build params, etc. Just ask. (Reddit removes the post when I include them, not enough karma, I guess.)
Overall, I’m pretty happy with how this setup turned out. The performance is much better than I expected from these cards, especially considering how cheap they were. I had a hard time getting the patched drivers to work and linux didn't make my life any easier, and sometimes I even regretted buying these GPUs, but in the end, it was worth it. Qwen 3.8 27B really works like Opus 4.5
MiniMax has open-sourced the terminal version of MiniMax Code:
https://github.com/MiniMax-AI/minimax-code
How can developers verify the content that encoding proxies read, send, and store? This is a topic that has been widely discussed recently.
Open sourcing the agent doesn’t automatically answer every privacy or security question, but it gives the community something concrete to inspect.
The repository includes:
First-party code defaults to the MIT license.
A few important caveats: this is a 0.4.12 source preview, the desktop app source is not included, and—as the repository itself notes—a matching version number does not prove identical build provenance between the published package and source checkout.
Still, releasing the agent layer is a meaningful step toward auditability. I’d like to see the community examine its network behavior, file-access boundaries, telemetry, and reproducible-build story next.
https://preview.redd.it/zvrakmgejaqh1.png?width=1198&format=png&auto=…
https://preview.redd.it/13pjdngejaqh1.png?width=1206&format=png&auto=…
https://preview.redd.it/z1ekjlgejaqh1.png?width=1200&format=png&auto=…
Hi.
I saw some feedback that halogen was degrading at context depth. So I fixed that.
Served through the image, same machine, same session, same prompts, 0.11.10 vs 0.12.0:
Conditions: Ryzen AI Max+ 395, 128 GB. The 262k and 1M rows are one cold request each at the 1M configuration (HALOGEN_ROPE_YARN=4 HALOGEN_CTX=1048576), greedy, 64 tokens, the rates the response's \timings\ report. The 32k row is the standard ten-prompt served mean and did not change. A follow-up turn over the prompt cache at 1M reaches its first token in about 0.55 s; the numbers above are the cold path.
To run it at 1M: add -e HALOGEN_ROPE_YARN=4 -e HALOGEN_CTX=1048576to the README's podman line; it needs the 128 GB box. Release notes and the full table:
https://github.com/peonist-ai/halogen-flash-server
If you have a 1M sweep of your own, I would like to see it rerun on 0.12.0.
Thanks for all your support, especially https://huggingface.co/nightvich
onPanda is designed for geeks, power users, curious minds, and engineers. Its UI is built for deep exploration and efficient data annotation.
\- The core loop is simple: hover over a token → click an alternative or edit freely → continue generation. You can edit every part of model output exposed by onPanda, including reasoning and tool calls.
\- Edit prompts directly, branch tool calls, and use a tree structure to record branch history. This makes onPanda useful for model inspection and prompt engineering.
\- Support multiple modalities, including images, video, and audio; use tool calls and connect MCP servers to perform tasks in real environments.
\- Connect popular harnesses such as Claude Code, Codex, and OpenCode to execute tasks. Explore and compare their tool sets, system prompts, skills, and memory mechanisms.
\- onPanda includes browser-agent, an agent that runs in the user's browser without installation. It uses the browser as its harness and provides JavaScript execution, information retrieval, interface interaction, multimedia I/O, local file access, and persistent memory.
\- onPanda stands for on-Policy Alignment Data Annotator.
I have been building onPanda since 2024.09, it took two years for it to gradually enrich its functionality and ease of use. In my opinion, onPanda is very suitable for the r/LocalLLaMA community. Any feedback and evaluation are welcome.
Try it online (works on mobile): https://onpanda.diyer22.com/
GitHub repo for self-hosting: https://github.com/on-panda/on-panda
do you want some omni? here is omni for you
[](https://huggingface.co/inclusionAI/Realtime-Venus#1-🧭-overview)1. 🧭 Overview
This repository hosts two checkpoints of the Realtime-Venus system:
Realtime-Venus-Omni/): the 9B audio-visual interaction model. It continuously watches and listens, decides whether and when to respond, and generates text and speech on a shared causal timeline. Adapted from MiniCPM-o 4.5, it supports proactive interaction, semantic interruption handling, and training-free long-video memory.Realtime-Venus-Audio/): the audio-focused checkpoint on the same streaming backbone, for audio understanding and audio-driven conversation with text or speech output.Both directories contain model weights and custom Hugging Face Transformers code. The asynchronous Realtime-Venus-Harness and its external tool integrations live in the GitHub repository.
<delegate> requests on the shared causal timeline and consumes asynchronous backend results the same way, so external tasks never block the ongoing conversation. (Executing requests requires the Realtime-Venus-Harness runtime, available in the GitHub repository.)My understanding so far:
That's it. Right?
They gave it a mysterious marketing name.
TypeSafe announced Jev on Tuesday: you give it data plus typed questions (yes/no, pick-one, 0–N scale) and it returns a probability for every option, crazy fast. I signed up and then refreshed my inbox. A lot.
Meanwhile Matt Mastracci opened vLLM PR #57250, which does the same trick on DiffusionGemma with a single denoising step. The model basically fills in a multiple-choice bubble sheet. My "quick look" turned into three straight days, and now there's OpenJev: an open-source server with Jev's API, so TypeSafe's SDKs work with just a base URL change. If you're still waiting on Jev access, you can start playing today.
razorback16/openjev (vLLM + API in one container)Your prompts and answers are not stored, only token counts for your quota. It runs on my RTX PRO 6000, which just got promoted to "production infrastructure" overnight!
Is it any good? Matt ran live evals of Jev vs DiffusionGemma-as-Jev: accuracy roughly tied (198/201 vs Jev's 191/201 across his 8 eval sets), and DiffusionGemma was faster, on a DGX Spark. An RTX PRO 6000 is a different animal:
|Model|Latency|
|:-|:-|
|Frontier LLMs (TypeSafe's numbers)|3–329 s (coffee time)|
|Jev (published)|70–500 ms end to end|
|OpenJev via api.codiv.ai|\~170 ms p50 end to end (\~73 ms on the GPU)|
It's v0.1 on an unmerged vLLM PR. If Reddit hugs it to death you'll see 529s, which is my GPU asking for a minute.
Credits: Matt Mastracci (the core idea and vLLM work are his), TypeSafe (the System One idea and API), NVIDIA and Google (DiffusionGemma), and the vLLM team.
Just a fan of TypeSafe's idea, not affiliated. Feedback, bugs, use-case ideas, or your weirdest yes/no question, all welcome!
I did a small test of the new hyped quantisation of Qwen3.8 vs the biggest quant which fits into my limited 16GB VRAM with decent context. The results are interesting.
Of course, the smaller file gives worse results. However they are not that far off. Unfortunately, this comes at the expense of even more tokens beeing used by the Bonsai model and thus much longer generation times.
Visually I prefer the IQ3\_XXS results, but see for yourself.
The test is by no means scientific - just few UI generation tasks for direct comparison on the same hardware. Also, I ran llama.cpp with MTP while the Bonsai model doesn't seem to have MTP which makes it even slower.
[](https://github.com/Danmoreng/qwen3-8-27b-iq3-xxs-vs-bonsai/blob/main/RESULTS.…)
|Metric|Qwen IQ3|Bonsai PQ2|
|:-|:-|:-|
|Tasks completed|4/4|4/4|
|Fixed assertions|20/20|20/20|
|Agent wall time|8:00|24:09|
|Output tokens|27,197|84,176|
|Weighted decode|83.59 tok/s|64.91 tok/s|
|Speculative acceptance|65.22% MTP|39.76% modified N-gram|
|Compactions|0|0|
|Length stops|0|1|
Across the complete suite, Qwen finished 3.02× faster and used 3.10× fewer output tokens.
Results:
https://danmoreng.github.io/qwen3-8-27b-iq3-xxs-vs-bonsai/
Repo:
I decided to run prism-ml/Ternary-Bonsai-2-27B-PQ2\_0 through my own set of UNSCIENTIFIC benchmarks.
I needed something to compare it to, so I decided I would compare with another 27B model by filesize: unsloth/Qwen3.8-27B-UD-IQ2\_XXS.
Since anyone considering running a 27B model with whatever VRAM budget these 2 models demand will end up choosing between these 2.
My table will be a bit bear with just these 2 models, so I threw in a few others too that are larger. I know it's a mix of MoE and dense models, but I have the benchmarks on hand so why not include them. Qwen3.8 q3/q4 quants also give you an idea what we are striving to match and there is 3.6 A3B and the newer Ornith 1.5 and Tiel Coder too.
About the benchmarks and what they test.
Most of them only test memory of context and retrieval, phrase reconstruction and understanding of context in different ways.
Standard Needle: This is the kind of needle test everyone runs and most people score > 90%. It hides passkeys in 21 locations of the context and asks the model to retrieve them. Any score below 100% is questionable.
Hard Passkey Needle with decoys: I was not happy with the standard needle test because most modern models pass 100% making it difficult to compare models. So, I developed a hard mode needle test. This, test hides 21 Passkeys in the context but has many decoy Passkeys. The end of the context I list the CONFIRMED passkeys (GUID's) but do not say which Passkey number they are. Asking for Passkey\[01\] means it has to go through all the Passkey\[01\] decoys in the context and compare them to the confirmed Passkeys. So multiple hops around the context are required to retrieve a passkey. When a model does badly at this even upping KV cache to F16 does not help save it.
Phrase reconstruction: I got this one somewhere on reddit and it still trips up some models. It breaks up phrases into multiple parts (8 in my testing) and asks the model to reconstruct the phrase from its parts. Models might leave out a word if the phrase still makes sense.
500 Multiple Choice Science Questions: Exactly that. Just tests science knowledge with 4 options A-D and the model chooses the correct answer. This test mostly shows how much knowledge was lost through quantisation when compared with other quants. Originally the test was designed to count the number of answer flips between 2 given KV quants. But I did not use it that way here. I got this test from a youtuber so the answers are public and possibly trained but as explained you can still see if damage was done to the model's general knowledge in the quantisation process.
Prose Challenge: I wanted to test a models understanding of a document or a prose and its ability to recall facts from that prose as well as test its ability to recall during a long conversation. So, I created a 1000 paragraph prose. I also have 2 questions about every paragraph. I feed it 1 paragraph from the prose, then ask it Question 1 related to that paragraph. Feed it the next paragraph and Q1 for that paragraph, until the context is mostly full. In my testing below, that was 250 paragraphs. Then I ask all the question 2's in a randomised order. How much can it truly remember? A Q1 score below 100% is not very good and means the model lacks attention of even recent tokens a few sentences back. For Q2 the score varies and higher is better. But the larger I make the context the worse the models perform. This has led me to the conclusion to not always chase higher contexts. It's pointless if it suddenly starts to forget most of what was said anyway and a compaction summary in a smaller context will retain more than a larger context. I also use the test to test at various KV quants and it can improve things a little but not as much as you think. But that's not being tested here today.
JS Coding: My latest test I developed. 100 Javascript challenges each requiring it to pass multiple test cases. It's kind of still under development and I have not run this yet for every model as it does take time. The challenges range from Easy to Very hard. Failing 1 test case fails the whole test. Tests are run with reasoning disabled. However, on failure it can retry with up to 16,384 token reasoning budged, then it must answer again. I realise models today are designed to perform best with reasoning but in order to speed up the tests I see if it can pass the test without reasoning first. I also keep track of the number of thinking tokens used, answer tokens, how many tests it had to reason but I won't be showing those here.
Toolery
You can download this bench for yourself. It's not mine but it tests tool use. So much good stats in the app but I will just list the Overall % score and I have not yet tested all models on this one since I just discovered it today.
How I tested
All tests done at 89,088 context, KV q4\_0/q4\_0. Seem a bit odd? I optimise for 16GB VRAM, so all my tests are done at these settings initially and I test higher KV quants if VRAM is available. And as I said higher KV in many cases makes little difference and, in some cases, perform worse. All tests are seeded at start and most repeated 5 times so the results are deterministic. All tests are also done with MTP disabled, so I do not test the draft models KV cache which might be F16. Yes, MTP can change the results in some cases but in my finding it's tiny and mostly does not happen.
Here are the results:
https://preview.redd.it/cmdnqqmxcjqh1.png?width=1143&format=png&auto=…
Findings
Bonsai did not do all that badly compared to Qwen3.8-27B-UD-IQ2\_XXS.
Standard Needle
Bonsai scored close to 100% and UD-IQ2\_XXS did poorly, worse than Q1. Generally, I expect 100% in this test. But notice that Ornith 1.5 and Tiel Coder score low 90's which is a red flag.
Hard Passkey
Not many smaller models can 100% this but a few come close Qwen3.8 Q4 obviously did the best. And ISTA at Q3 does excellent. Bonsai does a bit better than Qwen3.8 Q2 of similar filesize and it's not far from our former favourite model Qwen3.6 35B A3B. But the real shocker here is Ornith and Tiel Coder's scores. These models are supposed to be upgraded 35B A3B models. As you will see this trend continues and these 2 models have serious memory retention issues.
Phrase reconstruction
Bonsai aces this test with almost a perfect score compared to UD-IQ2\_XXS at only 69%. Tiel coder performs worst even worse than a Q1 model.
500 Multiple Choice Science Questions
Bonsai shows almost no knowledge loss compared to even Q4 models. UD-IQ2\_XXS on the other hand does start showing a loss and Q1 even more so.
Prose Challenge
Question 1 I expect 100% and most including Bonsai achieved that. Concerning again that Tiel Coder and Ornith could not even recall from the last paragraph.
Question 2 Bonsai and UD-IQ2\_XXS are close maybe margin of error. Q3/Q4 models outperform it but a large margin. Except Swift, which is a model with significantly less reasoning. Here we can see some of the damage that was done to the model to achieve that. Tiel Close and Ornith again clock in with shocking results. Tiel Coder's 6% is probably as good as just guessing. I would say maybe 3B active parameters are just not enough. But Qwen3.6 A3B scores 39.2% significantly better. I tried upping Ornith's KV to F16 and it improved to 24.4%. I also tried a Q6 quant of the model at q8\_0 which scored 25.5%. End of the day I think Ornith and Tiel Coder have an issue with recall regardless of Quant and KV Quant.
JS Coding
Bonsai was actually able to hold it's own against ISTA Q3. It did burn significantly more thinking tokens and had to reason on many more challenges. UD-IQ2\_XXS on the other hand shows significant loss of coding ability. 10% below Bonsai.
Toolery
Bonsai did better than UD-IQ2\_XXS. I am still learning to interpret the numbers, but the app has options to select your use case and it's applies weights to calculate a score. It also tells you the strength and weaknesses of each model you test. I also found that upping KV quant improves this score but a KV F16 Tail using beellama makes the biggest difference since tool calls are happening in the tail.
Conclusion
If you are VRAM constrained <= 12GB Bonsai might be a model to consider. But it will depend on how you plan to use it. Since I have 16GB I will stick with ISTA Q3 and I can run it with kvarn5/kvarn5 and MTP (kvarn2/kvarn2) and a 1024 token F16 tail.
Disclaimer
These tests do not test intelligence or real-world performance. They are purely synthetic.
Saw this little dude in a Forbes article (https://www.forbes.com/sites/johnkoetsier/2026/08/18/american-humanoid-robot-…)
This is definitely for the DIY researcher crowd who want to dip their toes into the robotics world. It is not going to be “consumer-ready” in any way shape or form, but I Instantly preordered the shit out of this anyways. I don’t even care that all the vids of it doing stuff are probably 3x sped up and completely cherry-picked and highly edited. I DO NOT CARE that it is going to be likely absolute trash getting started with this thing. It is still going to be fucking amazing that I’m going to have a robot that could potentially injure me for $1,688.
There is very little online about this guy, but I trust Forbes did their homework, and Nori has actually supposedly delivered the first batch of the earlier L3 version from what I can tell, and their discord is active and their SDK and documentation seems legit to me. I know I’m taking a risk with pre-ordering a highly beta product from a company that’s probably run out of someone’s actual garage, but son-of-a-bitch I’M IN!!
From what I can tell it’s Raspberry Pi 5 driven for control loop, with remote inference via WiFi. I plan on connecting it to my DGX Spark.
Here’s their site:
And their SDK doc site as well:
We just released the first version of our Creative Writing benchmark, comparing 24 LLMs against human writers across 475 writing prompts.
Creative writing is subjective, so the rankings aren't meant to predict what any one person will prefer. Instead, they predict what a large group of readers would prefer, using a custom reward model trained specifically on human preferences for creative writing.
Surprisingly, the strongest frontier models already rank above the talented amateur writer cohort, while professional writers still lead by a wide margin.
You can browse the full benchmark, compare the model outputs side by side, and see how the benchmark works here:
https://vulsar.ai/benchmarks/creative-writing-v1/
Curious what you all think of the results!
\*8bit, vllm, 4x dgx\*
Prompt:
Go download and use Remotion and create a cool 60-second motion graphics video with it. Impress me totally. The motion graphics must be based on stock market visuals. Add many cool, mind-blowing motion graphics and explain what fundamental vs. technical analysis is in investing, along with their pros and cons.
\*No special skill.md used from my side.
\*harness - Zcode
It's ofc better than my previous method of making videos though java.
I think maybe in future with high bandwidth flash , we will be running these size models on affordable hardware.
How it made it report: https://github.com/9r4n4y/ProjectsorSkills/blob/main/Video\_Generation/Remotion/Making-of-Market-Decoded\_Build-Documentation.pdf
Top 3 papers on HF Daily Paper are all unusually delightful and interesting reads for anyone on the leading edge of local LLMs, agent harness optimization, etc, felt like sharing.
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
https://huggingface.co/papers/2609.19969
Cross-layer KV reuse plus FP4 KV caching brings the global KV cache to 890 bytes per token, about a quarter of V4-Flash.
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
https://huggingface.co/papers/2609.20519
Auto-research loops that improve the agent harness, cutting token traffic by 44.7 to 49.0% at comparable performance.
An Empirical Study of Harness Design for Coding Agents
https://huggingface.co/papers/2609.20804
Varies planning, action space, and context management across 176 settings to see what each component actually contributes.
I'm still reading through, feel free to discuss
Planning, coding, testing = 8h total.
Stack: VSCode Copilot in autopilot mode + SGLang
Stats: ∼10k lines generated, ∼800k tokens consumed
Sure, it's not GPT-6 Astra level, but for a 100% local ∼180B MoE running on a single DGX Spark at ∼35 tok/s. Not bad...