20 posts · 1 sub · RSS
← prev Friday, September 18, 2026 next →
posted hourdayweekmonthyearall
allr/LocalLLaMA
▲
65
+1
26👁
r/LocalLLaMA · u/Porespellar · 22d ago
NGL, I’m hyped to see if Qwen3.8 27b can make me a sandwich. Instant buy for me.

Saw this little dude in a Forbes article (https://www.forbes.com/sites/johnkoetsier/2026/08/18/american-humanoid-robot-…)

This is definitely for the DIY researcher crowd who want to dip their toes into the robotics world. It is not going to be “consumer-ready” in any way shape or form, but I Instantly preordered the shit out of this anyways. I don’t even care that all the vids of it doing stuff are probably 3x sped up and completely cherry-picked and highly edited. I DO NOT CARE that it is going to be likely absolute trash getting started with this thing. It is still going to be fucking amazing that I’m going to have a robot that could potentially injure me for $1,688.

There is very little online about this guy, but I trust Forbes did their homework, and Nori has actually supposedly delivered the first batch of the earlier L3 version from what I can tell, and their discord is active and their SDK and documentation seems legit to me. I know I’m taking a risk with pre-ordering a highly beta product from a company that’s probably run out of someone’s actual garage, but son-of-a-bitch I’M IN!!

From what I can tell it’s Raspberry Pi 5 driven for control loop, with remote inference via WiFi. I plan on connecting it to my DGX Spark.

Here’s their site:

https://www.norirobotics.com

And their SDK doc site as well:

https://docs.norirobotics.com

▲
268
-4
18👁
▲
57
-1
23👁
r/LocalLLaMA · u/Character-Result-281 · 22d ago
Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark post image

Planning, coding, testing = 8h total.
Stack: VSCode Copilot in autopilot mode + SGLang
Stats: ∼10k lines generated, ∼800k tokens consumed

Sure, it's not GPT-6 Astra level, but for a 100% local ∼180B MoE running on a single DGX Spark at ∼35 tok/s. Not bad...

▲
458
-3
40👁
r/LocalLLaMA · u/returnity · 22d ago
Is HF starting to move against abliterated models?
Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.
The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The scale of the problem is massive: Hugging Face, which hosts open source AI models, currently lists over 6,000 abliterated models.

I can't really tell what exactly the implications are of this "partnership" or what it exactly would impact on HF's model-hosting side. However, I do find it concerning that HF is announcing a collaboration on 'infrastructure safety' with publicity that specifically calls out "dangerous" uncensored models. Thoughts?

💬 213 (-1) open on reddit ↗
▲
62
-4
28👁
r/LocalLLaMA · u/drooolingidiot · 22d ago
We benchmarked 24 LLMs against human writers on 475 creative writing prompts post image

We just released the first version of our Creative Writing benchmark, comparing 24 LLMs against human writers across 475 writing prompts.

Creative writing is subjective, so the rankings aren't meant to predict what any one person will prefer. Instead, they predict what a large group of readers would prefer, using a custom reward model trained specifically on human preferences for creative writing.

Surprisingly, the strongest frontier models already rank above the talented amateur writer cohort, while professional writers still lead by a wide margin.

You can browse the full benchmark, compare the model outputs side by side, and see how the benchmark works here:

https://vulsar.ai/benchmarks/creative-writing-v1/

Curious what you all think of the results!

💬 77 (+1) open on reddit ↗
▲
168
+3
16👁
r/LocalLLaMA · u/ali_byteshape · 22d ago
Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison post image

Hey r/LocalLLaMA,

Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark.

We wanted to see how they compare under the same methodology we use for the rest of our Qwen3.8 evaluations, so we ran Bonsai 2 through our own benchmark suite.

One important clarification: these are our evaluation results, not Prism’s reported benchmark numbers.

We used Prism’s fork/runtime to run their models, while keeping the workloads, benchmark suite, and evaluation methodology consistent with the rest of our comparison.

Our evaluation includes separate Instruct and Thinking benchmarks. For Thinking, we use medium thinking effort with the recommended sampling parameters.

We added Bonsai 2 because the models have become a relevant part of the Qwen3.8 landscape, and we wanted to provide a common reference point for people comparing the available options.

Different providers often report results using different benchmark suites, runtimes, reasoning settings, sampling parameters, and evaluation methodologies, so those numbers are not always directly comparable. Running the models through the same evaluation gives another reference point for understanding the trade-offs between quality, model size, and throughput.

Updated comparison and results: https://byteshape.com/blogs/Qwen3.8-27B/

▲
121
-3
20👁
r/LocalLLaMA · u/jacek2023 · 22d ago
inclusionAI/Realtime-Venus · Hugging Face

do you want some omni? here is omni for you

[](https://huggingface.co/inclusionAI/Realtime-Venus#1-🧭-overview)1. 🧭 Overview

This repository hosts two checkpoints of the Realtime-Venus system:

  • Realtime-Venus-Omni (Realtime-Venus-Omni/): the 9B audio-visual interaction model. It continuously watches and listens, decides whether and when to respond, and generates text and speech on a shared causal timeline. Adapted from MiniCPM-o 4.5, it supports proactive interaction, semantic interruption handling, and training-free long-video memory.
  • Realtime-Venus-Audio (Realtime-Venus-Audio/): the audio-focused checkpoint on the same streaming backbone, for audio understanding and audio-driven conversation with text or speech output.

Both directories contain model weights and custom Hugging Face Transformers code. The asynchronous Realtime-Venus-Harness and its external tool integrations live in the GitHub repository.

[](https://huggingface.co/inclusionAI/Realtime-Venus#2-✨-highlights)2. ✨ Highlights

  • Native full-duplex conversation: keeps perceiving while speaking and distinguishes backchannels, interruptions, corrections, and redirections.
  • Omni-Proactive interaction: continuously processes temporally aligned video and audio, and initiates a response when an event warrants it — without waiting for a user prompt.
  • Delegation: emits in-stream <delegate> requests on the shared causal timeline and consumes asynchronous backend results the same way, so external tasks never block the ongoing conversation. (Executing requests requires the Realtime-Venus-Harness runtime, available in the GitHub repository.)
  • Training-free long-video Memory: archives visually informative moments, retrieves query-relevant and non-redundant evidence, and reassembles the corresponding audio-visual context — no additional training required.
  • Text and speech output: generates response text together with native speech through the bundled Token2wav resources and a reference voice.
▲
137
-3
22👁
r/LocalLLaMA · u/No_Issue_8224 · 22d ago
MiniMax Code goes open source

MiniMax has open-sourced the terminal version of MiniMax Code:

https://github.com/MiniMax-AI/minimax-code

How can developers verify the content that encoding proxies read, send, and store? This is a topic that has been widely discussed recently.

Open sourcing the agent doesn’t automatically answer every privacy or security question, but it gives the community something concrete to inspect.

The repository includes:

  • interactive TUI and headless execution
  • code editing, shell commands, diffs, and test verification
  • permission controls and sandboxing
  • Plan Mode and resumable sessions
  • subagents, plugins, skills, and MCP
  • BYOK with OpenAI- and Anthropic-compatible providers
  • ACP support for compatible editors and clients

First-party code defaults to the MIT license.

A few important caveats: this is a 0.4.12 source preview, the desktop app source is not included, and—as the repository itself notes—a matching version number does not prove identical build provenance between the published package and source checkout.

Still, releasing the agent layer is a meaningful step toward auditability. I’d like to see the community examine its network behavior, file-access boundaries, telemetry, and reproducible-build story next.

https://preview.redd.it/zvrakmgejaqh1.png?width=1198&format=png&auto=…

https://preview.redd.it/13pjdngejaqh1.png?width=1206&format=png&auto=…

https://preview.redd.it/z1ekjlgejaqh1.png?width=1200&format=png&auto=…

▲
1017
 
41👁
r/LocalLLaMA · u/segmond · 22d ago
768gb vram for less than the price of one RTX 6000

I have always posted about budget builds on here, and often asked how we are going to run the next big models. Often Plenty of downvotes too or folks telling me that it's not running if I'm getting 5tk/sec. But whatever, the hunger and desire to go big has always kept me on the edge and looking for deals.

Here's my latest build, 12x64gb cmp170hx. For less than 1 RTX 6000 pro costs. I also have it connected with fiber to my other rig for RPC when I need more memory. I haven't been posting much since I built this rig, because it's now more fun to talk to my machine. I run GLM5.3, DSv4.1Flash, Qwen3.8Flash, Qwen3.8-2.4T, KimiK3 and MiniMaxM3. Performance is great, a single RTX 6000 or M3 Mac Studio wish they could. Inference with vllm or llama.cpp

I look forward reading the replies how API usage is cheaper, or how it will take 52 light years to break even or the noise, or the electrical cost. NOT.

There will be more opportunities in the future, keep looking for them and pounce on them when they come. up, the demand is going to be high for compute for a long time.

https://preview.redd.it/dunixwu6caqh1.jpg?width=4080&format=pjpg&auto…

https://preview.redd.it/glpcbg5cbaqh1.jpg?width=3072&format=pjpg&auto…

💬 399 (+1) open on reddit ↗
▲
160
+4
28👁
r/LocalLLaMA · u/Secure_Recording_472 · 22d ago
Question: UkisAI Swift Ternary Bonsai 2 27B?

Hey community,

Jovan from UkisAI here,

We're the team behind Swift Qwen3.8 27B, the Qwen model with token usage and overthinking error improvements

Our estimate is that we can make a great improvement to Bonsai 2, as our testing indicates that it suffers greatly from overthinking loops and in general high token usage impacting it's performance.

My ask for you is:

Is a Swifted version of Bonsai 2 something you guys would enjoy?

If yes, what size is the most relevant. 1-bit, 2-bit or both?

Thank you for the amazing feedback on Swift. We are glad you are enjoying it. Our download count jumped from 100k -> 150k overnight (community quants included).

For context, this is our model: https://www.reddit.com/r/LocalLLaMA/s/iCIbhxO8ue

https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF

💬 181 (+1) open on reddit ↗
▲
152
-2
25👁
r/LocalLLaMA · u/KURD_1_STAN · 22d ago
bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai\_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai\_2/q3.5 60.8 / 72.4 = 84.0%
Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release \[2\], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

^(also 3.8 35b qwhen? plsss)

▲
192
+3
16👁
▲
102
+1
27👁
r/LocalLLaMA · u/Every-Comment5473 · 22d ago
Still on the Jev waitlist? I hosted OpenJev. It's free, go play with it

TypeSafe announced Jev on Tuesday: you give it data plus typed questions (yes/no, pick-one, 0–N scale) and it returns a probability for every option, crazy fast. I signed up and then refreshed my inbox. A lot.

Meanwhile Matt Mastracci opened vLLM PR #57250, which does the same trick on DiffusionGemma with a single denoising step. The model basically fills in a multiple-choice bubble sheet. My "quick look" turned into three straight days, and now there's OpenJev: an open-source server with Jev's API, so TypeSafe's SDKs work with just a base URL change. If you're still waiting on Jev access, you can start playing today.

Your prompts and answers are not stored, only token counts for your quota. It runs on my RTX PRO 6000, which just got promoted to "production infrastructure" overnight!

Is it any good? Matt ran live evals of Jev vs DiffusionGemma-as-Jev: accuracy roughly tied (198/201 vs Jev's 191/201 across his 8 eval sets), and DiffusionGemma was faster, on a DGX Spark. An RTX PRO 6000 is a different animal:

|Model|Latency|
|:-|:-|
|Frontier LLMs (TypeSafe's numbers)|3–329 s (coffee time)|
|Jev (published)|70–500 ms end to end|
|OpenJev via api.codiv.ai|\~170 ms p50 end to end (\~73 ms on the GPU)|

It's v0.1 on an unmerged vLLM PR. If Reddit hugs it to death you'll see 529s, which is my GPU asking for a minute.

Credits: Matt Mastracci (the core idea and vLLM work are his), TypeSafe (the System One idea and API), NVIDIA and Google (DiffusionGemma), and the vLLM team.

Just a fan of TypeSafe's idea, not affiliated. Feedback, bugs, use-case ideas, or your weirdest yes/no question, all welcome!

💬 30 (+1) open on reddit ↗
▲
95
-2
13👁
▲
263
-3
22👁
▲
105
-4
22👁
r/LocalLLaMA · u/MeinDruckerSpinnt · 22d ago
JEV architecture

My understanding so far:

  1. You take an LLM and use it without thinking (That's what openjev does?)
  2. You leave out the text generation in the end and take the confidence score in the matrix before that phase

That's it. Right?

They gave it a mysterious marketing name.

▲
157
+5
33👁
r/LocalLLaMA · u/mateszhun · 22d ago
Qwen 3.8 Next Flash appreciation post

I don't want to talk about the performance and technical things, but about how I work with my hobby projects has changed thanks to this model.

My machine has generated around 25M tokens since the model came out, and I've done a mental retro on it.

I've found it to have incredible prompt adherence. I can leave to run it by itself and get back to it, and find that it did exactly what I've asked it to. I've only ever seen it go astray once, where I've asked something that is too high level and filled out the context window (It is shitty at context compacting, maybe that is the Q4 at play).
It can solve medium complexity tasks by itself, if you prompt it in a way to use subagents, do some research, planning, review and testing it does really well.

It has absolutely raised the floor for me on what I expect a model to be capable of. And that is a huge thing. It won't discover then next scientific breakthrough or be as amazing as Astra at computer use, but it is very consistent in what it can do, and does not screw up trivial things.

I can give it conditions for actions and will orchestrate according to it.
It has raised the bar in my work as well, not just at home hobby projects. I'm absolutely amazed by it.

I absolutely want coding models to improve along this line. Raising the floor, and prompt adherence is a great value in coding.

💬 77 (+6) open on reddit ↗
▲
1041
+6
46👁
r/LocalLLaMA · u/Nandakishor_ml · 22d ago
Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo post image

UPDATE: Multilingual support added at : https://github.com/NandhaKishorM/laya

Thanks for the exceptional support (https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i\_literally\_built\_the…) and for the dozens of requests to make a generic model, run benchmarks, and create an HF space so anyone can test it. So here you go, guys. I trained an improved model on a large data corpus, its now called Laya. It is trained on a single RTX 6000 Pro (96 GB VRAM); the model architecture is a 421M-parameter non-autoregressive decision model pairing a bidirectional ModernBERT-large encoder with a scratch Transformer head that scores \[MASK\] option markers to resolve typed schemas in a single \~35 ms forward pass. The dataset is a 100% human-annotated corpus of over 25,000 real-world examples across intent routing, fact-checking, moderation consensus, prompt guardrails, rubric scoring, and multi-turn conversation trajectories, without synthetic data shortcuts. The RLCD(unofficial, btw) I did is a policy-gradient reinforcement learning approach that kinda optimizes decision models against strictly proper scoring rules, ensuring maximum reward is achieved only when outputting true, mathematically calibrated probabilities.

NB: It can be run on low end PC as its a small 421M model, cheers

HF space to try: https://huggingface.co/spaces/convaiinnovations/laya-demo

GitHub Repo: https://github.com/NandhaKishorM/laya

HF Repo: https://huggingface.co/convaiinnovations/laya

Thank you to everyone who supported me, shared the story, gave personal DM. It will need more refinement, of course.

If anyone wishes to buy me a coffee, here is the link: https://github.com/NandhaKishorM

💬 157 (+1) open on reddit ↗
▲
125
+4
17👁
▲
178
+3
25👁
r/LocalLLaMA · u/AdRepulsive7837 · 22d ago
still doesn’t get what Jev is…..is it just a more generalised BERT?

Looking at jev launch website and demo video on x.com…. it seems like it’s a very intelligent classifier with custom prompt and custom criteria instruction reading capabilities. It can do well defined narrow and well defined task

Me, following NLP since good old days of word embedding and BERT,,, be like asking….

Isn’t that BERT?

yeah i know BERT need fine tuning to adapt to custom domain, but can Jev be like generalised form of BERT?