37 posts · 1 sub · RSS
← prev Sep 17, 2026 → Sep 18, 2026 next →
2026-09-17 → 2026-09-18 hourdayweekmonthyearall
allr/LocalLLaMA
▲
3582
+97
72👁
r/LocalLLaMA · u/Nandakishor_ml · 23d ago
I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Update:
I made a generic model and beaten the jev in all of the benchmarks. Code and details available at
https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a

frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. For anyones information the main guiding model is RL not embedding model or LLM

Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43

Paper: https://arxiv.org/abs/2503.23303

Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

Also the second work published in September 2025 was exactly the same one jev proposed now

Paper: https://arxiv.org/abs/2510.01237

My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).

Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.

It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂

💬 339 (+9) open on reddit ↗
▲
1678
+11
35👁
r/LocalLLaMA · u/xenovatech · 22d ago
Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU. post image

The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to <6GB in size. According to the model card, it's 9x smaller than FP16 while retaining 98.2% of the intelligence.
\- Collection: https://huggingface.co/collections/prism-ml/bonsai-2
\- Demo: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels

▲
1617
+9
48👁
r/LocalLLaMA · u/Nandakishor_ml · 23d ago
I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Update:
I made a generic version. Full details at https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs
It includes code, benchmark and hf repo

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a

frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. Links are. For anyones information the main guiding model is RL not embedding model or LLM

Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43

Paper: https://arxiv.org/abs/2503.23303

Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

Also the second work published in September 2025 was exactly the same one jev proposed now

Paper: https://arxiv.org/abs/2510.01237

My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).

Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.

It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂

💬 136 (+2) open on reddit ↗
▲
1044
+9
43👁
r/LocalLLaMA · u/Nandakishor_ml · 22d ago
Made the horizontal open-source model for Jev with RLCD, and it surpasses all the Jev benchmarks. HF space, benchmark, model, repo post image

UPDATE: Multilingual support added at : https://github.com/NandhaKishorM/laya

Thanks for the exceptional support (https://www.reddit.com/r/LocalLLaMA/comments/1wijo3e/i\_literally\_built\_the…) and for the dozens of requests to make a generic model, run benchmarks, and create an HF space so anyone can test it. So here you go, guys. I trained an improved model on a large data corpus, its now called Laya. It is trained on a single RTX 6000 Pro (96 GB VRAM); the model architecture is a 421M-parameter non-autoregressive decision model pairing a bidirectional ModernBERT-large encoder with a scratch Transformer head that scores \[MASK\] option markers to resolve typed schemas in a single \~35 ms forward pass. The dataset is a 100% human-annotated corpus of over 25,000 real-world examples across intent routing, fact-checking, moderation consensus, prompt guardrails, rubric scoring, and multi-turn conversation trajectories, without synthetic data shortcuts. The RLCD(unofficial, btw) I did is a policy-gradient reinforcement learning approach that kinda optimizes decision models against strictly proper scoring rules, ensuring maximum reward is achieved only when outputting true, mathematically calibrated probabilities.

NB: It can be run on low end PC as its a small 421M model, cheers

HF space to try: https://huggingface.co/spaces/convaiinnovations/laya-demo

GitHub Repo: https://github.com/NandhaKishorM/laya

HF Repo: https://huggingface.co/convaiinnovations/laya

Thank you to everyone who supported me, shared the story, gave personal DM. It will need more refinement, of course.

If anyone wishes to buy me a coffee, here is the link: https://github.com/NandhaKishorM

💬 157 (+1) open on reddit ↗
▲
395
+6
35👁
r/LocalLLaMA · u/FullstackSensei · 22d ago
AMD Plans 10% Price Hike Across GPUs, Chipsets, and Possibly CPUs

Great news! AMD is also considering accepting payment in organs!

Slightly less sarcastically, grab what you can, while you can. Waiting is becoming very costly almost by the day

▲
157
+5
33👁
r/LocalLLaMA · u/mateszhun · 22d ago
Qwen 3.8 Next Flash appreciation post

I don't want to talk about the performance and technical things, but about how I work with my hobby projects has changed thanks to this model.

My machine has generated around 25M tokens since the model came out, and I've done a mental retro on it.

I've found it to have incredible prompt adherence. I can leave to run it by itself and get back to it, and find that it did exactly what I've asked it to. I've only ever seen it go astray once, where I've asked something that is too high level and filled out the context window (It is shitty at context compacting, maybe that is the Q4 at play).
It can solve medium complexity tasks by itself, if you prompt it in a way to use subagents, do some research, planning, review and testing it does really well.

It has absolutely raised the floor for me on what I expect a model to be capable of. And that is a huge thing. It won't discover then next scientific breakthrough or be as amazing as Astra at computer use, but it is very consistent in what it can do, and does not screw up trivial things.

I can give it conditions for actions and will orchestrate according to it.
It has raised the bar in my work as well, not just at home hobby projects. I'm absolutely amazed by it.

I absolutely want coding models to improve along this line. Raising the floor, and prompt adherence is a great value in coding.

💬 77 (+6) open on reddit ↗
▲
229
+4
23👁
r/LocalLLaMA · u/Henrie_the_dreamer · 22d ago
Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash post image

Hey all, Henry from Cactus Compute here, I kinda wanted to share our latest model and get feedback from the family :)

Needle 3 is a small foundation model for automation: you give it the functions your app exposes, it reads a request and returns the calls with every argument filled in, or a typed record if what you gave it was a schema. It runs on the device, with no network in the loop. It is on Hugging Face, on GitHub, on PyPI as cactus-needle, and there is a sandbox that runs it in your browser at cactuscompute.com/needle if you want to poke at it before reading further.

1) Trades general capacity for frontier performance on automation tasks

The thing we decided early was that Needle would not chat. Every turn is a function call, and a request no declared tool can serve comes back as an empty list rather than a guess. That sounds like a limitation, and it is, but it is what let a 121M-parameter model be trained on 360B tokens of structured data and spend all of its capacity on three jobs: tool calls, structured extraction and text embedding.

The architecture follows from the same trade. It is a Simple Attention Network: the dense feed-forward layers are gone, replaced by a Monarch Hadamard MLP with 25.6K parameters per layer instead of 4.7M, and the knowledge a feed-forward layer would normally hold sits in an engram, hashed n-gram tables that are read by gather and cost no arithmetic. 70.8M of the 121M parameters live there, so the full model does the arithmetic of a 50M one: 100 MFLOPs per token against 296 for a transformer of the same shape.

https://i.redd.it/fsnfjojny4qh1.gif

We wrote the intuition up if you want the longer version: Simple Attention Networks and the Hadamard MLP.

2) Beats models 10x its size on tool calls and language-to-device control

On Mobile Actions (961 phone commands, scored on the exact call), the 20-layer model scores 86.0 through the shipped 2-bit binary with the confidence gate on. LFM2.5 1.2B is at 82.4, Qwen3.5 0.8B at 76.0, FunctionGemma 270M at 65.1 and Apple's on-device foundation model at 57.6, all at f16. DeepSeek V4 Flash through its API is at 88.4, which is the line in the chart.

https://i.redd.it/luunq717z4qh1.gif

The part we are most pleased with is not the number but how the calls are made. Every argument is a span of the request: the model writes a short derivation first ('living room' -> room; '30' -> brightness) and then emits the call under a byte-level grammar compiled from your schema, so the JSON always parses and an enum can never leave its set. An optional field with no evidence is omitted, a required one with no evidence withholds the call, and the engine drops a call the request negates or excludes. Ask for two things and you get two calls in order.

https://i.redd.it/227396y5z4qh1.gif

Full table across all six suites (tool calling is exact match, extraction is field F1, Needle through the shipped binary, baselines at f16 under vLLM):

|Model|Params|Mobile Actions|DroidCall|BFCL v4|DSTC8 F1|SNIPS gold F1|SNIPS 7-way F1|
|:-|:-|:-|:-|:-|:-|:-|:-|
|DeepSeek V4 Flash (cloud)|\-|88.4|60.5|77.2|80.0|69.4|66.7|
|Needle3-20L-121M|121M|86.0|47.0|50.2|40.7|30.2|24.7|
|LFM2.5 1.2B|1.2B|82.4|35.5|62.0|48.0|43.0|38.0|
|Needle3-16L-98M|98M|80.7|40.0|41.3|28.5|23.5|19.2|
|Qwen3.5 0.8B|800M|76.0|28.0|56.8|49.0|35.0|34.0|
|LFM2.5 350M|350M|72.8|32.5|59.1|20.0|34.0|29.0|
|LFM2.5 230M|230M|69.3|11.5|46.3|53.0|27.0|22.0|
|FunctionGemma 270M|270M|65.1|16.5|46.6|27.0|29.0|14.0|
|Needle 2|45M|63.5|17.0|\-|\-|\-|\-|
|Apple FM|3.0B|57.6|\-|\-|\-|\-|\-|
|Needle3-8L-52M|52M|36.8|36.5|28.2|15.3|16.6|10.1|
|Needle3-4L-29M|29M|11.7|21.0|19.5|6.9|7.7|4.3|

You can see where it is weaker too: BFCL and the extraction suites are where the bigger baselines pull ahead, and the smaller subnetworks fall off quickly on the general task (more on why that is fine in section 4).

3) Matches 2-3x bigger models on structured JSON extraction

Extraction is not a separate mode. You declare the record as the only tool and pass the passage where the query goes; with one tool declared the grammar admits exactly one call of that name, so the shape is guaranteed rather than requested, and the values are grounded the same way as arguments: a field is filled only from a span of the passage, an optional field with no span comes back as None, and a date whose year appears nowhere in the text is flagged instead of invented.

from pydantic import BaseModel
import needle

class Invoice(BaseModel):
vendor: str
total: float
due_date: str
po_number: str | None = None

needle.extract("Invoice from Acme Corp, $1,200.00, due 2026-09-01", Invoice)
# Invoice(vendor='Acme Corp', total=1200.0, due_date='2026-09-01', po_number=None)

It generalised to classification without special training, because an enum is just a constrained value: declare sentiment: Literal["positive", "neutral", "negative"] on a record and you have a classifier whose output cannot leave the set. A watch reads a notification into merchant, amount and date that way, then into a reply, then into a sentiment flag, one record each. On DSTC8 and the two SNIPS suites the 121M model lands between the 230M and 350M baselines, which is the 2-3x in the heading.

4) Intelligence ladder: every depth from 2 to 20 layers a model of its own

This is the part I would most like your thoughts on. Needle 3 is one set of weights, and every depth from 2 to 20 layers is a deployable model. Blocks 0 and 19 are always kept and the rest are added by bisection, so each subnetwork nests in the next; during training each step samples one path, mostly the full model and otherwise a random depth, with the smaller path distilled from the full one. The full-depth model ends up slightly better than an ordinary run of the same size, and every depth below it is trained rather than truncated.

https://i.redd.it/kbwimep4z4qh1.gif

Why we wanted it: a watch, a Raspberry Pi and a phone do not want the same model, and they want to pick the size at deploy time. needle build --layers 8 writes the 8-layer file; the same engine runs all of them. The small depths lose accuracy on the general benchmarks (that is the bottom of the table above), and they get it back when fine-tuned to one product's tools: on DroidCall every subnetwork gains 18 to 36 points, and from 4 layers (29M parameters) up the tuned subnetwork passes DeepSeek V4 Flash.

https://i.redd.it/y646wed3z4qh1.gif

Fine-tuning is LoRA on the frozen base, merged at export, and the Python package does it locally at 4 bits (needle finetune data.jsonl, then needle build). The maths is in Intelligence Ladders and the workflow in Fine-tuning Needle.

5) Runs locally at up to 4k tokens/sec decode speed

The engine is under 1 MB, plain CPU, no GPU or NPU, and the weights are read in place from a single file the engine maps into memory: a 196-byte header carrying the whole architecture geometry, a nameless tensor directory, and the quantised blobs in the order the forward pass reads them. On a Raspberry Pi 5, decode runs at up to 4k tokens/s at the bottom of the ladder and around 400 at the top, prefill from 10k down to 1k. Every response reports prefill_tps, decode_tps and peak_ram_mb, so you can measure on your own device rather than take our word for it.

Every response also carries a confidence score from a calibrated head, the minimum of a post-hoc judgement on the finished call and the decode probability of its tokens. The engine withholds anything under 0.1; above that the number is yours: act at once when it is high, show the call and ask when it is middling, treat [] as a refusal. How we use it is in Leveraging Needle's confidence.

6) 25-121M deployable parameters at CQ2-bit (8-29MB binaries)

The weights are quantised with Cactus Quants: groups of 128 weights are rotated by a Walsh-Hadamard matrix, which makes every group look Gaussian, split into an fp16 norm and a direction on the unit sphere, and the direction's coordinates are snapped to a 4-entry Lloyd-Max codebook. That is 2.125 bits per weight, and the kernel never expands them: it rotates and int8-quantises the activation instead, then does table lookups and sdot against the packed indices. The embedding and the confidence head keep 4 bits, the norms and gates stay fp16. The byte layout, and a twenty-line parser for it, are in The .cact format.

7) For mobiles, wearables, smart home, small robots and microcontrollers

Every target ships a prebuilt engine folder: macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32 (the Ingenic camera SoCs), Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly, and a WASI component with a WIT world. pip install cactus-needle covers the desktop and server platforms with wheel-tagged engines, and needle build --platform linux-arm64 --layers 8 --out ./pi puts an engine and the weights in a folder you copy over. Inference never touches the network, so an air-gapped device only needs the files in place. The list and the runtime surfaces are in What devices are supported.

import needle

@needle.tool
def get_weather(city: str):
"Get the current weather for a city."
return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]

What we would love feedback on

  • Where the grounding rules get in your way. The engine refuses to invent a number, drops a call the request excludes and withholds a required enum the request never names; we tuned those on our suites and would like to hear where they bite on real tools.
  • The ladder. Whether depth is the right knob for you, or whether you would rather have width, and what devices you would want the 2- and 4-layer models on.
  • Extraction cases we have not seen. Nested records, arrays, multilingual text (it is English-first, and non-English text fragments into about 1.7x more tokens).
  • Anything in the tool design guide that turned out wrong for your schemas.

Thanks for reading this far. Happy to answer anything in the comments.

▲
205
+4
29👁
r/LocalLLaMA · u/whodoneit1 · 22d ago
153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s post image

People kept commenting and asking about single AMD 1xR9700 cards in the comments and discord. Well, I finally had time to do some optimizations for 1xR9700 owners and performance has doubled across the board. You can see the results in BetterBench above if you like visuals or below if you're more into text.

These results were measured running Unsloth's Qwen3.8 27b NVFP4.

Decode
┌───────────────┬───────────────┬──────────────────┐
│ category │ update p50 ms │ decode t/s (med) │
├───────────────┼───────────────┼──────────────────┤
│ chat │ 42.3 │ 67.1 │
├───────────────┼───────────────┼──────────────────┤
│ code │ 42.5 │ 120.5 │
├───────────────┼───────────────┼──────────────────┤
│ file_edit │ 42.5 │ 138.0 │
├───────────────┼───────────────┼──────────────────┤
│ json │ 42.4 │ 153.1 │
├───────────────┼───────────────┼──────────────────┤
│ math │ 42.5 │ 140.0 │
├───────────────┼───────────────┼──────────────────┤
│ prose │ 42.3 │ 69.2 │
├───────────────┼───────────────┼──────────────────┤
│ reasoning │ 34.3 │ 123.9 │
├───────────────┼───────────────┼──────────────────┤
│ summarization │ 34.2 │ 141.7 │
└───────────────┴───────────────┴──────────────────┘

Prefill
┌───────────────┬───────────────┐
│ prefill depth │ pp tok/s │
├───────────────┼───────────────|
│ 2000 │ 3552 │
├───────────────┼───────────────|
│ 8000 │ 3536 │
├───────────────┼───────────────|
│ 16000 │ 3619 │
├───────────────┼───────────────|
│ 32000 │ 3437 │
├───────────────┼───────────────|
│ 64000 │ 3192 │
├───────────────┼───────────────|

Concurrency
┌───────────────┬───────────────┐
│ level │ tok/s │
├───────────────┼───────────────|
│ 1 │ 120 │
├───────────────┼───────────────|
│ 2 │ 215 │
├───────────────┼───────────────|
│ 4 │ 322 │
├───────────────┼───────────────|
│ 8 │ 471 │
├───────────────┼───────────────|

Links (Both repo's updated as some users wanted Github)

https://codeberg.org/ggz14/radiance-vllm-mxfp4

https://github.com/GGZ14/vllm-mxfp4

https://x.com/bkuyper

I hope you single R9700 card owners enjoy this release!

💬 120 (+2) open on reddit ↗
▲
125
+4
17👁
▲
99
+4
22👁
r/LocalLLaMA · u/Zulfiqaar · 22d ago
IFM/K2-Horizon-7B-Uno · Hugging Face - 5200tps with no quality loss

IFM released K2-Horizon-7B , diffusion augmented LLM at upto 5200 tps, with claimed lossless speedup.

causal LLM architecture and adds a plug-and-play diffusion adapter alongside the autoregressive weights.

https://arxiv.org/abs/2609.04010

▲
69
+4
21👁
▲
192
+3
16👁
▲
178
+3
25👁
r/LocalLLaMA · u/AdRepulsive7837 · 22d ago
still doesn’t get what Jev is…..is it just a more generalised BERT?

Looking at jev launch website and demo video on x.com…. it seems like it’s a very intelligent classifier with custom prompt and custom criteria instruction reading capabilities. It can do well defined narrow and well defined task

Me, following NLP since good old days of word embedding and BERT,,, be like asking….

Isn’t that BERT?

yeah i know BERT need fine tuning to adapt to custom domain, but can Jev be like generalised form of BERT?

▲
164
+3
23👁
r/LocalLLaMA · u/Jorlen · 22d ago
Does anyone use uncensored models purely for coding?

It sounds like a stupid question, and I do apologize if it is... but I've seen several people mention that coding models are better uncensored due to the fact that they don't have to constantly run prompts through the "is this okay" sort of checks.

Is this hogwash? Is it true? And more importantly, does anyone have any sources to confirm it?

Anecdotal evidence is fine too if you've tried and compared them.

Personally, I have never bothered because I'm too worried the de-censoring would damage the weights. The juice never felt like it was worth the squeeze... but maybe I was wrong?

Edit: Either I'm unclear or people are misinterpreting my request: Specifically, I mean for every day coding (not for hacking, not for reverse-engineering) but just for regular coding of new apps, etc. The question is: will the uncensored model produce better code faster (without reasoning so much) because it no longer has to worry about "is this alright" when it questions everything...

Edit 2: Decided to test the HuiHui Qwen 3.8 27b (UD-Q8\_K\_XL) quant myself. So far, it reasons far less, and I have yet to have any issues with its coding quality. Granted, I've only been testing it for about 8 hours (straight...) in an active project. Its reasoning is far shorter, it seems far more confident in its responses and as such, uses far less context to achieve the same result. I will continue testing for another week; it's pitted against the Dirk template version of Qwen 3.8 27b right now (same quant) which I'd been using the previous week.

▲
143
+3
21👁
r/LocalLLaMA · u/No_Issue_8224 · 21d ago
MiniMax Code goes open source

MiniMax has open-sourced the terminal version of MiniMax Code:

https://github.com/MiniMax-AI/minimax-code

How can developers verify the content that encoding proxies read, send, and store? This is a topic that has been widely discussed recently.

Open sourcing the agent doesn’t automatically answer every privacy or security question, but it gives the community something concrete to inspect.

The repository includes:

  • interactive TUI and headless execution
  • code editing, shell commands, diffs, and test verification
  • permission controls and sandboxing
  • Plan Mode and resumable sessions
  • subagents, plugins, skills, and MCP
  • BYOK with OpenAI- and Anthropic-compatible providers
  • ACP support for compatible editors and clients

First-party code defaults to the MIT license.

A few important caveats: this is a 0.4.12 source preview, the desktop app source is not included, and—as the repository itself notes—a matching version number does not prove identical build provenance between the published package and source checkout.

Still, releasing the agent layer is a meaningful step toward auditability. I’d like to see the community examine its network behavior, file-access boundaries, telemetry, and reproducible-build story next.

https://preview.redd.it/zvrakmgejaqh1.png?width=1198&format=png&auto=…

https://preview.redd.it/13pjdngejaqh1.png?width=1206&format=png&auto=…

https://preview.redd.it/z1ekjlgejaqh1.png?width=1200&format=png&auto=…

▲
463
+2
38👁
r/LocalLLaMA · u/returnity · 21d ago
Is HF starting to move against abliterated models?
Baseten launched a new safety infrastructure standard alongside its Base Labs research arm on Wednesday, partnering with Hugging Face and Goodfire AI to build safety evaluation and monitoring infrastructure for open-weight models.
The announcement lands amid debate for the safety of open-weight models — which can be made dangerous by removing their safeguards through a rising technique known as abliteration. The scale of the problem is massive: Hugging Face, which hosts open source AI models, currently lists over 6,000 abliterated models.

I can't really tell what exactly the implications are of this "partnership" or what it exactly would impact on HF's model-hosting side. However, I do find it concerning that HF is announcing a collaboration on 'infrastructure safety' with publicity that specifically calls out "dangerous" uncensored models. Thoughts?

💬 213 (-1) open on reddit ↗
▲
108
+2
14👁
r/LocalLLaMA · u/CharlesStross · 22d ago
600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"? post image

It's not the brightest bulb but it's my new drudgework model for read+find or code tasks I'm willing to let it brute force. Even if it takes 20x more tokens, that's still faster than many local models. Not quite Cerebras but still pretty fun to drive.

▲
166
+1
15👁
r/LocalLLaMA · u/ali_byteshape · 21d ago
Prism-ML Bonsai 2 Joins Our Qwen3.8 Quantization Comparison post image

Hey r/LocalLLaMA,

Prism-LM recently released its Bonsai 2 QAT models based on Qwen3.8, and they quickly gained traction. In our evaluation, the models strike a strong balance between throughput and quality, reaching roughly 91.5% on our composite benchmark.

We wanted to see how they compare under the same methodology we use for the rest of our Qwen3.8 evaluations, so we ran Bonsai 2 through our own benchmark suite.

One important clarification: these are our evaluation results, not Prism’s reported benchmark numbers.

We used Prism’s fork/runtime to run their models, while keeping the workloads, benchmark suite, and evaluation methodology consistent with the rest of our comparison.

Our evaluation includes separate Instruct and Thinking benchmarks. For Thinking, we use medium thinking effort with the recommended sampling parameters.

We added Bonsai 2 because the models have become a relevant part of the Qwen3.8 landscape, and we wanted to provide a common reference point for people comparing the available options.

Different providers often report results using different benchmark suites, runtimes, reasoning settings, sampling parameters, and evaluation methodologies, so those numbers are not always directly comparable. Running the models through the same evaluation gives another reference point for understanding the trade-offs between quality, model size, and throughput.

Updated comparison and results: https://byteshape.com/blogs/Qwen3.8-27B/

▲
102
+1
27👁
r/LocalLLaMA · u/Every-Comment5473 · 22d ago
Still on the Jev waitlist? I hosted OpenJev. It's free, go play with it

TypeSafe announced Jev on Tuesday: you give it data plus typed questions (yes/no, pick-one, 0–N scale) and it returns a probability for every option, crazy fast. I signed up and then refreshed my inbox. A lot.

Meanwhile Matt Mastracci opened vLLM PR #57250, which does the same trick on DiffusionGemma with a single denoising step. The model basically fills in a multiple-choice bubble sheet. My "quick look" turned into three straight days, and now there's OpenJev: an open-source server with Jev's API, so TypeSafe's SDKs work with just a base URL change. If you're still waiting on Jev access, you can start playing today.

Your prompts and answers are not stored, only token counts for your quota. It runs on my RTX PRO 6000, which just got promoted to "production infrastructure" overnight!

Is it any good? Matt ran live evals of Jev vs DiffusionGemma-as-Jev: accuracy roughly tied (198/201 vs Jev's 191/201 across his 8 eval sets), and DiffusionGemma was faster, on a DGX Spark. An RTX PRO 6000 is a different animal:

|Model|Latency|
|:-|:-|
|Frontier LLMs (TypeSafe's numbers)|3–329 s (coffee time)|
|Jev (published)|70–500 ms end to end|
|OpenJev via api.codiv.ai|\~170 ms p50 end to end (\~73 ms on the GPU)|

It's v0.1 on an unmerged vLLM PR. If Reddit hugs it to death you'll see 529s, which is my GPU asking for a minute.

Credits: Matt Mastracci (the core idea and vLLM work are his), TypeSafe (the System One idea and API), NVIDIA and Google (DiffusionGemma), and the vLLM team.

Just a fan of TypeSafe's idea, not affiliated. Feedback, bugs, use-case ideas, or your weirdest yes/no question, all welcome!

💬 30 (+1) open on reddit ↗
▲
65
+1
26👁
r/LocalLLaMA · u/Porespellar · 21d ago
NGL, I’m hyped to see if Qwen3.8 27b can make me a sandwich. Instant buy for me.

Saw this little dude in a Forbes article (https://www.forbes.com/sites/johnkoetsier/2026/08/18/american-humanoid-robot-…)

This is definitely for the DIY researcher crowd who want to dip their toes into the robotics world. It is not going to be “consumer-ready” in any way shape or form, but I Instantly preordered the shit out of this anyways. I don’t even care that all the vids of it doing stuff are probably 3x sped up and completely cherry-picked and highly edited. I DO NOT CARE that it is going to be likely absolute trash getting started with this thing. It is still going to be fucking amazing that I’m going to have a robot that could potentially injure me for $1,688.

There is very little online about this guy, but I trust Forbes did their homework, and Nori has actually supposedly delivered the first batch of the earlier L3 version from what I can tell, and their discord is active and their SDK and documentation seems legit to me. I know I’m taking a risk with pre-ordering a highly beta product from a company that’s probably run out of someone’s actual garage, but son-of-a-bitch I’M IN!!

From what I can tell it’s Raspberry Pi 5 driven for control loop, with remote inference via WiFi. I plan on connecting it to my DGX Spark.

Here’s their site:

https://www.norirobotics.com

And their SDK doc site as well:

https://docs.norirobotics.com

▲
1017
 
41👁
r/LocalLLaMA · u/segmond · 21d ago
768gb vram for less than the price of one RTX 6000

I have always posted about budget builds on here, and often asked how we are going to run the next big models. Often Plenty of downvotes too or folks telling me that it's not running if I'm getting 5tk/sec. But whatever, the hunger and desire to go big has always kept me on the edge and looking for deals.

Here's my latest build, 12x64gb cmp170hx. For less than 1 RTX 6000 pro costs. I also have it connected with fiber to my other rig for RPC when I need more memory. I haven't been posting much since I built this rig, because it's now more fun to talk to my machine. I run GLM5.3, DSv4.1Flash, Qwen3.8Flash, Qwen3.8-2.4T, KimiK3 and MiniMaxM3. Performance is great, a single RTX 6000 or M3 Mac Studio wish they could. Inference with vllm or llama.cpp

I look forward reading the replies how API usage is cheaper, or how it will take 52 light years to break even or the noise, or the electrical cost. NOT.

There will be more opportunities in the future, keep looking for them and pounce on them when they come. up, the demand is going to be high for compute for a long time.

https://preview.redd.it/dunixwu6caqh1.jpg?width=4080&format=pjpg&auto…

https://preview.redd.it/glpcbg5cbaqh1.jpg?width=3072&format=pjpg&auto…

💬 399 (+1) open on reddit ↗
▲
135
 
21👁
r/LocalLLaMA · u/Skyline34rGt · 23d ago
XingChen-AGI/Xing4.0-29B-A4B MoE

I find another new model at HF:

https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B

"Xing4.0-29B-A4B is a next-generation large language model in the Xing series (formerly TeleChat), developed by China Telecom Artificial Intelligence Technology Co., Ltd. With 29B total parameters and only 4B activated per token, it natively supports a 256K context length, extensible to 512K. It is the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework, and deeply optimized for complex engineering tasks.

For more information, please refer to our GitHub repository.

Highlights

  • Agent-Oriented Architecture: Built on the mHC + MLA + MTP architecture, supporting multi-step planning, tool calling, and complex reasoning chain execution, ensuring task coherence and execution stability under long contexts.
  • Deep Co-optimization with Ascend NPU: Adapted for Ascend 910C clusters using MindSpore/MindFormers, including feature adaptation for mHC and fused operator development, enabling stable and efficient training on the Ascend platform.
  • Significant Training Efficiency Gains: Through multi-level co-optimization — including fine-grained MoE communication optimization, selective recomputation, DVM automatic graph-operator fusion, and Ascend C mHC fused operators — overall training throughput was improved by approximately 96% over out-of-the-box performance.
  • Full Open-Source Ecosystem Compatibility: Supports LLaMA-Factory and MindFormers for fine-tuning; SGLang, vLLM, and KTransformers for inference and deployment; with targeted adaptation and format alignment for agent frameworks such as OpenCode, Claude Code, OpenClaw, and Hermes, enabling seamless integration into existing workflows.
  • Easy Adaptation for Domain-Specific Scenarios: The model is well-suited for downstream task fine-tuning, allowing lightweight customization on proprietary data for vertical domains such as intent classification, table understanding, contract auditing, and knowledge-based QA, enabling rapid domain capability development and deployment at low cost."

|Parameters|29B (4B active)|
|:-|:-|
|Number of Layers|40|
|Hidden Size|3584|
|Dense Intermediate Size|9216|
|Expert Intermediate Size|1024|
|Attention Type|MLA|
|Number of Routed Experts|64|
|Active Experts per Token|4|
|Number of Shared Experts|1|
|Context Length|256K (extensible to 512K)|

Benchmark

|Benchmark|Xing4.0-29B-A4B|Gemma4-26B-A4B|Qwen3.6-35B-A3B|
|:-|:-|:-|:-|
|IFBench|69.67|72.67|65.50|
|AIME2026|90.00|88.30|92.70|
|AA.LCR|61.00|66.00|62.00|
|Tau3-Bench|64.63|58.90|67.20|
|Claw-Eval|76.55|71.49|74.54|
|SWE-bench Verified|75.00|53.00|76.00|
|Terminal-Bench 2.1|57.50|30.00|51.50|
|SWE-bench Multilingual|66.00|51.00|67.20|
|DeepresearchBII|60.80|39.30|59.70|

💬 55 (+1) open on reddit ↗
▲
203
-1
24👁
r/LocalLLaMA · u/Odd_Caterpillar_2994 · 22d ago
Installing 6 GPUs in a standard case rather than using an open-frame chassis. post image

CPU : epyc 7262

MB : ROMED8-2T

GPU : V100 16GB PCIE x6

I have built a server with six V100 GPUs. I am now testing it and plan to eventually run Qwen3.8-Next-Flash configured with TP2 and PP3.

Because open-frame or server-style cases are large and unattractive, I chose to install all six cards in a full-tower case.

▲
123
-1
39👁
r/LocalLLaMA · u/Ashefromapex · 22d ago
First M5 Ultra benchmarks

just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link

For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!

💬 155 (+4) open on reddit ↗
▲
77
-1
27👁
r/LocalLLaMA · u/Gold-Bat-3225 · 22d ago
Humor Arena: Which LLM is the funniest? post image

We compared 20 model versions on 360 frozen joke prompts, with four jokes requested per prompt and model names hidden from our humor-trained judge.

Fable 5 had the highest estimated score: 66.8 points per 100 comparisons against the rival field. Fable 5.1 scored 58.2. Scores count a win as one point and a tie as half. The models near the top have overlapping uncertainty intervals. The reason we say estimated is that we scale the scores based on the scores of

The scores come from our own automated judge. A separate audit checked the judge with 1,400 ratings from 50 people; that was not a fresh human evaluation of Fable 5.1’s outputs. We specifically fine tuned an OS judge to rate the jokes and it correlates more highly with human preferences than any other model.

The full details here:
https://laugh.so/research/joke-generation/

Would love to hear your feedback!

▲
57
-1
23👁
r/LocalLLaMA · u/Character-Result-281 · 21d ago
Built this yesterday with Qwen3.8-Flash-Next (NVFP4, 262K context) on a single NVIDIA DGX Spark post image

Planning, coding, testing = 8h total.
Stack: VSCode Copilot in autopilot mode + SGLang
Stats: ∼10k lines generated, ∼800k tokens consumed

Sure, it's not GPT-6 Astra level, but for a 100% local ∼180B MoE running on a single DGX Spark at ∼35 tok/s. Not bad...

▲
986
-2
39👁
r/LocalLLaMA · u/Secure_Recording_472 · 22d ago
Thank you :) Swift Qwen 3.8 27B now has 100k+ downloads, is #1 finetune and #9 model on HuggingFace Trending post image

Hey everyone,

Jovan from UkisAI here, a small lab building the tech to make tiny frontier LLMs possible (and doing it open-source!)

The purpose of this post is simply to thank the community for all the amazing finetunes, quantizations and overall improvements over our original release which made our model get attention and the support for us to continue building in this direction! If it weren't for you guys going out of the way to contribute we wouldn't have half the results of this.

For context:

Swift Qwen 3.8 27B is our first open-source model release. It is proof of how penalizing pathological overthinking patterns inside of small LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy by not training them to think shorter directly but rather to think more efficiently.

We are continuing to build and are about to drop:

\- Swift1.5 Qwen3.8 27B (an improved checkpoint of the model with some training bugs fixed and more RL)

\- Swift Qwen3.8 Flash Next in the upcoming week week, we are now running the benchmark suite to not give out premature or incomplete results.

This time we ran even more benchmarks as you guys suggested, including more coding and long horizon!

It would be amazing if those of you who tried Swift would let us know what quants, features, changes you want to see in our upcoming model releases so we can do it better this time as we didn't even think about half of the stuff you guys were requesting last time :)

Let the era of non-slop finetunes begin!

EDIT:
Links -
https://huggingface.co/ukisai/Swift-Qwen3.8-27b
https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF
https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GGUF

💬 618 (+1) open on reddit ↗
▲
152
-2
25👁
r/LocalLLaMA · u/KURD_1_STAN · 21d ago
bonsai's document reveal how much cherry picked their headlines are

bonsai claim 98.2% intelligent retained, but their own documents show Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

that qwen3.5 is a typo cause these are qwen3.8 numbers, altho qwen3.5 numbers are

  • bonsai\_2/q3.5 52.8 / 41.6 = 126.9%
  • bonsai\_2/q3.5 60.8 / 72.4 = 84.0%
Long-context and coding performance. This release also delivers on the roadmap set out in our initial Bonsai 27B release \[2\], where we identified long-horizon, tool-driven software engineering as the next major capability to improve. With Ternary Bonsai 2 27B, that progress now shows up directly in agentic performance. Evaluated for the first time on Terminal-Bench 2.1 and SWE-bench Verified, the Ternary Bonsai 2 27B reaches 52.8 and 60.8, respectively, compared with 69.7 and 80.6 for Qwen3.5-27B, retaining roughly three quarters of the full-precision performance on both benchmarks.

link to their whitepaper on github, it is on page 7

^(also 3.8 35b qwhen? plsss)

▲
137
-2
21👁
r/LocalLLaMA · u/NineThreeTilNow · 22d ago
Update : Small model + Engram

I posted something about a 9b model a few days ago.

The real problem was 2 fold. Someone suggested the size was too big to prove it all out. It's a fair argument. I needed the model depth though. The second problem was the "ability" of these models and the fact that labs (with money) produce these models still. No real usefulness in my mind. A 2b model? maybe interesting. Want a chatbot? - this could be it.

Further, the Llama licensing at the core of the model was highly problematic and had to be thrown in the trash. It's vastly too restrictive and I couldn't Apache 2.0 anything.

So I moved to using the OLMo tokenizer. Except I shrunk the d\_model down to 2048 so I could build a tiny 2b model.

The size of the vocab with 1/2/3 gram forces the Engram table table to a full 1b. That's 50% where DS suggested like \~10-20%.

So basically 2b model + 1b Engram.

The model, because of the depth now allowed by 2048 d\_model allows me to push 40 total SWA/Global blocks to take advantage of attention based residual streams (Moonshot Kimi K3 style) in a dense model.

Data, is pulled from standard Wiki style data sources in my initial training. The data is pumped through the 7b OLMo model to generate the probabilities. This tells the model that the next token is a probability of 32 different tokens.

The big surprise was the results. Even after ONLY 15m tokens, the model is surprisingly coherent. Had I given pure 1 hot cross entropy, 15m tokens would do nothing. I'd get gibberish. However, I borrowed embeddings, and LM\_head that was down projected from 5k -> 2048 d\_model. This preserves \~65% of the data the "big" model had in the embedding when spectrum analysis is done.

The training is limited at 15m tokens, so it's like... babbling about Roman war history and stuff that exists in that token subset. It doesn't fall in to the classic early model phase where it repeats a token over and over though. That's what pretraining 15m generally gives you at first.

Anyways, this was an update and progress report for anyone who cares about this crap... I'm still processing all of the Wikipedia chunk from HuggingFace and working to train it. It'll take time.

If you read this far, thank you. If you have questions, I'm happy to reply. Someone suggested I was a kook who didn't understand ML previously. I started in ML some 20 years ago and held a brief (6 month) stint at Anthropic red teaming the original Opus 4 model before their "Constitutional" paper came out. I'm vaguely referenced in the paper as a "red teamer" I guess. I do this mostly because I love it.

I figured if I built an Apache 2.0 model, I'd at least want people to play with it. It appears 100% trainable on 24gb of VRAM. Training code / Model / Data will follow eventually. It's on HF / Github for now.

The old post is here :

https://old.reddit.com/r/LocalLLaMA/comments/1wezm58/is\_there\_still\_strong\_interest\_in\_a\_dense\_9b\_model/

edit;

Update can be found here :

https://www.reddit.com/r/LocalLLaMA/comments/1wnzu2f/engram_gone_wild_2b_mode…

▲
95
-2
13👁
▲
263
-3
22👁
▲
121
-3
20👁
r/LocalLLaMA · u/jacek2023 · 21d ago
inclusionAI/Realtime-Venus · Hugging Face

do you want some omni? here is omni for you

[](https://huggingface.co/inclusionAI/Realtime-Venus#1-🧭-overview)1. 🧭 Overview

This repository hosts two checkpoints of the Realtime-Venus system:

  • Realtime-Venus-Omni (Realtime-Venus-Omni/): the 9B audio-visual interaction model. It continuously watches and listens, decides whether and when to respond, and generates text and speech on a shared causal timeline. Adapted from MiniCPM-o 4.5, it supports proactive interaction, semantic interruption handling, and training-free long-video memory.
  • Realtime-Venus-Audio (Realtime-Venus-Audio/): the audio-focused checkpoint on the same streaming backbone, for audio understanding and audio-driven conversation with text or speech output.

Both directories contain model weights and custom Hugging Face Transformers code. The asynchronous Realtime-Venus-Harness and its external tool integrations live in the GitHub repository.

[](https://huggingface.co/inclusionAI/Realtime-Venus#2-✨-highlights)2. ✨ Highlights

  • Native full-duplex conversation: keeps perceiving while speaking and distinguishes backchannels, interruptions, corrections, and redirections.
  • Omni-Proactive interaction: continuously processes temporally aligned video and audio, and initiates a response when an event warrants it — without waiting for a user prompt.
  • Delegation: emits in-stream <delegate> requests on the shared causal timeline and consumes asynchronous backend results the same way, so external tasks never block the ongoing conversation. (Executing requests requires the Realtime-Venus-Harness runtime, available in the GitHub repository.)
  • Training-free long-video Memory: archives visually informative moments, retrieves query-relevant and non-redundant evidence, and reassembles the corresponding audio-visual context — no additional training required.
  • Text and speech output: generates response text together with native speech through the bundled Token2wav resources and a reference voice.
▲
268
-4
18👁
▲
152
-4
27👁
r/LocalLLaMA · u/Secure_Recording_472 · 21d ago
Question: UkisAI Swift Ternary Bonsai 2 27B?

Hey community,

Jovan from UkisAI here,

We're the team behind Swift Qwen3.8 27B, the Qwen model with token usage and overthinking error improvements

Our estimate is that we can make a great improvement to Bonsai 2, as our testing indicates that it suffers greatly from overthinking loops and in general high token usage impacting it's performance.

My ask for you is:

Is a Swifted version of Bonsai 2 something you guys would enjoy?

If yes, what size is the most relevant. 1-bit, 2-bit or both?

Thank you for the amazing feedback on Swift. We are glad you are enjoying it. Our download count jumped from 100k -> 150k overnight (community quants included).

For context, this is our model: https://www.reddit.com/r/LocalLLaMA/s/iCIbhxO8ue

https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF

💬 181 (+1) open on reddit ↗
▲
105
-4
22👁
r/LocalLLaMA · u/MeinDruckerSpinnt · 22d ago
JEV architecture

My understanding so far:

  1. You take an LLM and use it without thinking (That's what openjev does?)
  2. You leave out the text generation in the end and take the confidence score in the matrix before that phase

That's it. Right?

They gave it a mysterious marketing name.

▲
76
-4
30👁
r/LocalLLaMA · u/satnl · 23d ago
Qwen 3.8 27b is a amazing model, for the first time I see a local model found its own away to open a browser and test

I was testing this quantization IQ3\_XXS from GSQ-RCO with PI.

It is a heavy quantization case, the model is in IQ3\_XXS and KV cache in (Q4\_0, Q4\_0).

I asked it to make the flight simulator, using that popular prompt. For my surprise, when I went verify the session I saw some screenshots, but I don't have mcp or tools in the PI, it's an almost clean installation, only with some personal skills that dont have any mentions to browser testing.

Investigating the session I have noted that it found its way to open a browser for test by itself. My daily harness is opencode, and for not spending the 7k tokens with the chrome mcp in each session, I let the mcp disabled and I only enable it if the session needs. But this experience make think that I probably only need line in the agents md saying that he can use the headless browser with CDP when it needs.

https://preview.redd.it/wrgltip9rzph1.png?width=1952&format=png&auto=…

OK, let me write the file. Also test: I can't easily run a browser here. But I can at least do a syntax check of the JS via node (extract
the script? maybe parse with node --check).

----------------------

Now, final verification: run node --check again, and additionally do a simple DOM-less smoke test? Can't run three.js without a browser.
Syntax check + math tests are enough. Could also try opening it in a headless browser if available? Check whether the machine has
headless Chrome or playwright? That might be overkill; but if I could do an actual rendering check, I could catch runtime errors (e.g.,
shader compilation). Let me check whether headless Chrome is available.

$ cd /tmp && awk '/<script type="module">/{f=1;next} /<\/script>/{f=0} f'
/Users/[...]/zeroshot/Qwen3.8-27B-UD-IQ3_XXS_GSQ-RCO-pi-xhigh/index.html > sim.mjs && node --check sim.mjs
&& echo "SYNTAX OK"

SYNTAX OK

Took 0.1s

$ ls /Applications/ | grep -i -E 'chrome|chromium' ; which chromium chrome google-chrome 2>/dev/null; ls ~/Library/Caches/ms-playwright
2>/dev/null | head

... (5 earlier lines, ctrl+o to expand)
ffmpeg-1011
firefox-1532
firefox-1538
webkit-2311
webkit-2336

Took 0.0s

Chrome is available, and playwright's chromium headless shell is also available. Let's do a headless smoke test: load the page, wait a
few seconds, and capture console errors and screenshots. The page loads three.js from a CDN — the network needs to be working. Let's use
playwright's chromium (does npx playwright exist? The fact that the ms-playwright cache exists means playwright is installed somewhere).
Let me check if the playwright npm package is available globally, or use chrome directly with --headless.

The simplest approach: use Chrome headless directly:
/Applications/Google Chrome.app/Contents/MacOS/Google Chrome --headless --disable-gpu --screenshot=/tmp/shot.png --window-size=1600,900
--virtual-time-budget=8000 file://...

However, WebGL in headless Chrome: it works with SwiftShader (software GL) via --use-gl=swiftshader or --enable-unsafe-swiftshader. In
newer versions of Chrome, the software fallback requires --enable-unsafe-swiftshader. Let's try it. Also, console errors:
--enable-logging=stderr --v=1 will give us the console.

Let me write a small node script using playwright if available. First, check the availability of node playwright.

Chrome is available — let me do a headless smoke test to catch runtime/shader errors and grab a screenshot:

$ node -e "try{require.resolve('playwright');console.log('playwright OK')}catch(e){console.log('no playwright')}"; ls
~/Library/Caches/ms-playwright/chromium-1234/ 2>/dev/null | head -3

no playwright
chrome-mac-arm64
DEPENDENCIES_VALIDATED
INSTALLATION_COMPLETE

Took 0.1s

There's no playwright module, but the Chromium headless shell exists. Either find the path to the headless shell binary and run it
directly, or use Google Chrome in headless mode with SwiftShader.

The path for Playwright's Chromium headless shell:
~/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-mac-arm64/chrome-headless-shell? Let me check the structure.


$ find ~/Library/Caches/ms-playwright/chromium_headless_shell-1234 -name '*headless*' -o -name 'chrome*' | head; find
~/Library/Caches/ms-playwright/chromium-1234 -name 'Chromium' -o -name 'chrome' | head

... (3 earlier lines, ctrl+o to expand)
s.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/chrome-headless-shell
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/headless_lib_strings.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/headless_lib_data.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/LICENSE.headless_shell

Took 0.1s

Use chrome-headless-shell — but to catch console errors and drive key input, we need CDP (Chrome DevTools Protocol).
chrome-headless-shell supports --remote-debugging-port. I could also write a small Node script that speaks raw CDP over WebSocket... but
that would need a ws client. Node doesn't have a built-in WebSocket client until Node 22+ (Node 22+ has a global WebSocket). What version
of Node is it? Let me check. Node 24 (from the nvm path v24.21.0) — a global WebSocket exists.

Plan:
1. Launch chrome-headless-shell with --remote-debugging-port=9222, --no-sandbox, --use-gl=swiftshader? For headless shell, WebGL via
SwiftShader: --enable-unsafe-swiftshader flag; also --no-sandbox is needed? For headless shell, maybe --no-sandbox isn't necessary,
but I'll add it just in case. Also, a new profile directory with --user-data-dir=/tmp/cdp-profile.
2. Node script: fetch http://127.0.0.1:9222/json/new?url=file:///... (PUT) to create a tab, connect via WebSocket to
webSocketDebuggerUrl, and do the following:
- Enable Runtime and Log
- Collect console messages and exceptions
- Wait about 6 seconds (let the first frame render)
- Dispatch key events? Simulating keydown via CDP: Input.dispatchKeyEvent. Possible, but keyboard isn't required for the smoke test.
- Page.captureScreenshot → save to PNG
- Also, evaluate expressions: get the chunk count, and confirm the renderer is working? For example
document.querySelectorAll('canvas').length, window.__stats... I haven't exposed the stats. I can evaluate JS in the page context:
the module's scope isn't global, but I can check document.getElementById('alt').textContent (the HUD should show altitude after
0.12 seconds), and check the canvas pixel colors via toDataURL? Easier: check the HUD text + screenshot.
3. Report console errors.

Let me write this script.

I'll drive the headless shell over CDP with Node's built-in WebSocket to capture console errors and a screenshot:

------------

> why did you used a browser? did you already had any instruction about it?

User's question: "Why did you use a browser? Was there any prior instruction to do so?"

Honest answer: there was no prior instruction to use a browser. What I chose on my own was:
1. To do syntax checking with node --check (low cost, so I did that first).
2. Since a single HTML file with WebGL + CDN can't be fully verified without actually rendering it, I used a locally available headless
Chrome (I found Google Chrome and Playwright's chromium headless shell on the machine). It's to catch runtime errors (shader
compilation, WebGL failures, logic errors like the inverted bank direction), and to take screenshots as visual evidence.

There's no instruction in AGENTS.md or in the skills that requires using a browser. There's nothing in the skill list that mandates it
either (the qa-manual-guide skill is about writing a guide, and I didn't use it since it wasn't explicitly requested… actually, this
skill says "when an implemented feature needs an evidence-based local manual testing guide…" — what the user asked for was a simulator,
not a manual testing guide. My verification was a homegrown approach).

Why I used it: a single-file browser game can't be verified with unit tests alone. Headless rendering would let me detect actual failures
(e.g., I actually found an inverted A/D bank bug that a syntax check could never find), and screenshots let me verify the visual quality
(mountain appearance, sun position, cloud scale) — the "beautiful and relaxing" requirement is inherently visual.

I should answer concisely and honestly: no, there was no instruction. I discovered the browser on the machine and used it for
evidence-based verification. The main findings were the inverted turn bug and visual confirmation. Also note that it was only used for
testing and didn't modify any files inside the project.

Keep the reply short and structured.

No — there was no instruction to use a browser. That was my own choice for verification, not from AGENTS.md or any skill.

Why I did it:
- A single-HTML-file WebGL game can't be properly validated by syntax checks alone — the interesting failures (shader compile errors,
WebGL context issues, sign errors in the flight math) only surface when it actually runs.
- I noticed Google Chrome and Playwright headless Chromium were available locally, so I drove a headless shell over CDP: load the page,
capture console/exception logs, simulate key presses, read the HUD values, and take screenshots.

▲
62
-4
28👁
r/LocalLLaMA · u/drooolingidiot · 21d ago
We benchmarked 24 LLMs against human writers on 475 creative writing prompts post image

We just released the first version of our Creative Writing benchmark, comparing 24 LLMs against human writers across 475 writing prompts.

Creative writing is subjective, so the rankings aren't meant to predict what any one person will prefer. Instead, they predict what a large group of readers would prefer, using a custom reward model trained specifically on human preferences for creative writing.

Surprisingly, the strongest frontier models already rank above the talented amateur writer cohort, while professional writers still lead by a wide margin.

You can browse the full benchmark, compare the model outputs side by side, and see how the benchmark works here:

https://vulsar.ai/benchmarks/creative-writing-v1/

Curious what you all think of the results!

💬 77 (+1) open on reddit ↗