37 posts · 1 sub · RSS
← prev Sep 16, 2026 → Sep 17, 2026 next →
2026-09-16 → 2026-09-17 hourdayweekmonthyearall
allr/LocalLLaMA
▲
3592
+107
76👁
r/LocalLLaMA · u/Nandakishor_ml · 24d ago
I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Update: I made a generic model and beaten the jev in all of the benchmarks. Code and details available at https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. For anyones information the main guiding model is RL not embedding model or LLM Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43 Paper: https://arxiv.org/abs/2503.23303 Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations Also the second work published in September 2025 was exactly the same one jev proposed now Paper: https://arxiv.org/abs/2510.01237 My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0). Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices. It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂

💬 339 (+9) open on reddit ↗
▲
121
-3
41👁
r/LocalLLaMA · u/Ashefromapex · 23d ago
First M5 Ultra benchmarks

just saw some benchmarks on the omlx website for the m5 ultra (don’t know how official they are but they seem reasonable): Link

For Qwen 3.8 27B q4 it gets 50 tok/s th and 1800 tok/s pp at8k context and without mtp. Seems very promising!

💬 155 (+4) open on reddit ↗
▲
1628
+20
49👁
r/LocalLLaMA · u/Nandakishor_ml · 24d ago
I literally built the Jev architecture one year back and completely open-sourced it with model, dataset and paper

Update:
I made a generic version. Full details at https://www.reddit.com/r/LocalLLaMA/s/bbwyiOprUs
It includes code, benchmark and hf repo

Everyone now talks about the architecture that's not auto regressive and does lightning fast probability prediction with a json schema. I worked on this literally one year back in March 2025, published an arxiv paper, pushed the model to huggingface along with the pypi package and training dataset. And then one year later, a

frontier lab came, proposing the same idea like literal breakthrough without technical papers, open weights and no open dataset. I posted my approach in this subreddit. Links are. For anyones information the main guiding model is RL not embedding model or LLM

Reddit post: https://www.reddit.com/r/LocalLLaMA/s/6eGEwsAz43

Paper: https://arxiv.org/abs/2503.23303

Model: https://huggingface.co/DeepMostInnovations/sales-conversion-model-reinf-learning

Dataset: https://huggingface.co/datasets/DeepMostInnovations/saas-sales-conversations

Also the second work published in September 2025 was exactly the same one jev proposed now

Paper: https://arxiv.org/abs/2510.01237

My model uses PPO over sequence embeddings to output turn-by-turn conversion trajectories (probabilities from 0.0 to 1.0).

Jev uses parallel sampling (trained via RLCD) to output confidence distributions and schema choices.

It's incredibly frustrating that the thing that you made with months of hard work, sweat and sleepless night is architecturally similar with the vertical use case and don't get the support you deserve because frontier lab build something horizontal. The open-source story in general 🙂

💬 136 (+2) open on reddit ↗
▲
519
+5
38👁
r/LocalLLaMA · u/GuiltyBookkeeper4849 · 24d ago
Qwen 3.8 27B Running for 63 hours on a RTX 3090 to solve the Riemann hypothesis

I let Qwen 3.8 27B 4bit quantized with 100K context window run autonomously for 63 hours (50 million+ tokens) to try to solve the RH.

Of course it did not solve it, but the experiment still shows it's internal work, memory organization, strategies used and more.

The interesting thing is that it never hallucinated an answer and never stopped trying new ideas to solve it.

Multiple times it corrected it's own mistakes.

I am really hopeful that one of the unsolved millenium prize problems will be solved by an agent or a swarm of agents powered by an open source model in the next 12 months.

If you want to check out it's internal memories, code, strategies and more I published everything on HF: https://huggingface.co/datasets/gr0010/artificium-riemannhypothesis-experiment

My next goal is to actually use an agent perhaps powered by a smarter open model like GLM 5.3 flash or a swarm of agents, to solve an open math problem.

Please let me know if you tried something similar, what problem you'd suggest to tackle next, and if you have any question.

If you have GPUs consider getting in touch with me, we could run multiple agents to create a swarm and get them to tackle a simple yet open math/coding problem.

💬 193 (+2) open on reddit ↗
▲
206
+5
30👁
r/LocalLLaMA · u/whodoneit1 · 23d ago
153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s post image

People kept commenting and asking about single AMD 1xR9700 cards in the comments and discord. Well, I finally had time to do some optimizations for 1xR9700 owners and performance has doubled across the board. You can see the results in BetterBench above if you like visuals or below if you're more into text.

These results were measured running Unsloth's Qwen3.8 27b NVFP4.

Decode
┌───────────────┬───────────────┬──────────────────┐
│ category │ update p50 ms │ decode t/s (med) │
├───────────────┼───────────────┼──────────────────┤
│ chat │ 42.3 │ 67.1 │
├───────────────┼───────────────┼──────────────────┤
│ code │ 42.5 │ 120.5 │
├───────────────┼───────────────┼──────────────────┤
│ file_edit │ 42.5 │ 138.0 │
├───────────────┼───────────────┼──────────────────┤
│ json │ 42.4 │ 153.1 │
├───────────────┼───────────────┼──────────────────┤
│ math │ 42.5 │ 140.0 │
├───────────────┼───────────────┼──────────────────┤
│ prose │ 42.3 │ 69.2 │
├───────────────┼───────────────┼──────────────────┤
│ reasoning │ 34.3 │ 123.9 │
├───────────────┼───────────────┼──────────────────┤
│ summarization │ 34.2 │ 141.7 │
└───────────────┴───────────────┴──────────────────┘

Prefill
┌───────────────┬───────────────┐
│ prefill depth │ pp tok/s │
├───────────────┼───────────────|
│ 2000 │ 3552 │
├───────────────┼───────────────|
│ 8000 │ 3536 │
├───────────────┼───────────────|
│ 16000 │ 3619 │
├───────────────┼───────────────|
│ 32000 │ 3437 │
├───────────────┼───────────────|
│ 64000 │ 3192 │
├───────────────┼───────────────|

Concurrency
┌───────────────┬───────────────┐
│ level │ tok/s │
├───────────────┼───────────────|
│ 1 │ 120 │
├───────────────┼───────────────|
│ 2 │ 215 │
├───────────────┼───────────────|
│ 4 │ 322 │
├───────────────┼───────────────|
│ 8 │ 471 │
├───────────────┼───────────────|

Links (Both repo's updated as some users wanted Github)

https://codeberg.org/ggz14/radiance-vllm-mxfp4

https://github.com/GGZ14/vllm-mxfp4

https://x.com/bkuyper

I hope you single R9700 card owners enjoy this release!

💬 120 (+2) open on reddit ↗
▲
991
+3
43👁
r/LocalLLaMA · u/Secure_Recording_472 · 23d ago
Thank you :) Swift Qwen 3.8 27B now has 100k+ downloads, is #1 finetune and #9 model on HuggingFace Trending post image

Hey everyone,

Jovan from UkisAI here, a small lab building the tech to make tiny frontier LLMs possible (and doing it open-source!)

The purpose of this post is simply to thank the community for all the amazing finetunes, quantizations and overall improvements over our original release which made our model get attention and the support for us to continue building in this direction! If it weren't for you guys going out of the way to contribute we wouldn't have half the results of this.

For context:

Swift Qwen 3.8 27B is our first open-source model release. It is proof of how penalizing pathological overthinking patterns inside of small LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy by not training them to think shorter directly but rather to think more efficiently.

We are continuing to build and are about to drop:

\- Swift1.5 Qwen3.8 27B (an improved checkpoint of the model with some training bugs fixed and more RL)

\- Swift Qwen3.8 Flash Next in the upcoming week week, we are now running the benchmark suite to not give out premature or incomplete results.

This time we ran even more benchmarks as you guys suggested, including more coding and long horizon!

It would be amazing if those of you who tried Swift would let us know what quants, features, changes you want to see in our upcoming model releases so we can do it better this time as we didn't even think about half of the stuff you guys were requesting last time :)

Let the era of non-slop finetunes begin!

EDIT:
Links -
https://huggingface.co/ukisai/Swift-Qwen3.8-27b
https://huggingface.co/ukisai/Swift-Qwen3.8-27B-GGUF
https://huggingface.co/bartowski/ukisai_Swift-Qwen3.8-27b-GGUF

💬 618 (+1) open on reddit ↗
▲
139
+4
24👁
r/LocalLLaMA · u/Skyline34rGt · 23d ago
XingChen-AGI/Xing4.0-29B-A4B MoE

I find another new model at HF:

https://huggingface.co/XingChen-AGI/Xing4.0-29B-A4B

"Xing4.0-29B-A4B is a next-generation large language model in the Xing series (formerly TeleChat), developed by China Telecom Artificial Intelligence Technology Co., Ltd. With 29B total parameters and only 4B activated per token, it natively supports a 256K context length, extensible to 512K. It is the first model of this scale trained entirely on the Ascend NPU platform with the MindSpore framework, and deeply optimized for complex engineering tasks.

For more information, please refer to our GitHub repository.

Highlights

  • Agent-Oriented Architecture: Built on the mHC + MLA + MTP architecture, supporting multi-step planning, tool calling, and complex reasoning chain execution, ensuring task coherence and execution stability under long contexts.
  • Deep Co-optimization with Ascend NPU: Adapted for Ascend 910C clusters using MindSpore/MindFormers, including feature adaptation for mHC and fused operator development, enabling stable and efficient training on the Ascend platform.
  • Significant Training Efficiency Gains: Through multi-level co-optimization — including fine-grained MoE communication optimization, selective recomputation, DVM automatic graph-operator fusion, and Ascend C mHC fused operators — overall training throughput was improved by approximately 96% over out-of-the-box performance.
  • Full Open-Source Ecosystem Compatibility: Supports LLaMA-Factory and MindFormers for fine-tuning; SGLang, vLLM, and KTransformers for inference and deployment; with targeted adaptation and format alignment for agent frameworks such as OpenCode, Claude Code, OpenClaw, and Hermes, enabling seamless integration into existing workflows.
  • Easy Adaptation for Domain-Specific Scenarios: The model is well-suited for downstream task fine-tuning, allowing lightweight customization on proprietary data for vertical domains such as intent classification, table understanding, contract auditing, and knowledge-based QA, enabling rapid domain capability development and deployment at low cost."

|Parameters|29B (4B active)|
|:-|:-|
|Number of Layers|40|
|Hidden Size|3584|
|Dense Intermediate Size|9216|
|Expert Intermediate Size|1024|
|Attention Type|MLA|
|Number of Routed Experts|64|
|Active Experts per Token|4|
|Number of Shared Experts|1|
|Context Length|256K (extensible to 512K)|

Benchmark

|Benchmark|Xing4.0-29B-A4B|Gemma4-26B-A4B|Qwen3.6-35B-A3B|
|:-|:-|:-|:-|
|IFBench|69.67|72.67|65.50|
|AIME2026|90.00|88.30|92.70|
|AA.LCR|61.00|66.00|62.00|
|Tau3-Bench|64.63|58.90|67.20|
|Claw-Eval|76.55|71.49|74.54|
|SWE-bench Verified|75.00|53.00|76.00|
|Terminal-Bench 2.1|57.50|30.00|51.50|
|SWE-bench Multilingual|66.00|51.00|67.20|
|DeepresearchBII|60.80|39.30|59.70|

💬 55 (+1) open on reddit ↗
▲
95
+1
15👁
r/LocalLLaMA · u/MrMrsPotts · 25d ago
What's the next local model you are excited about?

And why?

💬 215 (+1) open on reddit ↗
▲
73
-2
35👁
r/LocalLLaMA · u/BullfrogScary8947 · 24d ago
[Release] SOTA GGUFs for Qwen3.8-Flash-Next: GSQ-RCO Providing Near Baseline Performance

https://preview.redd.it/e5wn8eyh7vph1.png?width=1080&format=png&auto=…

https://preview.redd.it/8ov5gl8j7vph1.png?width=1080&format=png&auto=…

New Qwen3.8-Flash-Next quantization using GSQ-RCO. Cuts the size of Qwen3.8 Flash Next from around 80-95GB to 68-76GB, while still preserving near baseline quality. Also their Q2\_0 variant claims to be much faster offering 6.2x better prompt throughput in coding.

"Q2\_0 is built for speed. It avoids the quantization formats that rely on large lookup tables: those formats pack more accuracy into a given bit-width, but decoding them costs real time, and on this model that cost dominates inference. Q2\_0 delivers 3.4x the prompt throughput and 1.9x lower end-to-end latency than IQ2\_XS at a slightly smaller file size, and its decode rate stays flat across workloads instead of varying with the content. The trade is a little quality: 89.07 task average against 89.16 for IQ2\_XS, and 3.5 points below IQ3\_XXS. Pick it when throughput matters most, and see *Performance* for the measurements.

The IQ3\_XXS model is the strongest operating point: it matches the base model exactly on AIME25 (100.00) and is within 0.51 points on GPQA-Diamond and 1.14 on LiveCodeBench v6, at roughly one fifth of the BF16 size."

Model link: https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF

💬 54 (+1) open on reddit ↗
▲
62
-1
29👁
▲
1682
+15
36👁
r/LocalLLaMA · u/xenovatech · 23d ago
Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU. post image

The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to <6GB in size. According to the model card, it's 9x smaller than FP16 while retaining 98.2% of the intelligence.
\- Collection: https://huggingface.co/collections/prism-ml/bonsai-2
\- Demo: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels

▲
1273
+6
21👁
▲
536
-4
37👁
r/LocalLLaMA · u/RishiFurfox · 24d ago
Hey, Meta. Where's those Muse Spark weights? post image

It was well over a month since Meta promised to release the weights for Muse Spark.

Back then (10th August), they were on Spark 1.2. Now we're on 1.3 and still nothing's been released. So it begs the question: will they be releasing the 1.2 weights when 1.4 drops? Or will we get whatever's then-current as open weights?

It's ironic given Mark Zuckerberg said at the same time that we can't delay the release of models by "even a month," due to the competition with China. It's been well over a month. He was arguing in the context of new regulations delaying models, but I think it applies equally to the open weights contest as it does to the closed models one.

After all, the Chinese models are all open. That's the competition and point of comparison.

Have Meta given any sort of explanation for why they're sitting on the weights or how much longer it'll take for them to honour their promise? Will we even get them in light of all the attempts at regulatory capture and dire warnings about how AI is dangerous?

▲
463
+6
24👁
r/LocalLLaMA · u/skeole · 24d ago
Xiaomi MiMo 2.6 Live Training Dashboard

Cool to see this as it happens!

▲
422
-5
25👁
r/LocalLLaMA · u/Porespellar · 24d ago
Frontier LLM development simplified for politicians: post image

Nobody is buying this “Pace the frontier” nonsense. It makes no logical sense at all. Are American labs really going to take a pause and lose any small lead they still may have over Chinese labs? Does anyone really believe this? This seems like some performative virtue signaling BS. Why are they bothering with this pacing campaign? Someone please explain.

▲
395
+6
35👁
r/LocalLLaMA · u/FullstackSensei · 23d ago
AMD Plans 10% Price Hike Across GPUs, Chipsets, and Possibly CPUs

Great news! AMD is also considering accepting payment in organs!

Slightly less sarcastically, grab what you can, while you can. Waiting is becoming very costly almost by the day

▲
229
+4
23👁
r/LocalLLaMA · u/Henrie_the_dreamer · 23d ago
Cactus Needle 3: A Sliceable 8-29MB Automation Foundation Model That Matches DeepSeek v4 Flash post image

Hey all, Henry from Cactus Compute here, I kinda wanted to share our latest model and get feedback from the family :)

Needle 3 is a small foundation model for automation: you give it the functions your app exposes, it reads a request and returns the calls with every argument filled in, or a typed record if what you gave it was a schema. It runs on the device, with no network in the loop. It is on Hugging Face, on GitHub, on PyPI as cactus-needle, and there is a sandbox that runs it in your browser at cactuscompute.com/needle if you want to poke at it before reading further.

1) Trades general capacity for frontier performance on automation tasks

The thing we decided early was that Needle would not chat. Every turn is a function call, and a request no declared tool can serve comes back as an empty list rather than a guess. That sounds like a limitation, and it is, but it is what let a 121M-parameter model be trained on 360B tokens of structured data and spend all of its capacity on three jobs: tool calls, structured extraction and text embedding.

The architecture follows from the same trade. It is a Simple Attention Network: the dense feed-forward layers are gone, replaced by a Monarch Hadamard MLP with 25.6K parameters per layer instead of 4.7M, and the knowledge a feed-forward layer would normally hold sits in an engram, hashed n-gram tables that are read by gather and cost no arithmetic. 70.8M of the 121M parameters live there, so the full model does the arithmetic of a 50M one: 100 MFLOPs per token against 296 for a transformer of the same shape.

https://i.redd.it/fsnfjojny4qh1.gif

We wrote the intuition up if you want the longer version: Simple Attention Networks and the Hadamard MLP.

2) Beats models 10x its size on tool calls and language-to-device control

On Mobile Actions (961 phone commands, scored on the exact call), the 20-layer model scores 86.0 through the shipped 2-bit binary with the confidence gate on. LFM2.5 1.2B is at 82.4, Qwen3.5 0.8B at 76.0, FunctionGemma 270M at 65.1 and Apple's on-device foundation model at 57.6, all at f16. DeepSeek V4 Flash through its API is at 88.4, which is the line in the chart.

https://i.redd.it/luunq717z4qh1.gif

The part we are most pleased with is not the number but how the calls are made. Every argument is a span of the request: the model writes a short derivation first ('living room' -> room; '30' -> brightness) and then emits the call under a byte-level grammar compiled from your schema, so the JSON always parses and an enum can never leave its set. An optional field with no evidence is omitted, a required one with no evidence withholds the call, and the engine drops a call the request negates or excludes. Ask for two things and you get two calls in order.

https://i.redd.it/227396y5z4qh1.gif

Full table across all six suites (tool calling is exact match, extraction is field F1, Needle through the shipped binary, baselines at f16 under vLLM):

|Model|Params|Mobile Actions|DroidCall|BFCL v4|DSTC8 F1|SNIPS gold F1|SNIPS 7-way F1|
|:-|:-|:-|:-|:-|:-|:-|:-|
|DeepSeek V4 Flash (cloud)|\-|88.4|60.5|77.2|80.0|69.4|66.7|
|Needle3-20L-121M|121M|86.0|47.0|50.2|40.7|30.2|24.7|
|LFM2.5 1.2B|1.2B|82.4|35.5|62.0|48.0|43.0|38.0|
|Needle3-16L-98M|98M|80.7|40.0|41.3|28.5|23.5|19.2|
|Qwen3.5 0.8B|800M|76.0|28.0|56.8|49.0|35.0|34.0|
|LFM2.5 350M|350M|72.8|32.5|59.1|20.0|34.0|29.0|
|LFM2.5 230M|230M|69.3|11.5|46.3|53.0|27.0|22.0|
|FunctionGemma 270M|270M|65.1|16.5|46.6|27.0|29.0|14.0|
|Needle 2|45M|63.5|17.0|\-|\-|\-|\-|
|Apple FM|3.0B|57.6|\-|\-|\-|\-|\-|
|Needle3-8L-52M|52M|36.8|36.5|28.2|15.3|16.6|10.1|
|Needle3-4L-29M|29M|11.7|21.0|19.5|6.9|7.7|4.3|

You can see where it is weaker too: BFCL and the extraction suites are where the bigger baselines pull ahead, and the smaller subnetworks fall off quickly on the general task (more on why that is fine in section 4).

3) Matches 2-3x bigger models on structured JSON extraction

Extraction is not a separate mode. You declare the record as the only tool and pass the passage where the query goes; with one tool declared the grammar admits exactly one call of that name, so the shape is guaranteed rather than requested, and the values are grounded the same way as arguments: a field is filled only from a span of the passage, an optional field with no span comes back as None, and a date whose year appears nowhere in the text is flagged instead of invented.

from pydantic import BaseModel
import needle

class Invoice(BaseModel):
vendor: str
total: float
due_date: str
po_number: str | None = None

needle.extract("Invoice from Acme Corp, $1,200.00, due 2026-09-01", Invoice)
# Invoice(vendor='Acme Corp', total=1200.0, due_date='2026-09-01', po_number=None)

It generalised to classification without special training, because an enum is just a constrained value: declare sentiment: Literal["positive", "neutral", "negative"] on a record and you have a classifier whose output cannot leave the set. A watch reads a notification into merchant, amount and date that way, then into a reply, then into a sentiment flag, one record each. On DSTC8 and the two SNIPS suites the 121M model lands between the 230M and 350M baselines, which is the 2-3x in the heading.

4) Intelligence ladder: every depth from 2 to 20 layers a model of its own

This is the part I would most like your thoughts on. Needle 3 is one set of weights, and every depth from 2 to 20 layers is a deployable model. Blocks 0 and 19 are always kept and the rest are added by bisection, so each subnetwork nests in the next; during training each step samples one path, mostly the full model and otherwise a random depth, with the smaller path distilled from the full one. The full-depth model ends up slightly better than an ordinary run of the same size, and every depth below it is trained rather than truncated.

https://i.redd.it/kbwimep4z4qh1.gif

Why we wanted it: a watch, a Raspberry Pi and a phone do not want the same model, and they want to pick the size at deploy time. needle build --layers 8 writes the 8-layer file; the same engine runs all of them. The small depths lose accuracy on the general benchmarks (that is the bottom of the table above), and they get it back when fine-tuned to one product's tools: on DroidCall every subnetwork gains 18 to 36 points, and from 4 layers (29M parameters) up the tuned subnetwork passes DeepSeek V4 Flash.

https://i.redd.it/y646wed3z4qh1.gif

Fine-tuning is LoRA on the frozen base, merged at export, and the Python package does it locally at 4 bits (needle finetune data.jsonl, then needle build). The maths is in Intelligence Ladders and the workflow in Fine-tuning Needle.

5) Runs locally at up to 4k tokens/sec decode speed

The engine is under 1 MB, plain CPU, no GPU or NPU, and the weights are read in place from a single file the engine maps into memory: a 196-byte header carrying the whole architecture geometry, a nameless tensor directory, and the quantised blobs in the order the forward pass reads them. On a Raspberry Pi 5, decode runs at up to 4k tokens/s at the bottom of the ladder and around 400 at the top, prefill from 10k down to 1k. Every response reports prefill_tps, decode_tps and peak_ram_mb, so you can measure on your own device rather than take our word for it.

Every response also carries a confidence score from a calibrated head, the minimum of a post-hoc judgement on the finished call and the decode probability of its tokens. The engine withholds anything under 0.1; above that the number is yours: act at once when it is high, show the call and ask when it is middling, treat [] as a refusal. How we use it is in Leveraging Needle's confidence.

6) 25-121M deployable parameters at CQ2-bit (8-29MB binaries)

The weights are quantised with Cactus Quants: groups of 128 weights are rotated by a Walsh-Hadamard matrix, which makes every group look Gaussian, split into an fp16 norm and a direction on the unit sphere, and the direction's coordinates are snapped to a 4-entry Lloyd-Max codebook. That is 2.125 bits per weight, and the kernel never expands them: it rotates and int8-quantises the activation instead, then does table lookups and sdot against the packed indices. The embedding and the confidence head keep 4 bits, the norms and gates stay fp16. The byte layout, and a twenty-line parser for it, are in The .cact format.

7) For mobiles, wearables, smart home, small robots and microcontrollers

Every target ships a prebuilt engine folder: macOS, Linux on x86-64, ARM64, ARMv7, RISC-V and MIPS32 (the Ingenic camera SoCs), Windows x64 and ARM, Android, iOS, watchOS, tvOS, the browser as WebAssembly, and a WASI component with a WIT world. pip install cactus-needle covers the desktop and server platforms with wheel-tagged engines, and needle build --platform linux-arm64 --layers 8 --out ./pi puts an engine and the weights in a folder you copy over. Inference never touches the network, so an air-gapped device only needs the files in place. The list and the runtime surfaces are in What devices are supported.

import needle

@needle.tool
def get_weather(city: str):
"Get the current weather for a city."
return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]

What we would love feedback on

  • Where the grounding rules get in your way. The engine refuses to invent a number, drops a call the request excludes and withholds a required enum the request never names; we tuned those on our suites and would like to hear where they bite on real tools.
  • The ladder. Whether depth is the right knob for you, or whether you would rather have width, and what devices you would want the 2- and 4-layer models on.
  • Extraction cases we have not seen. Nested records, arrays, multilingual text (it is English-first, and non-English text fragments into about 1.7x more tokens).
  • Anything in the tool design guide that turned out wrong for your schemas.

Thanks for reading this far. Happy to answer anything in the comments.

▲
208
-3
21👁
▲
203
-1
24👁
r/LocalLLaMA · u/Odd_Caterpillar_2994 · 23d ago
Installing 6 GPUs in a standard case rather than using an open-frame chassis. post image

CPU : epyc 7262

MB : ROMED8-2T

GPU : V100 16GB PCIE x6

I have built a server with six V100 GPUs. I am now testing it and plan to eventually run Qwen3.8-Next-Flash configured with TP2 and PP3.

Because open-frame or server-style cases are large and unattractive, I chose to install all six cards in a full-tower case.

▲
168
+2
22👁
r/LocalLLaMA · u/Mysterious_Hearing14 · 24d ago
Openjev post image

https://huggingface.co/AlexWortega/openjev

I build an openjev, it can play games and do everything what jev can. and yes - it's trained as crossencoder

▲
164
+3
23👁
r/LocalLLaMA · u/Jorlen · 23d ago
Does anyone use uncensored models purely for coding?

It sounds like a stupid question, and I do apologize if it is... but I've seen several people mention that coding models are better uncensored due to the fact that they don't have to constantly run prompts through the "is this okay" sort of checks.

Is this hogwash? Is it true? And more importantly, does anyone have any sources to confirm it?

Anecdotal evidence is fine too if you've tried and compared them.

Personally, I have never bothered because I'm too worried the de-censoring would damage the weights. The juice never felt like it was worth the squeeze... but maybe I was wrong?

Edit: Either I'm unclear or people are misinterpreting my request: Specifically, I mean for every day coding (not for hacking, not for reverse-engineering) but just for regular coding of new apps, etc. The question is: will the uncensored model produce better code faster (without reasoning so much) because it no longer has to worry about "is this alright" when it questions everything...

Edit 2: Decided to test the HuiHui Qwen 3.8 27b (UD-Q8\_K\_XL) quant myself. So far, it reasons far less, and I have yet to have any issues with its coding quality. Granted, I've only been testing it for about 8 hours (straight...) in an active project. Its reasoning is far shorter, it seems far more confident in its responses and as such, uses far less context to achieve the same result. I will continue testing for another week; it's pitted against the Dirk template version of Qwen 3.8 27b right now (same quant) which I'd been using the previous week.

▲
146
-3
24👁
r/LocalLLaMA · u/sadnessdevil · 24d ago
You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

I'm pretty sure it can be done with any model based on qwen4exp, which Qwen's next local models will be based on. You can use a quant that barely fits in VRAM and still run at the model's maximum context length without kv cache quantization, since most of the KV cache can live in system RAM.

I actually made it working on vLLM and now I get 1M context with 3x 3090. I get \~80 tok/s at short context, dropping to \~60 tok/s once QSA reaches its 2048-token budget, after which decode speed stays flat as total context grows. The throughput is pretty good too, and I get like 150tk/s @ 4 concurrent requests. Prefill at 248k reaches 3,701 tok/s. (The patches and the model are available on my huggingface page if you're interested)

Decode speed is a bandwidth problem. Each decode step produces one token, and to produce it the GPU reads every weight and every piece of attention state that the step needs. On a single stream the card spends most of the step waiting for memory rather than computing. So the size of that per-step read sets the token rate.

This is why a normal model keeps its KV cache in VRAM. Take Qwen3.8-27B, which is built on the Qwen3-Next architecture and shares most of its properties with Qwen3.8-Flash-Next (\qwen4\_exp\). It still has one full attention layer every few layers, and a full attention layer reads its entire KV cache on every step. That read grows with the context, so decode gets slower as the conversation gets longer. It also grows past what any host link (such as PCIe) can carry, so the cache has to sit next to the compute.

The numbers of this model show the size of the problem. One QSA layer holds 2 key/value heads of 256 dimensions, as K and as V, in 2 bytes each, which is 2,048 B per token. At 262,144 tokens that is 512 MiB for one layer, and 6 GiB for all 12 layers on every single step. A PCIe 4.0 x16 slot carries about 32 GiB/s, so a host-resident cache of that shape allows about 5 tokens per second.

Here's an interesting part, Qwen3.8-Flash-Next avoids this in two ways:

Only 12 of the 48 layers have a KV cache at all. The other 36 layers are gated delta-net layers, a linear attention whose recurrent state has a fixed size. That state does not grow with the context.

Those 12 layers also do not attend over the whole context. QSA runs a cheap indexer over a pooled, compressed key, where \indexer\_head\_dim=128\ divided by \indexer\_compress\_ratio=4\ gives the pooled width. The indexer selects at most \indexer\_budget=2048\ positions. The layer reads the main KV rows only for the positions that the indexer selects.

So \indexer\_budget\ bounds the bytes that a decode step reads, and the context length does not:

\\\`

2048 selected x 2 kv heads x 256 dim x 2 (K and V) x 2 B = 4 MiB per layer

x 12 layers = 48 MiB per token

\\\`

Take an example, at 80 tok/s that is about 3.9 GB/s across the link. It is a small fraction of a PCIe 4.0 x16 slot, and most of it overlaps with compute.

Only few things need to stay on the GPU. The model itself, and a 2-byte slot plus the pooled index key, which is \1 x (128 / 4) x 2 B = 64 B\. Together they are 66 B per token per layer, against 2,048 B for a full row.

▲
137
-2
21👁
r/LocalLLaMA · u/NineThreeTilNow · 23d ago
Update : Small model + Engram

I posted something about a 9b model a few days ago.

The real problem was 2 fold. Someone suggested the size was too big to prove it all out. It's a fair argument. I needed the model depth though. The second problem was the "ability" of these models and the fact that labs (with money) produce these models still. No real usefulness in my mind. A 2b model? maybe interesting. Want a chatbot? - this could be it.

Further, the Llama licensing at the core of the model was highly problematic and had to be thrown in the trash. It's vastly too restrictive and I couldn't Apache 2.0 anything.

So I moved to using the OLMo tokenizer. Except I shrunk the d\_model down to 2048 so I could build a tiny 2b model.

The size of the vocab with 1/2/3 gram forces the Engram table table to a full 1b. That's 50% where DS suggested like \~10-20%.

So basically 2b model + 1b Engram.

The model, because of the depth now allowed by 2048 d\_model allows me to push 40 total SWA/Global blocks to take advantage of attention based residual streams (Moonshot Kimi K3 style) in a dense model.

Data, is pulled from standard Wiki style data sources in my initial training. The data is pumped through the 7b OLMo model to generate the probabilities. This tells the model that the next token is a probability of 32 different tokens.

The big surprise was the results. Even after ONLY 15m tokens, the model is surprisingly coherent. Had I given pure 1 hot cross entropy, 15m tokens would do nothing. I'd get gibberish. However, I borrowed embeddings, and LM\_head that was down projected from 5k -> 2048 d\_model. This preserves \~65% of the data the "big" model had in the embedding when spectrum analysis is done.

The training is limited at 15m tokens, so it's like... babbling about Roman war history and stuff that exists in that token subset. It doesn't fall in to the classic early model phase where it repeats a token over and over though. That's what pretraining 15m generally gives you at first.

Anyways, this was an update and progress report for anyone who cares about this crap... I'm still processing all of the Wikipedia chunk from HuggingFace and working to train it. It'll take time.

If you read this far, thank you. If you have questions, I'm happy to reply. Someone suggested I was a kook who didn't understand ML previously. I started in ML some 20 years ago and held a brief (6 month) stint at Anthropic red teaming the original Opus 4 model before their "Constitutional" paper came out. I'm vaguely referenced in the paper as a "red teamer" I guess. I do this mostly because I love it.

I figured if I built an Apache 2.0 model, I'd at least want people to play with it. It appears 100% trainable on 24gb of VRAM. Training code / Model / Data will follow eventually. It's on HF / Github for now.

The old post is here :

https://old.reddit.com/r/LocalLLaMA/comments/1wezm58/is\_there\_still\_strong\_interest\_in\_a\_dense\_9b\_model/

edit;

Update can be found here :

https://www.reddit.com/r/LocalLLaMA/comments/1wnzu2f/engram_gone_wild_2b_mode…

▲
123
 
22👁
▲
114
+4
17👁
r/LocalLLaMA · u/EcstaticDentist · 25d ago
Qwen3.8 27b Game Dev Part 2 post image

Qwen3.8 27b may not be able to whip up 3d models & GLB’s but it will sure do with them as you please once you drop them in the game repo. Absolutely fascinating

▲
108
+2
14👁
r/LocalLLaMA · u/CharlesStross · 23d ago
600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"? post image

It's not the brightest bulb but it's my new drudgework model for read+find or code tasks I'm willing to let it brute force. Even if it takes 20x more tokens, that's still faster than many local models. Not quite Cerebras but still pretty fun to drive.

▲
94
-3
24👁
▲
94
-3
22👁
r/LocalLLaMA · u/SomewhereAtWork · 24d ago
LocalJev?

Jev is a model to produce structured output (choices) from input text. It apparently can play (not run!) Doom.

https://typesafe.ai/blog/introducing-system-one-models-and-jev

Is there already a open implementation of this kind of model?

▲
99
+4
22👁
r/LocalLLaMA · u/Zulfiqaar · 23d ago
IFM/K2-Horizon-7B-Uno · Hugging Face - 5200tps with no quality loss

IFM released K2-Horizon-7B , diffusion augmented LLM at upto 5200 tps, with claimed lossless speedup.

causal LLM architecture and adds a plug-and-play diffusion adapter alongside the autoregressive weights.

https://arxiv.org/abs/2609.04010

▲
82
+2
31👁
r/LocalLLaMA · u/satnl · 24d ago
Qwen 3.8 27b is a amazing model, for the first time I see a local model found its own away to open a browser and test

I was testing this quantization IQ3\_XXS from GSQ-RCO with PI.

It is a heavy quantization case, the model is in IQ3\_XXS and KV cache in (Q4\_0, Q4\_0).

I asked it to make the flight simulator, using that popular prompt. For my surprise, when I went verify the session I saw some screenshots, but I don't have mcp or tools in the PI, it's an almost clean installation, only with some personal skills that dont have any mentions to browser testing.

Investigating the session I have noted that it found its way to open a browser for test by itself. My daily harness is opencode, and for not spending the 7k tokens with the chrome mcp in each session, I let the mcp disabled and I only enable it if the session needs. But this experience make think that I probably only need line in the agents md saying that he can use the headless browser with CDP when it needs.

https://preview.redd.it/wrgltip9rzph1.png?width=1952&format=png&auto=…

OK, let me write the file. Also test: I can't easily run a browser here. But I can at least do a syntax check of the JS via node (extract
the script? maybe parse with node --check).

----------------------

Now, final verification: run node --check again, and additionally do a simple DOM-less smoke test? Can't run three.js without a browser.
Syntax check + math tests are enough. Could also try opening it in a headless browser if available? Check whether the machine has
headless Chrome or playwright? That might be overkill; but if I could do an actual rendering check, I could catch runtime errors (e.g.,
shader compilation). Let me check whether headless Chrome is available.

$ cd /tmp && awk '/<script type="module">/{f=1;next} /<\/script>/{f=0} f'
/Users/[...]/zeroshot/Qwen3.8-27B-UD-IQ3_XXS_GSQ-RCO-pi-xhigh/index.html > sim.mjs && node --check sim.mjs
&& echo "SYNTAX OK"

SYNTAX OK

Took 0.1s

$ ls /Applications/ | grep -i -E 'chrome|chromium' ; which chromium chrome google-chrome 2>/dev/null; ls ~/Library/Caches/ms-playwright
2>/dev/null | head

... (5 earlier lines, ctrl+o to expand)
ffmpeg-1011
firefox-1532
firefox-1538
webkit-2311
webkit-2336

Took 0.0s

Chrome is available, and playwright's chromium headless shell is also available. Let's do a headless smoke test: load the page, wait a
few seconds, and capture console errors and screenshots. The page loads three.js from a CDN — the network needs to be working. Let's use
playwright's chromium (does npx playwright exist? The fact that the ms-playwright cache exists means playwright is installed somewhere).
Let me check if the playwright npm package is available globally, or use chrome directly with --headless.

The simplest approach: use Chrome headless directly:
/Applications/Google Chrome.app/Contents/MacOS/Google Chrome --headless --disable-gpu --screenshot=/tmp/shot.png --window-size=1600,900
--virtual-time-budget=8000 file://...

However, WebGL in headless Chrome: it works with SwiftShader (software GL) via --use-gl=swiftshader or --enable-unsafe-swiftshader. In
newer versions of Chrome, the software fallback requires --enable-unsafe-swiftshader. Let's try it. Also, console errors:
--enable-logging=stderr --v=1 will give us the console.

Let me write a small node script using playwright if available. First, check the availability of node playwright.

Chrome is available — let me do a headless smoke test to catch runtime/shader errors and grab a screenshot:

$ node -e "try{require.resolve('playwright');console.log('playwright OK')}catch(e){console.log('no playwright')}"; ls
~/Library/Caches/ms-playwright/chromium-1234/ 2>/dev/null | head -3

no playwright
chrome-mac-arm64
DEPENDENCIES_VALIDATED
INSTALLATION_COMPLETE

Took 0.1s

There's no playwright module, but the Chromium headless shell exists. Either find the path to the headless shell binary and run it
directly, or use Google Chrome in headless mode with SwiftShader.

The path for Playwright's Chromium headless shell:
~/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-mac-arm64/chrome-headless-shell? Let me check the structure.


$ find ~/Library/Caches/ms-playwright/chromium_headless_shell-1234 -name '*headless*' -o -name 'chrome*' | head; find
~/Library/Caches/ms-playwright/chromium-1234 -name 'Chromium' -o -name 'chrome' | head

... (3 earlier lines, ctrl+o to expand)
s.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/chrome-headless-shell
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/headless_lib_strings.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/headless_lib_data.pak
/Users/[...]/Library/Caches/ms-playwright/chromium_headless_shell-1234/chrome-headless-shell-mac-arm64/LICENSE.headless_shell

Took 0.1s

Use chrome-headless-shell — but to catch console errors and drive key input, we need CDP (Chrome DevTools Protocol).
chrome-headless-shell supports --remote-debugging-port. I could also write a small Node script that speaks raw CDP over WebSocket... but
that would need a ws client. Node doesn't have a built-in WebSocket client until Node 22+ (Node 22+ has a global WebSocket). What version
of Node is it? Let me check. Node 24 (from the nvm path v24.21.0) — a global WebSocket exists.

Plan:
1. Launch chrome-headless-shell with --remote-debugging-port=9222, --no-sandbox, --use-gl=swiftshader? For headless shell, WebGL via
SwiftShader: --enable-unsafe-swiftshader flag; also --no-sandbox is needed? For headless shell, maybe --no-sandbox isn't necessary,
but I'll add it just in case. Also, a new profile directory with --user-data-dir=/tmp/cdp-profile.
2. Node script: fetch http://127.0.0.1:9222/json/new?url=file:///... (PUT) to create a tab, connect via WebSocket to
webSocketDebuggerUrl, and do the following:
- Enable Runtime and Log
- Collect console messages and exceptions
- Wait about 6 seconds (let the first frame render)
- Dispatch key events? Simulating keydown via CDP: Input.dispatchKeyEvent. Possible, but keyboard isn't required for the smoke test.
- Page.captureScreenshot → save to PNG
- Also, evaluate expressions: get the chunk count, and confirm the renderer is working? For example
document.querySelectorAll('canvas').length, window.__stats... I haven't exposed the stats. I can evaluate JS in the page context:
the module's scope isn't global, but I can check document.getElementById('alt').textContent (the HUD should show altitude after
0.12 seconds), and check the canvas pixel colors via toDataURL? Easier: check the HUD text + screenshot.
3. Report console errors.

Let me write this script.

I'll drive the headless shell over CDP with Node's built-in WebSocket to capture console errors and a screenshot:

------------

> why did you used a browser? did you already had any instruction about it?

User's question: "Why did you use a browser? Was there any prior instruction to do so?"

Honest answer: there was no prior instruction to use a browser. What I chose on my own was:
1. To do syntax checking with node --check (low cost, so I did that first).
2. Since a single HTML file with WebGL + CDN can't be fully verified without actually rendering it, I used a locally available headless
Chrome (I found Google Chrome and Playwright's chromium headless shell on the machine). It's to catch runtime errors (shader
compilation, WebGL failures, logic errors like the inverted bank direction), and to take screenshots as visual evidence.

There's no instruction in AGENTS.md or in the skills that requires using a browser. There's nothing in the skill list that mandates it
either (the qa-manual-guide skill is about writing a guide, and I didn't use it since it wasn't explicitly requested… actually, this
skill says "when an implemented feature needs an evidence-based local manual testing guide…" — what the user asked for was a simulator,
not a manual testing guide. My verification was a homegrown approach).

Why I used it: a single-file browser game can't be verified with unit tests alone. Headless rendering would let me detect actual failures
(e.g., I actually found an inverted A/D bank bug that a syntax check could never find), and screenshots let me verify the visual quality
(mountain appearance, sun position, cloud scale) — the "beautiful and relaxing" requirement is inherently visual.

I should answer concisely and honestly: no, there was no instruction. I discovered the browser on the machine and used it for
evidence-based verification. The main findings were the inverted turn bug and visual confirmation. Also note that it was only used for
testing and didn't modify any files inside the project.

Keep the reply short and structured.

No — there was no instruction to use a browser. That was my own choice for verification, not from AGENTS.md or any skill.

Why I did it:
- A single-HTML-file WebGL game can't be properly validated by syntax checks alone — the interesting failures (shader compile errors,
WebGL context issues, sign errors in the flight math) only surface when it actually runs.
- I noticed Google Chrome and Playwright headless Chromium were available locally, so I drove a headless shell over CDP: load the page,
capture console/exception logs, simulate key presses, read the HUD values, and take screenshots.

▲
77
-1
27👁
r/LocalLLaMA · u/Gold-Bat-3225 · 23d ago
Humor Arena: Which LLM is the funniest? post image

We compared 20 model versions on 360 frozen joke prompts, with four jokes requested per prompt and model names hidden from our humor-trained judge.

Fable 5 had the highest estimated score: 66.8 points per 100 comparisons against the rival field. Fable 5.1 scored 58.2. Scores count a win as one point and a tie as half. The models near the top have overlapping uncertainty intervals. The reason we say estimated is that we scale the scores based on the scores of

The scores come from our own automated judge. A separate audit checked the judge with 1,400 ratings from 50 people; that was not a fresh human evaluation of Fable 5.1’s outputs. We specifically fine tuned an OS judge to rate the jokes and it correlates more highly with human preferences than any other model.

The full details here:
https://laugh.so/research/joke-generation/

Would love to hear your feedback!

▲
76
+1
26👁
r/LocalLLaMA · u/Hefty_Wolverine_553 · 24d ago
What's the current best LLM uncensoring method?

With the recent Nvidia Huggingface acquisition and frontier AI labs screaming about safety and putting guardrails everywhere, I think it's important that we have local models that aren't affected by arbitrary guardrails set during training. To be clear, this post NOT about Enterprise Resource Planning (ERP). Censorship in an LLM can highly affect its abilities to do many legitimately useful things (note GPT-OSS, Fable 5), and going forth I believe censorship will only get more and more strict.

There have been many resources and posts about uncensored models using abliteration, heretic, and probably many other methods that I'm not aware of. However, it seems like all of this information is scattered about the place, and Huggingface is essentially flooded with "uncensored" variants of basically every popular open source model, many of which don't work well, affect the model's intelligence greatly, and have "KLD 0.0001" presumably from measuring against Wikitext datasets. I'm hoping that this post can gather some more useful information to serve as a starting point/discussion of which uncensoring methods work best.

Please share your experiences with specific uncensoring methods (not just a single uncensored model) and how well they work (both good and bad), as well as any notable people doing consistent/high quality work on uncensoring models.

▲
69
+4
21👁
▲
64
 
35👁
r/LocalLLaMA · u/Qwen30bEnjoyer · 25d ago
Open Source Appreciation Post

It's late at night in the lab, I've been working on a basic script for a virology project, and holy hell the safeguards have been pissing me off.

Mirroring detectEVE data over rsync to my laptop by making a zip file first? No no no, great safety mogul DARIO demands there be NO file transfer today. Request blocked, reported, labeled [cyber]. Yet, Deepseek V4.1 does it with no complaint.

I got tired of reading papers - so I ask Claude - "Does this PDF go over binary host virus infections?" Immediately blocked for biological safety risk. Deepseek V4.1 tells me it doesn't have the data I need without drama.

Bioinformatics server goes down and I need help getting it back up by getting the outputs of my diagnostic scripts to the mounted usb drive? Oops, its named exfil. Looks scawy. No transfer of output logs for you due to CYBER risk.

Would CNNs be a good architecture to start on phage-host prediction? Claude wouldn't tell me because information you can find in a google search is too dangerous for me to handle apparently - but once again Deepseek v4.1 tells me that GCNs are where I should start.

I get that Virology is a particularly sensitive topic, but come on. Imagine if Google had taken the same safety approach in the early days of search. Like if Google made it so that you either had to give up your identification and where you work to them, or go to the library and search by hand. It's almost unthinkable, yet in the name of the almighty Safety, Dario and Altman continue working to keep scientific knowledge locked away.

I think for the sake of all scientists, open source AI must win because we need a tool that just WORKS without egomaniacs micromanaging us or shaking us down for ID.

▲
60
-1
26👁
r/LocalLLaMA · u/No_Run8812 · 24d ago
Upgraded my local setup with 2 rtx pros and it's amazing. post image

Follow up post of https://www.reddit.com/r/LocalLLaMA/s/nGMyKswrch.

Thanks everyone who replied. I didn't change the specs. Might be loosing some of the memory bandwidth but will scale in future if I need to.

It took me 2.5 days to build it because one of the GPU connected to PSU was loosing power whenever I load anything on the GPU, so I had to rewire every connection again to identify the fault. I am glad the system is working because I was apprehensive if this will work (I am software dev, getting my hands dirty with hardware for the 3rd time in life). My finger tips still hurt from pulling the cables from motherboard and PSUs.

To enable the full potential of the system, I had to enable peer to peer communication between the GPUs, cuda graph, tensor parallelism. I have capped both the GPUs at 500W (no reason, just didn't want GPUs to run on its full capacity).

Also, I had to open my box, because temps were shooting high, and fans were making weird noises.

I am running:

  1. Qwen 3.8 flash next 8 bit
  1. Deepseek v4 flash 0731 (official)

I have a M3 ultra 512, LLMs run on it, but I personally find it useless for inference. My head just hurts watching it work slow. On the other hand this new system is killing it, decode 150 tk/s and prefill 10K tk/s.

Qwen is good, but most of the context is consumed by thinking tokens, I was checking if it's a good idea to not use the thinking token. I barely have vram left for concurrent requests with full context window. Loving the Deepseek 1M context, and I also have room for 4 concurrent requests. Both of them are okay model, even if they make mistakes, I don't notice because of the speed. It's just fast, makes an error, corrects it moves on.

Finally the day is here when I can save on monthly subscriptions and not worry about the weekly or 5 hours limit. I have already setup my server with openclaw, opencode, openweb UI and Tailscale.

Has anyone experience excluding the thinking tokens of Qwen from the context and keep the final result? Was there any impact on the performance or accuracy of the model?

Any suggestions, what else I should install on it? Any new models to try?

▲
60
+2
35👁
▲
54
-3
26👁
r/LocalLLaMA · u/WebAssemblyMan · 24d ago
Recurrent Looped Transformer post image

Recurrent Looped Transformer (RLT)passes the decoder's final hidden state to the next token, together with that token's causal encoder representation. The decoder reads encoder-derived global KV memory and maintains a sliding-window attention (SWA) cache at every layer. The same update runs over prompt and response tokens.

More effective reasoning depth!

https://github.com/yifanzhang-pro/recurrent-looped-tranformer