Like.. yea bro, I knew GLM was cool. Now everyone does.
Like.. yea bro, I knew GLM was cool. Now everyone does.
Here are some inspirational quotes you can put into the comments:
You may think that jev, OpenClaw or TurboQuant are super cool, but actually the coolest LLM invention happened two years ago
As we all know, the best source of reliable information about LLMs is YouTube:
https://preview.redd.it/4tv0ctbwyfsh1.png?width=2544&format=png&auto=…
Back in September 2024, Reflection 70B appeared out of nowhere and was announced as an open-source model that supposedly destroyed GPT-4o
There was only one small problem. People downloaded it. And tested it :(
https://preview.redd.it/kamhrp9izfsh1.png?width=1514&format=png&auto=…
It turned out that Reflection 70B was basically a Llama 3.1
https://preview.redd.it/ghmof4l50gsh1.png?width=1524&format=png&auto=…
but at the end the mystery was solved
https://preview.redd.it/44eb3jw90gsh1.png?width=1530&format=png&auto=…
Let this be a moment of reflection on the current hypes in LocalLLaMA.
great summary by Maziyar PANAHI https://x.com/MaziyarPanahi/status/1838559480658710982
I played a game of a custom chess variant against an claude opus 5.5, and it beat me.
The game combined several rule changes: the board wraps around from the h-file to the a-file, knights move three squares in one direction and one sideways, and captured pieces can be dropped back onto the board, as in crazyhouse. I gave the model the rules, the starting position, and a board diagram, then asked it to reply with one legal move at a time.
Also I recreated this board game https://nika-game.com/ and played with claude, and it still win, although I believe it is really unpopuplar and old game without much training data available (claude didn't even know the rules initially)
I used to think chess exposed a fundamental limitation of LLMs: keeping track of a changing board, following exact rules, and planning ahead seemed like a poor fit for a language model. This game made me reconsider that assumption.
So I’m curious: what broad classes of problems do you think LLMs still can’t solve reliably? Are there limitations you consider fundamental/archiectural, rather than problems that might improve with better models, more computation, or tools? What would be a good test?
Anthropic released the GLM article today. Trump is getting very involved.
Do you foresee Chinese open weight models getting banned soon?
I was optimizing my diet using a cloud AI provider, and I found my Pi agent stuck like this. Looks like I had too many mushrooms lol.
Here's a metric dashboard giving an idea of the last few days.
Been testing with a variety of different agentic coding use-cases, mostly using a pi harness.
qwen3.8-flash-next has seriously exceeded my expectations (used https://huggingface.co/tcclaviger/Qwen3.8-Flash-Next-MXFP4-FP8)
Both speed and quality have surprised me, given that I can get 3-5 concurrent streams going with \~100t/s gen each, and single stream easily gets to 150+t/s. Prefill is 10k+t/s
ninfer, dwarfstar, Splash, llamAmpere, gufo, etc.
We've all seen them popping up, great tok/s, people loving them. Forks of llama.cpp or another engine, or made from scratch.
For better or worse, the list will continue to grow
They work so well because they dodge a main difficulty of software, generality, and just implement for a single model/hardware combo (or a few), and then optimize kernels/compute graph for that one case. Highly 'overfit' codebases that beat well-known inference engines (llama.cpp, vLLM, etc) and incidentally will be completely forgotten in 6 months.
But new ones will take their place...
THESIS
One-off engines will become the norm. Llama.cpp, vllm, etc, will not make sense for most people to use, because they're slower
A few axioms you probably accept:
String these axioms together, and I arrive at
Thank you for coming to my ted talk
IMPLICATIONS
Most obvious example: it would be annoying to have a different usage API for every engine, so we already standardized on OpenAI API compatibility years ago.
Is that also true for cli arguments/configs? The packaged gui (llama-server)? Benchmarking tools (llama-bench)? Logging, model format, Etc?
One-off engines that replicate the experience of everything wrapping the inference itself will be more seamless to adopt. Case in point, the main reason I haven't tried any of these new one-off engines myself is it was annoying enough to figure out how to drive llama.cpp properly. Don't want to do that again unless it's really worth it.
There's probably a place for an open source project that standardizes all of this and makes it easy for one-off engines to adopt.
ALTERNATIVE FUTURES
Scenarios where the one-off future doesn't happen:
P.S. there's growth in a dimension separate from single model/hardware engines which is more like "frontrunning a general inference engine's features because it's slower to pull in PRs". Freetoken, BeeLlama, etc. Not as model- or hardware- specific as the other examples I've given. Haven't thought much about that dimension.
IQuest-Q1 is a Mixture-of-Experts (MoE) model developed by IQuest for agentic coding, reasoning, and multi-step tool use. It comprises approximately 320B total parameters, with an estimated 15B parameters activated per token.
I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.
Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.
And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.
The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.
All figures here: https://github.com/0xShug0/audio.cpp/tree/main/assets/figure/
We have released Qwen3.8-Flash-Next quantized with GSQ and RCO, together with a second, capability-targeted build in which half of the model's experts have been removed.
Flash-Next is a sparse mixture-of-experts model: 512 routed experts per layer across 48 layers, 176.9B parameters, 354 GB at BF16.
What's inside
Results:
At 3.50 bpw the model matches the BF16 base on every benchmark evaluated.
Coder (capability pruned model):
Instead of storing every parameter at lower precision, half of the routed experts are removed from the model: 256 of 512 per layer, selected by RCO optimising the KL divergence against the unpruned model. The retained weights remain at 3.5 bpw. Pruning and quantization compound, and the combined effect is an average of 1.89 bits per parameter of the original transformer. The averaged bitwidth amortises the removed experts over the original parameter count, and therefore expresses the joint effect of pruning and quantization. No individual weight is stored at 1.89 bits.
The practical consequence is that a 176.9B-parameter model has a resident working set of 29.6 GB, since the n-gram shard is a lookup table and may be served from disk. This is within the capacity of a single 32 GB accelerator.
Both measured at xhigh reasoning effort.
Links
Both repositories ship the complete per-tensor RCO allocation.
The Coder build is an experimental release and feedback is welcome, particularly on capabilities that were not represented in the calibration mixture. Requests for models to quantize or prune are also welcome.
From the ISTA Deep Algorithms and Systems Lab.
​
I created this repo to help the DGX Spark users that have a spare 10-24 GB GPU at home to squeeze some extra memory out of a single Spark or a Sparks cluster.
It moves the spec-decode draft model off your Sparks onto that GPU: the freed GB of memory can be used for extra context, or better quant quality. Supports both TCP and RDMA, shipped as eugr-vllm compatible mods:
https://github.com/ciprianveg/gb10-vllm/tree/main/remote-dspark
Hi there!
My name is Jerry and I recently built Enclosure, e deterministic environment for testing LLMs. It's a simulated office enfironment https://github.com/openconstruct/Enclosure with tools agents are used to, like slack, calendar, mail, chat and more. For conversations is uses AIML instead of a model, so it is totally deterministic.
The first test I made(had Claude make) is the one I had in mind when I designed Enclosure, a personality test for models. It's not a winnable benchmark, but rather a tool to help people find the mdoels that fit their workstyle best. It's called disposition and you can find it here: https://github.com/openconstruct/disposition
I tested the cheapest models on Alibab modelstudio already, after I refill my opencode go next month, I will probably test more. I am asking the community to test some models if you think it's a cool project.
Instructions are in the Disposition repo, and the submission repo is here: https://github.com/openconstruct/disposition
We did a study: does post training actually make LLMs funnier?
We used open models that publish every stage of post training, so we could compare a base model with the future models it became: Tulu 3 (on Llama 3.1 70B), OLMo 3.1 32B and Qwen2.5. We tracked 11 stages, 100 joke prompts, 64 human raters and 2,330 head-to-head judgments.
What we found: post training makes models funnier, but reduces diversity of response.
\- In 5 of 7 training steps, the later model's jokes were judged funnier. Jokes also got 10–20 words shorter after early post training, so they get to the punchline faster.
\- In 6 of 7 steps, the jokes a model wrote for the same prompt got more similar to each other. Ask for eight jokes on one premise and you get eight versions of the same joke. The biggest drop was Qwen2.5 base to instruct.
\- Asking the model to plan a line or two before the joke cut variety in all 4 models we tried, with no reliable gain in funniness.
\- A comedian persona won back a little variety in all 4 models, but only made the jokes funnier in 2 of them.
Humans judged the base versus final. A model judge calibrated on those votes compares the stages in between.
Full report and paper below. Which open models should we run through this next?
For example, directly using claude code or code, which is then hooked up to automatically delegate the actual code writing tasks to a local model like qwen 3.8 flash next, to save on cloud usage limits.
I’m imagining the loop would be:
User writes prompt
Claude/codex thinks about it and the plan
Claude/codex sends the specific and bounded coding instructions to the local model+harness (opencode, pi, etc) via api endpoint or MCP, with clear instructions on a defined endpoint
One the local model+harness hits the clear endpoint/“done” step, it sends a ping back to claude/codex
Claude/codex then verifies the output and then thinks about next steps to instruct the local model+harness on
Does this actually lead to improved savings on the cloud model usage while preserving code quality? Or does this end up being unnecessarily complex and not saving on any cloud usage
Forlinx Embedded has listed an M.2 AI accelerator card based on Rockchip’s RK1820 and RK1828 processors, providing 20 TOPS of INT8 computing performance and up to 5GB of integrated DRAM. The module uses an M.2 2280 interface and is designed to handle local AI inference, including large language models, vision-language models, and computer vision workloads on embedded Linux and Android systems.
I wanted a setup where I can compare all the harnesses with all the local models.
It turned out to be a rabbit hole. For instance, you would not only have to test all the harnesses (being sure that they are well configured), but also all models, with all their flavors, and this for all kinds of hardware.
Everyone can do their share, but no one can pretend to do all possible tests extensively.
To solve this, I created a website where everyone can test the configurations they want and share the results if they want. You can try it here: airbench.ai
There is a leaderboard where I share the tests I'm making, but I hope I can populate it with tests from others. https://airbench.ai/leaderboard?k=poL
My hope is to turn this into a fully distributed Agent Benchmark.
Let me know what you think.
I’ve been working on an open-source project called WebBrain that gives LLMs the ability to see and operate a browser.
One thing I really wanted to avoid was making the browser agent dependent on a single cloud model/provider.
So WebBrain can work with local models through things like LM Studio and Ollama, as well as cloud APIs if you want them.
I also ended up training a small vision model specifically for browser tasks:
webbrain-vl-2-450M
It’s based on LFM-2.5-VL-450M and fine-tuned on browser screenshots/tasks. The idea is that instead of sending every screenshot to a giant multimodal model, some browser perception can happen with a very small model locally.
It can run through WebGPU directly on the user's machine.
The agent itself combines screenshots with the browser accessibility tree rather than relying entirely on DOM parsing.
Current architecture is roughly:
• screenshot + accessibility tree for perception
• browser-specialized tiny VLM where useful
• model-agnostic planner
• local models via LM Studio/Ollama
• Chrome / Edge / Firefox / Chromium support
• optional cloud execution
• open source
I'm especially interested in figuring out how far browser agents can realistically go with small local models rather than GPT/Claude-scale models.
Repo: https://github.com/webbrain-one/webbrain
Model: https://huggingface.co/webbrain-one/webbrain-vl-2-450M
Dataset: https://huggingface.co/datasets/webbrain-one/webbrain-vl-2-450M-dataset
Would be very interested in feedback from people here running smaller Qwen/LFM/MiniCPM/etc. models locally — particularly what model you would try as the planner.
Anyone have a good config they've found for QFN on a A40 (or similar Ampere GPU(s) with 48GB VRAM)? The system the card is in has 128GB of RAM (DDR3); right now it's running 27B but I'm curious if there's a way to move to QFN to maybe get better speed/a little more smarts. TY in advance!
TypeSafe's Jev answers typed questions (yes/no, pick one label, pick a rubric level) with a probability per answer instead of text. Its docs say it takes text only: "Images, audio, and video are not supported (yet)." I wanted the same kind of answer about photos, from an open model on my own card, so I built typevet (MIT, Python 3.12).
How it works. No sampling, no parsing. typevet composes the native Gemma 4 turn itself, ends the prompt with the empty-thought no-thinking prefill, and reads the next-token distribution over the allowed answer tokens only. What comes back is a label plus a probability for every option, or an error. Images go to llama.cpp's /completion as base64 in prompt.multimodal_data with one media marker per image; on vLLM they go as image_url blocks. There's also a JSON path that returns an object passing your JSON Schema, or raises.
Local setup. Gemma 4 31B, my 24 GiB vramfit pack (byte-identical to the file on HF, projector sidecar for vision), llama.cpp b11223, one RTX 4090. Nothing leaves the machine.
The receipt test. 6 real receipts from CORD v2 (CC BY 4.0), 3 synthetic expense claims each: right total, two digits swapped, digits masked with ?. One three-label Choice: match / mismatch / insufficient_evidence. Each claim sent as text only, then with the receipt photo.
Example: claim says 646329, receipt says 664,329. Text only: match at 0.99964. With the photo: mismatch at 0.99999. Every swapped total was caught at 0.99998 or higher, and the masked ones abstained every time. The tests also check the image actually arrived: each photo added 228 to 1,108 prompt tokens here, and the gate fails if the count doesn't grow.
Hosted. Same code against vLLM 0.30.0, BF16 Gemma 4 31B, one H100: the receipt test went 18/18, and reversing the label order flipped 0 of 18 answers. Throughput on 480 Banking77 records with two questions each: 0.24 s median per record at 1 in flight, 39.6 records/s at 64 in flight, 0 errors.
Prior art, credit where due. The text-side decision model comes from TypeLLM (SGLang), which added its own image input on 9/24; typevet's image path is separate code on llama.cpp's request shape. allanrbo posted a Jev-like single script for Gemma 4 12B with webcam images on 9/25. VQAScore has read the probability of "Yes" from VLMs since 2024. typevet's angle: a library, not a script, the 31B on one 24 GiB card, the same code on vLLM, and image-arrival checks.
Scope: 18 claims, one run per server. The probabilities are the model's confidence, not calibrated.
So I got this new local model on my system & wanted a mini test to see if it was actually smart or just confidently wrong (as I like to do with new models I haven't yet tried). I asked the cloud assistant to cook up a pretty rigorous 10 point diagnostic suite: reasoning traps, Python semantics, SQL fluency, strict instruction following, the whole gauntlet. At first it felt like DeepSeek V4-Flash frontier cloud intelligence vs my little local quantized guy. Classic quick test.
Then I ran it on the local model & shared the results back. The assistant was grading it & found out it got item #4 wrong. That item was a logic puzzle. The assistant thought one statement had to be false, but the local model was like nah, the set of statements is logically consistent, so your question is built on a false premise. It literally refused the leading question. I was like "wait wtf". That was the turning point fr. The local model solved a trap that the cloud model just completely keyed wrong.
The local model is a Q2\_K quantization of Nex-N2.5-mini, which is a fine-tuned Qwen3.5-MoE architecture. A 2-bit quant. Normally people call that low quality. But it outperformed a frontier model on a logic trap. The assistant went from “I am grader” to genuine admiration, saying resisting a leading question is high-level reasoning. Lowkey pretty humbling for the cloud side.
The whole thing made me think about emergent intelligence & AI democratization. Less giant centralized compute, more efficient specialized local stuff. The student corrected the teacher. Efficiency & MoE architecture maybe can beat raw parameter count sometimes. The mini test felt like a rite of passage for the local model. Its kinda like it became a validated thinker instead of just software. Q2 compression is also symbolic resilience, because despite being squished, the reasoning circuits stayed intact. And the assistant admitting it was wrong made the local model’s win feel more real.
And the craziest part? It went 10/10. This wasn't some easy benchmark either. It was a deliberately nasty little diagnostic with multiple ways for a heavily quantized model to screw up, and it didn't.
I need to make it clear that I am not claiming that Q2 ORCA model is generally superior to DeepSeek V4-Flash. But rather as an anecdotal demonstration that an extremely compressed local model can sometimes catch a reasoning failure in a frontier model & maintain much of it's reasoning power when done correctly & skillfully. It is a testament to how even under Q2 compression, it still preserved the model's “reasoning circuits."
Final verdict from the assistant: model is in excellent shape & ready for real work. So yeah, a local model dunked on the cloud AI. I’m happy for what this means for the future of local ai.
MODELS USED for quick test:
Local Model: Nex N2.5 Mini Uncensored
Frontier Model: DeepSeek V4.1
EDIT / CORRECTION bc I fucked this part up 😅
Small but important correction to the post. I originally called the frontier model I tested DeepSeek V4-Flash. That's not the right model name for the one I actually used on the DeepSeek website. It was DeepSeek V4.1-Flash.
Also, V4-Flash itself is a local/open-weight model, so my original wording made it sound like I was comparing my local model against some cloud-only AI. That's not accurate & that's on me.
The actual comparison was my local Q2\_K Nex-N2.5-mini vs DeepSeek V4.1-Flash through the DeepSeek website.
So yeah, V4.1-Flash is the model I should have named in the original post.
I'm leaving this correction here instead of quietly changing the post bc I don't wanna bullshit anybody or make it look like I didn't make the mistake. I got the model name wrong, someone pointed it out & I'm correcting it. 🤷♂️
The actual 10/10 result & the logic trap part of the test are unchanged.
\*\*TL;DR:\*\* I tested a local Q2\_K Nex-N2.5-mini on a 10-part mini test, it caught a logic trap the cloud assistant got wrong, & the assistant basically certified it as ready for real work. It got 10/10 correct.
That's why we have to fight against closed source centralized AI and closed recipes open weights overcoming the frivolous "gifts" exchanged with them by giving away our souls to those greedy entities if we want to be free in the future instead of being squeezed like lemons/at mercy/enslaved in the paws of these soulless folks that have a id of any single one of us at their disposal to switch us on/off at their leisure/convenience.
I was just wondering because I see it mentioned in threads here a lot... When people talk about LLMs used for roleplay, that's a euphemism for dirty/sexy chats right? Kind of like how "torrenting linux ISOs" is really just pirating copywritten media. Or are you guys really burning tokens pretending to talk to a medieval shopkeeper?
I know I sound crazy, but i’m using a strix halo and finding that GLM 4.7 flash just runs significantly better than qwen 3.6 35 a3b. But its obviously quite old at this point, i wish we had a new glm that was 4.7 flash sized but is anyone using a model that they have found better than this.
My brief testing with qwen shows that glm is better at tool calling and better at world knowledge, but maybe there’s a chance my qwen set up is wrong
I'm making a local roleplaying model for my girlfriend's community server.
It's supposed to do roleplay, have a specific talking style(dry while answering to normal stuff and extensive when talking about lore), never talk out of roleplay, have hundreds of pages of lore and information and their rank in lore( discord roles maybe?).
It's basically supposed to be just an LLM you can converse with that answers in a specific talking style and has all the lore info.
For now I implemented: 10 ish% of the written lore, and it recognizes 3 people, but by discord ID that i inserted in the system prompt.
I'm using GPT 5.6 sol(and well 6 sol now) for doing stuff, but i keep running into a problem.
When i reach a nice point where the model has a nice talking style and knows information well enough, i tell Sol to add this new lorebook and this info, here now everything breaks.
Talking style is fucked, It doesen't recognise people individually anymore, when asked about other unrelated lore it just gets it wrong or hallucinates, or even starts roleplaying as one of the characters in it's lore book out of nowhere.
It's connected to discord through a discord bridge that Sol made and a developer dashboard bot.
I'm using GPT OSS20b on MXFP4, single 9070xt and 32gb ddr4.
Should I maybe fine tune it?
PS. im a beginner in AI so if it wasn't obvious i do NOT know what im doing but im trying my best for her.
Thanks to tcclaviger, vLLM now has expert RAM offloading support (link). This makes frontier models much more accessible on a local setup!
I was able to run the original DeepSeek-V4-Flash-Vision-Exp on four R9700s.
podman run --rm -it \
--init \
--network host \
--ulimit memlock=-1:-1 \
-v /models:/models:ro \
-v ~/.vllm-cache:/cache \
-e VLLM_ROCM_USE_AITER=0 \
-e ROCR_VISIBLE_DEVICES=0,1,2,3 \
-e VLLM_CACHE_ROOT=/cache/vllm \
-e TORCHINDUCTOR_CACHE_DIR=/cache/inductor \
-e TRITON_CACHE_DIR=/cache/triton \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--annotation run.oci.keep_original_groups=1 \
--security-opt label=disable \
--security-opt seccomp=unconfined \
--shm-size 160g \
docker.io/tcclaviger/vllm@sha256:ef99b3d07c3f15e7978528c7510762ba024df9ab4242070d8ed092cd4cc1a694 \
/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp \
--served-model-name DeepSeek-V4-Flash-Vision-Exp \
--tensor-parallel-size 4 \
--enable-expert-offload \
--expert-offload-mem 160 \
--reasoning-parser deepseek_v4 \
--reasoning-config '{"reasoning_parser":"deepseek_v4","reasoning_start_str":"","reasoning_end_str":""}' \
--tool-call-parser deepseek_v4 \
--enable-auto-tool-choice \
--max-num-seqs 8 \
--enable-prefix-caching \
--enable-chunked-prefill \
--max-num-batched-tokens 2048 \
--kv-cache-dtype fp8 \
--block-size 256 \
--max-model-len 256000 \
--gpu-memory-utilization 0.97 \
--mm-processor-cache-gb 4.0 \
--override-generation-config '{"max_tokens": 128000, "temperature": 1.0, "top_p": 0.95}' \
--speculative-config '{"method":"dspark","model":"/models/deepseek-ai/DeepSeek-V4-Flash-Vision-Exp","num_speculative_tokens":3,"draft_sample_method":"probabilistic","enable_adaptive_verification":false}' \
--compilation-config '{"cudagraph_capture_sizes": [4,8,12,16], "max_cudagraph_capture_size": 16}' \
--host 0.0.0.0 \
--port 8090
I am pretty dumb in this kinda stuff so please don't blame me for it.
My llama serve kept crashing with Cublas errors on first prompt with some models, but my VRAM usage was 3000MB/8k. I chatted with sonnet 5.5 for a bit and it told me it was a CUDA version issue (and made it work by not selecting any CUDA device)... I don't know if this makes sense, please tell me if it doesn't/what problem you think there is.
I realized that the best way was to delete llama.cpp (installed with curl install script) and install a clean CUDA 12 Ubuntu version.
\\- \*\*How do I properly uninstall llama.cpp (I don't wanna mess with Ollama files)?\*\*
\\- \*\*How do I install new version from .tar.gz archive without messing with system packages?\*\*
\\-Does my error diagnosis make sense to you? Would CUDA make generation actually faster? (I am getting 8tk/s with Qwen35B Q2 on 4060Ti 8GB due to no CUDA selected)
Edit: You guys saved me! Thanks! I had to install NVIDIA toolkit and switch to CUDA 13 llama.cpp tarball build
TL;DR: I'm a solo dev who wanted a simple, private way to have local LLMs watch my screen and do simple logging/notifying. After a year of building, I released v3.0.0 and my mom was able to use it for the first time and I wanted to say thank you!
Hey r/LocalLLaMA,
It is designed to monitor anything, some use cases:
It's a micro-agent framework controlled by an MCP (Agent which I call Observer). So you type in Observer what you want monitored, and it'll control the framework to monitor it.
The desktop app uses llama.cpp as an inference engine, the webapp uses transformers.js, and they both support your v1/chat/completions endpoints :DD
You can try it out in your browser with zero setup!... running gemma-4-e2b ONNX in the browser, crazy stuff! Thanks to Xenova/HuggingFace for transformers.js c:
You guys told me that the framework was cool, but it was very manual to setup agents/workflows. So I've spent the last year slowly making it more accessible so anyone from any technical background can use it.
Every couple of months I ask my mom to use the App. And for the first time she actually was able to setup a monitoring agent with a local LLM! Which makes me think the app is ready for general public adoption (wuuuu!).
I hope this makes local LLMs useful for everyone! Tutorial/Demo Which is the whole point of the project.
The core Observer AI platform is, and will always be, free and open-source. That's non-negotiable. The code is all on GitHub for you to use, fork, and inspect.
The line in the sand which I have is "if it's free for me, it should be free for the user", that won't change ever.
This project wouldn't exist without the inspiration I've drawn from this community. You are the people I'm building this for.
I'll be hanging out here all day to answer any and all questions. Thank you again for everything!
Cheers,
Roy
I ran Swift-1.5-Qwen3.8-27b-oQ8e-mtp through the llm-bench.io a few times today: oMLX on M5 Max 64 GB, thinking on at xhigh, 262k context window.
The big difference between Swift 1.5 and the base Qwen3.8 27B is how much it writes. Per full run (agent workflow, code generation, research, role play) Swift averages 51k generated tokens and Qwen3.8 averages 77k.
| Scenario|Swift 1.5|Qwen 3.8 27B|
|:-|:-|:-|
|Code generation|28.6k|40.1k|
|Research|11.1k|21.2k|
|Agent workflow|7.5k|11.9k|
|Role play|4.0k|3.8k|
The full benchmark run duration: avg. 24 min for Swift 1.5, avg. 38 min for base Qwen 3.8 27B
Everything else is about equal:
Still 3/4 of what Swift generates is reasoning, it just does less than Qwen 3.8 27B. The output is still very usable. I'll for sure give it a try to be my daily driver for a few day.
Runs Swift 1.5:
Runs Qwen 3.8 27B:
Hi, so I was working on building a local server for an older online game client a few years ago, and the amount of data I had to synthesize was intense. I ended up shelving the project. I had managed to build a basic login server, sorted out some crypt stuff, but it was just way too slow of progress for me. I've got some decrypted packets and lots of data to work with, so the LLM is not going to be forced to do this entirely blind.
I restarted it recently with Fable, and as expected, it has been a huge help. However, I'm consistently hitting the safety guardrails now that I am further into the project. I am wondering what you all would suggest for a local setup? I have 16GB of VRAM (RTX 4060 Ti) and 128GB of DDR4. If that is entirely insufficient, I might be willing to pay for hosting a more powerful local model. I'm fine with it being slow and chugging along all day and night - I am primarily concerned with it actually figuring out the client functions, and doing things as accurately as possible.
I appreciate any insight into specific models, harnesses, and other stuff I should be looking into. Thanks!
Would increasing the number of ram channels (eg dual channel to quad channel) help improve inference speeds with freetoken in hybrid gpu/ram setups or not? Or is increasing the vram capacity and vram speed the only optjon
So Codacus created Jev mode for Lllama.cpp , and I thought Why not extend this concept further and ask questions about images and have the constrained answer be an image selection? So i spun up an agent and added image support and a harness. Now you can use images as your prompt without the decode step, no caption pause, just a decision based on an image or group of images. Ask the same question for a batch of images, like clasification. OR hand 1 context a whole group of images and ask it to pick on. like which of these 20 images has a ruber duck?
https://github.com/thecodacus/llama.cpp/pull/17
Youtube explainer using Codacus own RenderDiv framework to create the video.
https://youtu.be/Xuw3la2zVpg?si=rtSAydhuF9n3SYWV
I’m building a Polish General purpose legal Model that drafts documents, answers questions using legal sources, and has enough coding ability to handle some automation. The workflow is very tool-heavy:
Question → many sequential tool calls → final answer/document
Think Claude Code/Codex-style execution, but for legal workflows. Reliable tool selection, correct arguments, and recovering from errors matter as much as writing a good final answer.
I’ve had decent results with a dense 27B Qwen 3.8 custom made fine-tune for complex legal document summarization and classification. I’m already familiar with the smaller Qwen A3B and Gemma options. What interests me is the tier above those: 50B+ total-parameter MoEs with a relatively small active parameter count.
The question is, the small dense ones are great, but slow for agentic stuff (afaik), and i wonder if theres some middle ground maybe 70-120B models that would be able to be fine-tuned for the law stuff but be MoE so the agentic ClaudeCode style inference would also be lightning fast, and also low-ish cost for fine-tuning and inference.
Basically: Does the larger-total/small-active MoE approach actually buy you meaningfully stronger reasoning and tool reliability while retaining low latency,and at what hardware cost?
I understand that small active parameter counts don’t mean small VRAM requirements: the weights still need to live somewhere, alongside context and serving overhead. I also don’t assume that more total parameters automatically means a better model. I’m interested in where that tradeoff works in practice.
There are three things I’m trying to pin down:
For context, fine-tuning would target Polish language, document conventions, and successful tool trajectories. The actual legal sources would remain in retrieval/tools rather than relying entirely on memorized law.
I’m not looking for someone to compile a model shortlist (althought would be nice, but i dont expect anyone to break their back over this).
I’m looking for pointers, and firsthand experience with this particular size/architecture tradeoff. A configuration like “model + quantization + GPU(s) + serving engine + context length + concurrency + measured latency,” along with whether you successfully fine-tuned it, would be much more useful than a leaderboard score.
Has moving from a \~30B model to a 50B+ low-active-parameter MoE actually improved your agent’s successful tasks per minute, or did the memory, interconnect, and training requirements erase the advantage? Thanks for reading
I read textbooks on my PC and ideally i want a program with a normal chat environment where i can just hammer in questions about what's currently on my screen. Example: I'm working on a PDF, underline or circle things... and then just type "what does this sentence mean?", you get it.
I do NOT want to manually screenshot, navigate to the folder, drag the picture into the environment and then also have to type the question. It should also naturally be aware that the conversation is about what's on the screen, so i don't have to steer it with "make a screenshot; use your vision capabilites" etc.
I tried multiple different MCP in LM Studio, none of them were great...or worked :/
Someone said AnythingLLM has this function but i couldn't find it.
Is there a good solution?
Thank you in advance! :)
What (and if) would you buy if you had $5k ... $10k ... $20k to spend on local AI?
So, I'm a contractor and a solo dev working on some products/saas/apps. Basically I usually run up to three cline/opencode sessions in the same time, long running software engineering tasks, often run out of 128k context, so 256k ctx is prefferable. Pretty much every day for multiple hours so I could burn quite a lot of $ daily on OpenRouter. Fortunately, since Qwen3.8 i rarely need to delegate to larger models like Kimi K3 or MiniMax M3.
Apart from code I often work on confidential documents so a local AI or an approved remote AI is a must.
My current AI setup is:
\- dev server with RTX3090 running Qwen3.8 UD Q4\_K\_XL with 128k ctx q8
\- lab server with 2x RTX3090 + 128gb ram running Qwen3.8 Flash Next with 256k ctx
Both machines are fine to run up to three sessions, one on dev and 2 concurrent 128k ctx tasks on the lab server. The performance i get from the 2x RTX3090 with Qwen3.8 Flash Next is close to a single DGX Spark GB10 (according to numbers).
Soon I may need to run more agents and work with other people so started thinking of an upgrade.
Does it make sense to invest in either a single GB10 machine or two and cluster them to get more space for more context and therefore more concurrent sessions? Would you consider other options?
Any feedback appreciated.
Another day, another refusal. Apparently asking opus "what are your thoughts on this?" is an attempt at a 'distillation attack'
Our precinct is hugging face, we work at breakneck speed, we're up against art thieves, code thieves, extortionists, we're on call around the clock. The people of LocalLLaMA -- our finetunes is our job (and we write our own emdashes thank you)
WHERE ARE THE DATASETS PEOPLE. What happened to us? We used to throw pies at Dario Altman and now, what, we're penniless, downtrodden, what happened?
Ornith-1.5-9B-DFlash pairs the Ornith-1.5-9B model with a DFlash draft model for speculative decoding.
https://huggingface.co/ornith-ai/Ornith-1.5-9B-DFlash
Ornith-1.5-397B-DFlash pairs the Ornith-1.5-397B model with a DFlash draft model for speculative decoding.
https://huggingface.co/ornith-ai/Ornith-1.5-397B-DFlash
Ornith-1.5-35B-A3B-DFlash pairs the Ornith-1.5-35B-A3B model with a DFlash draft model for speculative decoding.
#
https://huggingface.co/emperorofrome/Gmcoder
Beats Ornith 1.5 and Oxcoder on HumanEval+ Mini — and does it without the overthinking. It gets to the answer using 40–68% fewer tokens. Built as a finetuned merge.
Edit: Best in the world is just hype. The coder is comparable to Ornith 1.5 but more token efficient by about 50% on average up to 78% at times and can be faster.
I'm building a new rig to host the mission control application that should sit on its own node and I immediatly thought of a cheap single socket ddr4 solution,
does anyone have some suggestions which bundle is cheapest or gives best value for price...?
\-no need for gpus
\-no need for much ram (8gb ram sticks)
\-many threads and max core speed would be important (maybe if it doesnt spike the price)
Not an LLM, but the ternary findings should carry over, and we hadn't seen Sherry-style 3:4 weights run in a browser before. Disclosure: this is our work at Precisit, everything is MIT.
What it is
|Model|File size|vs depth-4 bot|vs depth-6 bot|
|:-|:-|:-|:-|
|dense (fp32)|29.7 MB|0.92|0.89|
|T34, trained ternary|1.59 MB|0.93|0.91|
|T34, fine-tuned from dense|1.59 MB|0.89|0.9|
|T34, converted after training|1.59 MB|0.13|0.11|
|Base243 (TQ1\_0 style), trained|1.93 MB|0.89|0.88|
200 games each, both sides play a random move 5% of the time, a win counts 1 and a draw ½.
What we learned
Play it:
https://precisit.github.io/onepass-web/demo/c4-size/
Code, models, every result:
https://github.com/precisit/onepass-webgpu-ternary
The write-up:
https://precisit.com/en/blog/onepass-c4-size/
Has anyone gotten post-training 3:4 conversion to work on models, or does it need training?
My ninfer-ext fork of the famous ninfer now supports exl3 , and have released two models as well. https://huggingface.co/jabbatheduck/ninfer-ext-models
What FOSS/self-hosted tools or AI tools have you found genuinely helpful for students? Also interested in anything non-AI that has helped with studying, notes, organization, research, etc. Curious what you guys actually use or found useful. (written fully by human)
Trying to see if I can use a local model to generate game assets, or even 3D printer models. Any suggestions? Something that would fit in 64GB of ram, speed is not an issue. Just need it to work well.
I'm running Qwen 3.8 27B Q4 XS with \~30 t/s decode (no MTP) at 100k context with q8/q5 KV cache on a 16GB AMD GPU. I wanted to share my setup since many believe Qwen 3.8 27B to be infeasible on 16GB VRAM.
Build llama.cpp with Vulkan:
cmake -B build -DGGML_VULKAN=ON && cmake --build build --config Release -j
Grab Qwen3.8-27B-UD-IQ4_XS.gguf and mmproj-F16.gguf from unsloth/Qwen3.8-27B-GGUF, then:
llama-server \
--model Qwen3.8-27B-UD-IQ4_XS.gguf \
--mmproj mmproj-F16.gguf --no-mmproj-offload --image-max-tokens 2400 \
--n-gpu-layers 999 --ctx-size 100096 --parallel 1 --no-kv-unified \
--batch-size 2048 --ubatch-size 512 \
--flash-attn on --cache-type-k q8_0 --cache-type-v q5_1 \
--load-mode none --fit off \
--cache-ram 4096 --ctx-checkpoints 4 --checkpoint-min-step 8192 \
--no-context-shift --jinja --reasoning-format deepseek \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--threads 6 --threads-batch 6 --host 127.0.0.1 --port 8080
---
Edit: The process is documented here: https://zenodo.org/records/23088880. Feedback will be integrated into the upcoming Qwen 4 setup.