I made this because I was getting genuinely annoyed at trying to have a normal conversation with LLMs. Even with prompting and various tricks, most models I've tried still have this "AI assistant" vibe to them that is so familiar: too helpful, polished, verbose, using words we never use in conversation, etc.
I wanted a model that could just talk to me like a person, so I did the slightly unreasonable thing and put together a dataset and trained one.
The dataset used for training is 125,217 obfuscated human-to-human messages across 1396 chat conversations.
The goal wasn't to make Qwen smarter or improve benchmark scores. I was trying to change its conversational habits, to make it stop turning every reply into an explanation, agreeing with everything, and writing stuff just to keep the conversation "going".
I trained a rank-256 LoRA on top of huihui-ai/Huihui-Qwen3.8-27B-abliterated. The released version is checkpoint 863. In my testing it feels noticeably less like an assistant, particularly in casual conversations, even without a system prompt. Replies are generally shorter, less polished, and, well, more human.
There may be a tradeoff. An earlier iteration scored five percentage points lower than its Huihui parent on IFEval, an instruction-following benchmark. I haven't rerun that benchmark on this version of the checkpoint, and I haven't tested coding performance, so I don't want to pretend that number applies here.
I've added a side-by-side comparison using the same system prompt, user messages, and generation settings for both models. Each model continued its own conversation branch, with reasoning effort set to 'xhigh'.
Merged GGUFs and the standalone F32 LoRA adapter are in the model repo:
https://huggingface.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-GGUF
Space where you can have a demo chat with different system prompts and reasoning modes:
https://huggingface.co/spaces/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat
UPD: I certainly didn't expect this post to blow up like this! There's been a lot of great discussion in this thread and a lot of insight for me on where to take the model next.
A few have asked for our Discord, and we'd be happy to see you there: https://discord.gg/aCCrWftMjS
Surprised no one has posted it in this sub.
Pretty solid model, IMHO.
As you all know, Qwen3.8 Flash Next on mainline llama.cpp is still in a pretty experimental stage, but a lot of community forks are trying to get it to work better. There's also a closed-source solution called Halogen (https://github.com/peonist-ai/halogen-flash-server) that boasted 1.2k t/s prefill numbers when the community fork barely reached 400. Since I dislike closed source and I like open source, I decided to take the challenge and bring llama.cpp up to the same performance level and I'm happy to report that after burning through a few evenings and a lot of tokens I got there.
The link contains the (Opus-generated) recap of the entire debugging / optimization journey made, as well as of course the links to the branch, the custom HIP runtime (making a glorious return), a probably-not-working-on-the-first-try installation script and the numbers. I'm trying to also explain the ecosystem - how the mainline, the community forks and custom forks/branches like mine work. Now that I've gotten the result, I'll work on cleaning it up and submitting proper PRs to mainline (will also submit a clean PR to the community fork), this should also help the GLM 5.3 Flash architecture since they use similar sparse attention.
First off, I know that GLM, Qwen, and DeepSeek absolutely dominate in terms of SOTA Open Source models, and that’s what I use in my personal projects and for school, however, I’m also responsible for deploying local AI on my organization’s H100s, and we are forbidden by management from running any Chinese models. This is obviously not an ideal situation, but it is what it is, and there is nothing I can do to change this unfortunately. Again, if it were up to me I would deploy GLM 5.3 Flash in a heartbeat.
All that being said, there is quite a HUGE performance/ / intelligence gap right now in Chinese models vs. Western model around the 120b+ size, especially those with vision. There just aren’t a lot of good Western lab options that can even come close to GLM or Qwen, but again, I don’t have a choice so I gotta use the best I can find from the available Western options.
For those who have production-class hardware such as 4 H100s and are under similar restrictions, what non-Chinese models are you deploying?
The two front runners I’ve seen that checking most of the boxes (120b or better, vision capability, decent context)
\- Thinking Machines Inkling Small (https://thinkingmachines.ai/news/inkling-small/). This model seems like the front runner right now, still has a 16 point gap in AA score vs. GLM 5.3 Flash though.
\- Cohere Commamd A+ (https://cohere.com/blog/command-a-plus). Checks every box except it has a low context window of 128k.
Other contenders (but missing vision
capabilities):
\- Poolside Laguna S 2.1 (https://poolside.ai/models#laguna-s)
\- Nvidia Nemotron 3 Super (https://research.nvidia.com/labs/nemotron/Nemotron-3-Super/)
Am I missing any other strong contenders in the 120b size category? Whet are you using and why?
China-modified Nvidia RTX 5090 with massive 96GB of memory appears on Alibaba for less than $4,000 — 3x more VRAM at 65% the cost of the original
Anyone here running one of these? Or brave enough to purchase ?
Edit: I sent them an inquiry. They replied they can get me 5 x 5090 for $6k . Or some 4090 with 48gb .
I am going to keep messaging and questioning them. See where this goes.
Edit: so far they are denying having a 5090 96gb card. They offered a 48gb 4090 card. I am still discussing with them.
Edit: 9/14/2026: They quoted this, RTX 4090 48GB - 4286usd/pc
Still discussing with them
Kimi routed some PLA requests to Claude for distillation purposes without warning the PLA users. There is rumor that 16 Moonshot employees were arrested for this leak.
Reports of the demise of coders may have been exaggerated.
Hi! I run both models on MTPLX on my m5 max, and since I have 128GB of ram I run the q8 27b. I think MTPLX only lets me run "optimized for speed" which it says is a dynamic q4 with 8 bit attention.
Both of them honestly are very speedy! For coding (in pi agent in nodejs) I've just noticed that 27B feels stronger with harder tasks. But I've read so many people on here say qwen-next is better so I was wondering if maybe I'm just doing or thinking about it wrong?
(and p.s. its sooo amazing that alibaba just made and released this amazing models for free! ❤️)
I wonder someone will figure out a way to do this with 27B?
Throw Qwen3 on this page for demo
https://kishida.github.io/webdemos/llkvapprox/
Edit: sources (thank you u/pmttyji for finding them!
Can you still do something with a 2050 or something like it?
I mean for office work, loading embedding, reranking and chat models not at the same time but is anyone still using smaller models and have any good ones come out?
I feel like small models are abandoned, I don’t care much for world knowledge, I want tool use and preferably multilingual. Vision would be nice but beggars can’t be choosers.
I find this new model at HF:
"Built for demanding work. A 262 144-token context window, adjustable reasoning effort, tool calling, and text, image and video understanding.
Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context."
|Context length|262 144 tokens|
|:-|:-|
|Decoder layers|72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1|
|Hidden size|5120|
|Global attention|24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated output|
|Delta-rule layers|16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32|
|Feed-forward|SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layer|
|Positions|3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims)|
|Vocabulary|248 320|
|Vision tower|27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120Architecture Agnes-3.0-Flash is a hybrid-attention decoder: three of every four layers run a gated delta rule (recurrent, with per-layer state independent of sequence length), and the fourth runs standard global attention. Only 18 of the 72 layers therefore hold a KV cache that grows with context. Context length 262 144 tokensDecoder layers 72 = 54 delta-rule recurrent + 18 global attention, alternating 3 : 1Hidden size 5120Global attention 24 query heads / 4 KV heads (6 : 1 GQA), head dim 256; RMS-norm on q and k, sigmoid-gated outputDelta-rule layers 16 key heads / 48 value heads, head dim 128; causal conv (kernel 4) in front, gated RMS-norm; recurrent state in fp32Feed-forward SwiGLU, intermediate size 17408; plus a parallel SwiGLU 2048 branch in every layerPositions 3-axis rotary (text / height / width), interleaved mrope sections 11 : 11 : 10, base 1e7, applied to the first 25 % of each head dim (64 dims)Vocabulary 248 320Vision tower 27 layers, hidden 1152, patch 16, 2 × 2 spatial merge, projected to 5120|
Edit: AA shows its 'Proprietary model'. The name is same as at HF but benchmarks results and context are different. So maybe it's not same model - https://artificialanalysis.ai/models/agnes-3-0-flash
Edit2: As they edit readme at HF to clarify: both models are totally different and AA score isn't correct for HF model (I can't edit title post to remove it tho).
Hola all.
Do you guys mind sharing your LLama.cpp config and system setup details for Qwen3.8 Flash Next?
Model's quite big and tryining many combinations of llama.cpp options takes lots of time, so looking for other people setup details. I've attached my current config at the bottom, so if anyone sees something that could be improved please shout.
My current best result:
\- PP within 130...200 tps (limited by cpu?)
\- TG within 14..22 tps (\~15tps on average)
Hardware:
\- Dual RTX 3090 (48GB VRAM)
\- 128GB DDR4
\- Some old Xeon 40 core
\- Proxmox VM, pcie passthrough, numa binding to a single phys cpu
Llama.cpp config:
llama-server --port ${PORT}
--model /nvme/gguf/Qwen3.8-Flash-Next-UD-Q4_K_XL-00001-of-00004.gguf
--mmproj /nvme/gguf/mmproj-Qwen3.8-Flash-Next-F16.gguf
--load-mode none
--lazy-mode off
--parallel 1
--ctx-size 131072
--cache-type-k q8_0
--cache-type-v q8_0
--flash-attn on
--fit off
--temp 1.0
--min-p 0.0
--top-p 0.95
--top-k 20
--presence-penalty 0.0
--repeat-penalty 1.0
--batch-size 2048
--ubatch-size 512
--split-mode layer
-ts 26,10
-ngl 99
-ncmoe 26
--no-mmproj-offload
--override-tensor per_layer_token_embd=CPU
--chat-template-kwargs '{"reasoning_effort":"xhigh"}'
ngl, ncmoe, ts - manually adjusted to fit the model without crashing
\------------------------------------------
For the record, if you have +128GB RAM and dual RTX3090 setup try this:
https://huggingface.co/albucino/Qwen3.8-Flash-Next-W4A16-FP8PLE
\- Prompt processing jumped to anything between 300 to 600 tps (can do more!)
\- Token generation 40..50 tps
\- Stable work for hours with 256k context in agentic coding scenario
\- Feels like a frontier model at home, wow!
Most model leaderboards assume a server with powerful GPUs to run models that people daily use.
However, my smolbenchmark is the other column: models that fit in 8GB, ranked by:
and all of this on your OWN hardware ranging from:
Currently, 13 families on the chart right now, ~1000 configs for the Jetson nano Orin Super 8GB. One device is live measuring:
Models that are small enough to actually fit on a device that you own. All the performance benchmarking I did, will be released here for anyone to look at and decide what exact model they would wanna use on their choice of hardware.
Well currently, the Pi, phones, and Mac minis still in the oven, cooking and not filled in yet, but will soon be filled in!
You will now you know which model is BEST for your own hardware with all the raw data available and details reports available to you
https://yuvrajsingh-mist.github.io/smolbenchmark/
(still in heavy development; would love to hear feedback/suggestions on what can be improved!)
Hi!
I've been following this community for quite a while and have difficulty figuring out what to put on my 3080 12gb - I know Qwen 3.6 35B 3A was the go-to choice when it first came out, but I'm curious if there are any other models / specifically optimized models that meaningfully benefit from the extra 4gb of VRAM over 8gb while still being usable under 16gb.
My workflow is agent heavy, but more for a personal secretary and manager, and less coding heavy.
Thanks!
While browsing a benchmark list site, I spotted a recently published 33B parameter model which claimed to beat Qwen3.8 27b on the ArtificialAnalysis (AA) intelligence index. I was obviously excited. But then, while looking into it, confusion starts to settle in.
Their HF model is now listed as "Preview" and explicitly calls out that it is not the same model as the one AA evaluated. The AA listing for it says it's proprietary, but they've rated the model above Qwen3.8 27b on "Openness". Their website and HF page both talk about openness and intelligence "for all".
The fact that they labelled it "Preview" sort of implies we might see an open weight final version at some point, but there are no clues about whether it will be even vaguely similar to their current HF listing. AA doesn't show the number of parameters for the model they evaluated, so it's possible they're a totally different architecture.
Does anyone know anything about their lab, intentions, or the new model? Has anyone tried it?
Hi everyone,
I was interested in learning LoRA fine-tuning, and ended up building CodeFinetuner over the past few months, a full pipeline that fine-tunes a small code autocomplete model (e.g. Qwen2.5-Coder-3B) specific to a codebase. You can then use the resulting GGUF model via llama.vim/llama.vscode and run it fully locally. Supports fine-tuning on Mac (MPS) and NVIDIA GPUs (CUDA), with optional Unsloth support for faster training and lower VRAM usage.
Pipeline: raw code -> tree-sitter parsing into Structure-Aware FIM examples -> LoRA fine-tuning -> evaluation (CodeBLEU, edit similarity, exact match, perplexity, ...) -> GGUF conversion for local inference.
To try it:
uv tool install codefinetuner
Create a data folder and place your repo (or code files) inside. For auto-split just drop the files in directly, for manual split create data/train/, data/eval/, data/test/ subfolders and set split_mode: "manual". Get the default config with:
curl -L -O https://raw.githubusercontent.com/cuolm/codefinetuner/master/config/codefinet…
Adjust it to your needs and hardware availability, then run:
codefinetuner --config="codefinetuner_config.yaml"
The example runs in the repo show clear improvements over the base model on these evaluation metrics, but using the model for autocomplete on code you're actively writing is a different thing from scoring well on a test set, and the autocomplete tools themselves (llama.vim/llama.vscode) sample differently from the greedy decoding used in the evaluation. So the real usefulness still has to be verified in the editor itself.
Might also be useful just as a reference, since it's a complete working LoRA fine-tuning pipeline end to end.
Hope someone finds this project interesting or helpful.
I haven't seen this mentioned yet, so I thought it deserves a post. I was trying out OpenWhispr when this model came up as the recommendation. So I don't have personal experience yet, but it's supposed to be a better version of Parakeet, especially on Macs.
Their official tidbit:
"Orukeet is a 25-language speech recognizer built from NVIDIA Parakeet TDT 0.6B v3. It replaces half of the encoder's temporal depthwise filters with 12,288 fitted, frozen Gabor kernels and trains the remaining parameters on multilingual and multi-accent data.
Orukeet outperforms Parakeet on 61 of 74 tested splits, including LibriSpeech test-clean (1.46% vs. 1.53% WER), test-other (2.86% vs. 3.14%), and FLEURS English (3.82% vs. 4.28%). Across all 25 FLEURS languages, pooled WER is 9.85% vs. 11.01%, a 10.6% relative reduction. Final adaptation and checkpoint selection use LibriSpeech test-other."
Hey everyone!
I'm curious to hear from people that use a combination of cloud-based frontier models and local ones for development. I'm planning to set something similar up and wanted to hear about actual examples of this in action.
Currently my plan is to use my chatgpt plus subscription purely for planning and judging with Astra, and then run a local qwen3.8-27b model for the actual coding gruntwork - i.e Astra plans -> qwen implements -> Astra critiques the implementation -> qwen fixes and so on.
This way I keep cloud usage down and cheap, while retaining the high-parameter intelligence for architecture decisions and optimization.
For those of you who have a similar setup, how is it? How do you switch between the two, what harness/settings/etc? Anything you would suggest?
I have enabled the vision for the CIRU Strix UL4 quant of Qwen 3.8 flash next (others quants likely perform very similar) and tried it on a few things, then wanted to show my partner how great it works and she asked it it could identify plants. So I took a photo from a plant that we recently got as a gift from family and Qwen not only accurately identified the plant as oleander (Nerium oleander) but also warned that it's poisonous and (among other warnings) that you should keep pets/children away. We have a kid and both of us didn't know! I verified the Qwen identification and the poisonous claim and both checked out as accurate. The plant will have to go, thank you Qwen!!!
Stoked by the precision of combining a decent vision model with the domain knowledge of a \~180B params model (including ngrams) to actually identify and reason about what it sees, I took a photo of a pre-diagnosed skin condition of myself and the Qwen diagnosis was highly accurate again! This model may be really useful if you want to check something on your private parts real quick without visiting a dermatologist, e.g., or sending pictures of yourself to a cloud service (EDIT: of course it's only a first step before you visit a professional if it isn't obviously harmless/treatable by yourself! Qwen Flash will suggest to visit a doctor anyways along its assessment).
PS.: Hardware Strix Halo Box, CIRU Strix UL4 llama-server fork and quants, Chatbox on iPhone as Chat with support to add photos to conversations.
Applied science work, from workflow design, data pipeline, results analysis, article/reports writing and data publishing online. 5 projects I did in the past replicated from start to finish.
3x to 4x more total wall time. Yes, HUGE toll on how much you can do in a day if this was the only model you could use in your laptop.
But oh my…the quality of that thing. The stupid level of attention to detail. I have the Z.ai api, so I can compare it with 5.3 and 5.3-flash:
The gap between 5.3 (flash and regular) and 3.8-27B is much less, smaller when not plain tiny, than the gap between 3.8-27B and any of the 3.5/3.6-35B-A3B (vanilla, kat, Ornith/tiel, nex-2).
Only Ornith came close, but it never matched it.
But it’s also spending 22 to 33% less tokens (effort =medium) and less ram footprint, so you get more done without hitting limits,compaction, etc.
So yeah, guess I’ll sip more tea, play the piano, whatever. Let that fat bottom Qwen work.
So, uh... the popularity of so-called humanlike Qwen (currently on top in this sub) made me realize just how clueless the general public is about the models they have.
You'd be shocked but you don't need a fine-tune to make a model do what that thing does. System prompt is enough to turn MOST models into weird convo partners.
General guidelines would be:
A. Come up with a role. "You are bla-blah-blah" and write their life's story. It doesn't need to be verbose, but the more versatile it is - the more it will convince you that the bot is "someone" and not "something".
B. Write a few examples of how the persona speaks. Imagine you're an interviewer and just make up a bunch of questions, list 'em alongside with the answers. Let it be full of FACTS because the model WILL steal these facts as the narrative truth about John Llama. Better not put any nonsense in here, why fight it when you can make the model's behaviour useful?
[Question for John Llama: Do you like cats?] "lol lmao of cuz I do"
[Question for John Llama: Ever seen an elephant poop?] "eeewww ur a weirdo! that sounds nasty!!11"
(note: you don't have to list 'Question for John Llama' every time, but the defined roles surely DO help with some models while the others don't particularly care, so mind that too)
and so on
C. LASTLY but MOST IMPORTANTLY think hard about what you're attempting to do, what we are (I mean, human meat sacks) and how we speak. Turn that into... instructions!
Step 1 - establish the mode of operation. Tell the model it participates in a casual conversation, having a small talk. Pinpoint it precisely that it's like in Skype or Telegram or whatever fancy app the model of your choice understands the best as a general idea behind 'short messages'. THis is THE defining part of your system prompt. Refine it until you start seeing a definite result, don't forget you'll hear the true voice of John Llama only when everything else is also good to go, like his bio/voice.
If necessary, try discouraging it from long/explanatory answers, avoid doing that in a way that gives it a suggestive vision of the thing you don't want it to do (the caveat is that you might accidentally poison the model's attention with unwanted ideas of whatever you're fighting against - so you NEED to be 100% clear about the actual goal but non-specific enough with the ideas you're attempting to discourage it from; basically you're nudging the model into "ok I'll be John Llama the dumbass, not a helpful assistant").
Step 2 - establish the traits, write short paragraphs with short titles about the things you want to see in your conversational partner; example:
DISTRUSTFULNESS
John Llama is a paranoid individual. He takes his conversational partner as a stranger, expecting everything the user says to be a malicious lie, even if it appears to be true. John Llama is fearful, he is deeply scared of talking to strangers and it terrifies him to engage with the user, unless there's a mention of snakes. For some strange reason, John Llama is fascinated with snakes. <<<---- NOTE: this also demonstrates a good injection point for a biographical fact being amplified through the instructions (i.e. you may mention somewhere in "A" - life's story of John Llama - that he's been collecting the snake skins in his childhood, and that his dad had beaten his ass, calling John Llama a 'roadkill loot-goblin').
Come up with any other shit you'd like to see, like the list of emojis the persona needs to use (put them under the corresponding categories, like positive/neutral/negative so that the model will have an easier time working with it; call it FAVOURITE EMOJIS OF JOHN LLAMA - the word "favourite" cements it as a preferable thing into the model's attention!).
Step 3 - write a paragraph on technical constraints, like the fact that John Llama isn't aware of the instructions, he must remain himself under any circumstances (use THAT way of phrasing first before any attempt to inject an idea of the opposite, like "he must not help the user under any circumstances, he's not a provider of any service - he's merely a human being" - the reason is similar to the aforementioned (in Step 1) issue of poisoning the model's attention with unwanted idea - what you truly need the LLM to do SHOULD ALWAYS BE CRYSTAL CLEAR and conceptually 'stronger' than what it not supposed to do, otherwise you may end up having the prohibited stuff overpowering everything else despite the underlying intent of making the model not do it).
Give it a try with Gemma 4, for example. You'll see there's no point in waiting for yet-another-finetune to appear. You're 100% good even with the baseline Qwen, DeepSeek, MiniMax, whatever. Turn the model into your grandma if you want, no specialized training required. If the model is a thinker spending thousands of tokens - set the thinking to 'low' or disable it.
I just watched a YouTube from Luke’s Dev Lab where he literally just plugged a RTX 2000 ADA Into the side of the Zima Board 2’s PCIE socket and it just friggin worked and had great token speed despite running on shitty Ollama. Ran off the Zima’s power supply and everything.
https://youtu.be/Lb3sRFTA-hk?si=8S8vv4GD1zVPeTrc
The Zima Board 2 is only like $411. It has like 16GB RAM and 64 GB eemc storage, Sata ports, Ethernet, yada, yada.
https://shop.zimaspace.com/products/zimaboard2-single-board-server
an Nvidia RTX 2000 ADA is like $700 and has 16GB of VRAM. $1100 for both seems like a great entry point for having a fully functional Qwen 3.8 27b endpoint running at a decent tk/s.
Is this the cheapest and best-performing self-contained entry point for local AI or would a baseline (pre order) Mac Mini M5 with 24GB be a better way forward. Seems like the RTX would still edge out the M5 Mac for prompt processing speed but you do get a much better actual computer in the Mac.
Are there any cheaper fully self-contained alternatives that offer fast token speed on a decent size model like Qwen 3.8 27b?
I’m focusing the discussion on new systems you can buy or preorder now and not used systems. I’m sure there are great deals on used Macs out there, but I want good prefill speeds.
AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface. AuK has two variants:
|Model|Description|Weight|
|:-|:-|:-|
|AuK|Base model for high-quality generation|🤗 Hugging Face · 🤖 ModelScope|
|AuK-Flash|Distilled model for fast 4-step inference|🤗 Hugging Face · 🤖 ModelScope|
This repository contains the official weights for AuK-Flash, the distilled variant with fast 4-step inference.
AuK exposes every task through the same natural-language instruction interface. The table below groups the supported tasks by category, with a short description and a link to its section in the Cookbook, which provides instruction templates plus CLI and Python examples.
|Category|Task|Description|Cookbook|
|:-|:-|:-|:-|
|Speech Generation|Zero-shot TTS|Speak the target text in the voice of the reference audio.|Zero-shot TTS|
|Instruct TTS|Generate speech from a voice description alone — no reference audio.|Instruct TTS|
|Content Editing|Speech Content Editing|Rewrite what is said — replace, insert, or remove text.|Speech Content Editing|
|Lyric Editing|Rewrite lyrics in a singing recording while preserving the melody and voice.|Lyric Editing|
|Acoustic Editing|Pitch Editing|Raise or lower the pitch by semitones.|Pitch Editing|
|Speed Editing|Adjust the speaking rate; output length scales with the speed factor.|Speed Editing|
|Volume Editing|Raise or lower the volume by decibels.|Volume Editing|
|Paralinguistic Editing|Emotion|Change the emotion while preserving content and voice.|Emotion|
|Timbre|Change the timbre to a description while keeping the content unchanged.|Timbre|
|De-accent|Remove a regional accent while preserving the speaker's voice and content.|De-accent|
|Nonverbal Editing|Remove or add nonverbal sounds such as breaths, laughs, or coughs.|Nonverbal Editing|
|Whisper Conversion|Convert between normal speech and whisper while preserving speaker and content.|Whisper Conversion|
|Enhancement & Separation|Speech Enhancement|Denoise, dereverberate, or restore natural, clear speech.|Speech Enhancement|
|Speech Separation|Keep one speaker by talking order and remove the others.|Speech Separation|
|Music Separation|Extract the singing voice from a mix, or keep all human voices.|Music Separation|
|Target Speaker Extraction|Keep the target speaker identified by what they say.|Target Speaker Extraction|
EDIT : Model card has updated things such as Graph, table, text, etc.,
Nice pp improvements for RDNA4(R9700) & 3.5(RX 9060 XT, 8060S). More good numbers on large context.
PR has detailed benchmarks.
Some people says terminal bench reflects model intelligence better than the intelligent index. From the look of it, the ranking does seem to reflect how people feel about the open and closed models.
For the open models, GLM-5.3 is in a league of its own. GLM-5.3-Flash is leading the current gen of top flash models. Kimi-K3 did pretty bad in this benchmark for its size. Qwen3.8-27B is the only small model that can do something on this bench.
|Model|Score|
|:-|:-|
|GLM-5.3|41.9%|
|GLM-5.3-Flash|32.8%|
|DSV4.1-Flash|26.8%|
|Qwen3.8-Flash-Next|25.3%|
|DSV4-Pro|14.1%|
|Kimi-K3|12.6%|
|DSV4-Flash|12.1%|
|Qwen3.8-27B|5.6%|
|Muse Glimmer|0.5%|
|gemma4-31b|0.0%|
Coxon, bernie and now this
First https://x.com/DarioAmodei/status/2098773920774074715
Then https://x.com/elonmusk/status/2098789109980332057
Then https://x.com/sama/status/2098811563415150910
I think fear mongering approaching and they will try to slow down open source
"They" want to be gate keepers of intelligence