This is mostly for setting up for expectation, since personally without LLM i could take 3 days (15 hours of active programming) to debug or implement a feature, but with Qwen 27B (even before Qwen 3.8) it take 4 hours. And yes 0.5 tok/s is human, not accounting of deletion and pausing, that's also the reason i am fine leaving overnight code base wide analysis or fin tech and deep research.
I am still waiting for Llama 5, because Muse Spark will be too big for me, or just something between Glimmer and Spark https://x.com/finkd/status/2095232032896946311
I have been trying to squeeze TTS stack down far enough to run in a $3 chip which has 512kb of SRAM without NPU. While trying to get to that milestone i built sanoTTS which has
- 11 voices, 6 languages
- params size ranging from 294k - 2.2m. For comparison we are 244x smaller than kokoro, 9000x smaller than voxtral TTS
- 1.5m model has a SCOREQ of 4.13 and UTMOS of 4.10
- 337kb for 294k model when quantized into int8
- can be run in website with web assembly npm install sanotts-web
- there is a recipe to follow so that you can extend to more languages, voice
I can tell you with confidence that this family release contains the smallest neural TTS model ever with around 2% WER on whisper.
Please check it out on : https://github.com/ampixa/sanoTTS
for live demo: https://tts.ampixa.com/sanoTTS
HF: https://huggingface.co/ampixa/sanoTTS
on SCOREQ sanoTTS-Amy(1.51m) is better than Inflect Nano(4.63m) and KittenTTS(15m) i.e
4.13 vs 3.81 vs 3.02
on esp32 microcontroller we are getting RTF of 0.225 which in plain terms means 4sec of audio is generated in 1sec
Happy to answer your queries.
Defined as AI exceeding human cognitive abilities. 20 years in prison. Plenty of local models already fall under that big of an umbrella in some capacities. This is why it's not enough to say that you could torrent open models so who cares what the politicians do. They want you to not have access to anything good and will put you in prison for it.
ChatGPT is down r/ChatGPT Claude is down r/ClaudeCode Grok is down r/grok my local llama.cpp works as always
more sizes (probably still uploading):
https://huggingface.co/IFM/K2-Horizon-32B-GGUF
https://huggingface.co/IFM/K2-Horizon-7B-GGUF
https://huggingface.co/IFM/K2-Horizon-3.7B-GGUF
https://huggingface.co/IFM/K2-Horizon-0.9B-GGUF
from IFM:
K2-Horizon-MoVA-36B-A4B is the sparse member of the K2-Horizon family: a Mixture-of-Experts model with Mixture-of-Values attention (MoVA) that stores 36B parameters and runs 4B per token. We have released the final checkpoint; intermediate checkpoints, along with the data and the training code, will be released.
collection: https://huggingface.co/collections/IFM/k2-horizon
Looking into the new Qwen architecture, I was curious if you could modify the Ngram PLE Table to make it work like a long-term knowledge database. It turns out that, with some limitations, you can.
I coded a small modification to llama.cpp to modify the table in-memory, allowing you to patch it with new data in real time. The PLE table is updated on every prompt, so now you can hot-swap parts of it without reloading the model.
The limitation is that it’s hard to control the output reliably, as the embeddings are injected early in the layers. However, with some techniques, you can influence the model’s output with simple modifications, as the example shows.
I created two repos:
There are some limitations in the project, as the PLE table needs to be memory-mapped into memory (this is the default in llama.cpp), and I have only tested it with q8 quantization, so you need quite a bit of memory to test this.
Can this be used as a new way of low-cost training? Perhaps. Its not easy at the current state but with simple modifications, I think you could easily create models with long-term instantaneously hot-swappable memory.
I liked the Nvidia that focused on just GPUs for gaming, not on the Nvidia of today which seem want power consolidation. Modelscope is another platform for those that simply want to know an alternative if things go south. However, time will tell what happens to huggingface after the deal is finalized Link: https://modelscope.cn/home, and https://modelscope.ai/home
ik\_llama.cpp merged qwen4exp MTP support yesterday (PR #2369, mine, reviewed and tested by four other people on their own hardware). It's on main now, no fork or patch needed. Posting since the last couple threads had people saying MTP for this model only exists as an unsloth fork PR... there's another path.
Flash-Next ships a 2.6B MTP head that the public converters were dropping. With it loaded the model drafts its own next tokens and then verifies them, so output is identical to running without it. On code I get 93-99% draft acceptance, prose more like 60-65%.
Numbers, decode tok/s, no MTP → MTP. My 5090 + 128GB DDR5, experts on CPU: 45 → 90 on coding traffic with ngram-mod chained in front. treo on an RTX Pro 6000: 85 → 113 on code, but story went 83 → 59, so not a free win on prose yet. joelfarthing on a 12GB 4070: 9.5 → 12.5 on code at n\_max=1. Caveats: single slot for now (-np 1), and --jinja lowers acceptance because the template turns thinking on by default and reasoning text drafts like prose.
Stock CUDA build, then:
llama-server -m Qwen3.8-Flash-Next-MXFP4-ngramQ8-NextN.gguf -ngl 999 -ncmoe 38 -fa 1 -c 196608 -ub 512 -ctk q8\_0 -ctv q8\_0 -np 1 -t 24 -tb 32 --jinja --spec-type ngram-mod:n\_min=4 --spec-type mtp:n\_max=4 --spec-ckpt-mode gpu-fallback -rtr -muge
Already have an unsloth or other quant? The separate head route works on the same code, no re-pull: -md <head>.gguf --spec-type mtp:n\_max=4. dzannotti's and ji-farthing's heads were both tested during review. Haven't tried unsloth's "shared" shards yet, different layout.
PR: https://github.com/ikawrakow/ik\_llama.cpp/pull/2369
My integrated-head MXFP4 files: https://huggingface.co/jamesrogers/Qwen3.8-Flash-Next-MTP-MXFP4-GGUF
ji-farthing's ik\_llama KT quants + head: https://huggingface.co/ji-farthing/Qwen3.8-Flash-Next-ik-llama-GGUF
Curious what you measure, especially anything AMD!!
EDIT: Forgot to mention that multi-GPU has not been worked into this, just single GPU for now; getting multi setups addressed is on the to-do list and anyone with setups to help test would be great, so please DM me if that’s you!
I want to share a short paper just published exploring a simple but surprisingly effective optimization for sparse MoE reasoning models.
The idea: Instead of retraining anything, we just tweak the router at runtime. Specifically, we expand the expert selection budget (N≥K*N*≥*K*) only in the late transformer layers, with a linear decay factor applied to the extra experts. Early layers stay untouched. So Qwen 3.6 35B A3B becomes Qwen 3.6 35B A4B+ !
What we found — "Succinct Convergence":
When you give the model more expert capacity at the decision-critical final layers, it stops rambling. It reaches the same correct answer via significantly shorter reasoning trajectories.
Results on full MMLU-Pro (714 questions, Qwen3.6-35B-A3B):
Links:
there you can also take a look to my github repo (with beta version code) and the detailed json results of MMLU-Pro benchmark.
In the future i hope i can make same experimentation with a larger model like DeepSeek V4 Flash Q2.0
I'm a Non-native english speaker, part of this post was generated , for translation reason with the help of AI.
TimesFM-3 is the third generation of Google Research's zero-shot forecasting model, and the main change from 2.5 is that it handles multivariate inputs natively instead of being limited to a single series' own history. It supports multiple simultaneous targets, past-only covariates, and past-future covariates (things like holidays or planned promotions where future values are known), all without fine-tuning.
Architecturally it's a decoder-only transformer with 20 layers at model dim 1280 and 16 heads, patching 32 contiguous time steps per token, and alternating two attention types per layer: causal attention across time within a series, and full attention across series at a given time step. Forecasts are generated in one forward pass rather than autoregressively — the model appends masked placeholder tokens for the whole horizon and fills them in simultaneously, with past-future covariates left unmasked so their known values stay visible. It outputs 9 quantiles (10th–90th percentile) per target per horizon step.
Pretraining used GiftEvalPretrain (minus fev-bench overlaps), Wikipedia pageviews through Nov 2023, Google Trends queries through end of 2022, plus synthetic data, totaling over 1 trillion time points. Google reports best average rank on Gift-Eval, FEV-Bench, and Time against Chronos-2, Toto 2.0, and TimesFM-2.5, and claims the univariate-only mode already matches or beats those baselines before covariates are added.
Worth flagging: the weights are under the TimesFM Non-Commercial License v1.0, so this isn't a drop-in for production use the way some other releases are. PyTorch weights are on Hugging Face and GitHub now; BigQuery integration is listed as coming later.
I had to swing down to my local Microcenter yesterday and while I was browsing around the store I noticed something odd... Inventory. They must have had a few dozen 5090's on the shelf in various configurations/board partners (for comparison, the last time I was there a few months ago they had 1 available for purchase and it was a AIO liquid cooled model that was absolutely off the charts expensive). They also had a few prebuilts on the floor with 5090's in them. Granted, this is one market one store, but.. IDK, perhaps some hopium.... But for anyone who wants a 5090, Microcenter in Charlotte has a bunch of them in the mid 4K range for price. Yes, that price is ridiculous, I know.
They also had 2 Pro 6000's 96GB in the store, on "sale" for 14K a pop. In case anyone is looking to spend used car money on a card. ;) I'd never seen a 96GB 6000 at my local store before available for sale.
For a few days I've been working on creating a custom local-only harness for some work related research using Codex / GPT 5.6 Sol and the model feels not only dumber than usual, but straight up counter productive. It keeps adding unnecessary guardrails for the local agents, removes tools that I clearly specified I want them to have and always drifts from the original requirements. I need to ask it to change things multiple times, which ends up on some over-complicated final product.
This is not the first time either, for months I've been avoiding asking frontier llms for local AI advice as it always seems to be bad, obsolete, or clueless even with internet search. Sometimes it still recommends me Qwen3-Coder-Next for my set up when it's clearly an obsolete model. I'm pretty sure I'm not the only one either as I've heard from other people.
What have you been your experiences on this?
124B total parameters, 5.1B activated parameters, and a 256K context window
I would be really curious about this especially on unified memory devices.
https://preview.redd.it/tiiyv2u76bnh1.png?width=3980&format=png&auto=… x.com/jukan05/status/2095353082309972273 Will it more than double again next year and give us DRAM relief for our local llama builds? Update: Misleading because this is by revenue, not by DRAM volume.
I've been building Arcon around Qwen3-4B + LoRA. Instead of just making it a chatbot, I'm experimenting with persistent memory, personality/mood, internal state, tools, and eventually having it process things before replying.
I'm curious what people who've built local agents think - how far can you realistically push a small model with good architecture around it?
I put the whole thing on GitHub if anyone wants to poke around, roast the architecture, or tell me what I'm doing wrong, stars are always appreciated!
Ling-3.0-flash-Fin is the first finance-enhanced model in the Ant Ling family. Developed by Ant Group with leading financial institutions and domain experts, it extends Ling 3.0 flash through continued training on high-quality financial data.
With 124B total parameters, 5.1B activated parameters, and a 256K context window, the model combines financial expertise with efficient inference for long-horizon agent workflows