← r/LocalLLaMA
▲
218
+4
32👁
r/LocalLLaMA · u/Acceptable-Cycle4645 · 10d ago

Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models

post image

I started mapping the building blocks shared across all the models in audio.cpp. The result ended up being more interesting than I expected.

Qwen has become by far the most common language backbone in this collection: 32 audio model families use a Qwen-family architecture, and 20 of them use Qwen3 LLM specifically.

And it’s no longer just TTS. Qwen-based models now show up across speech synthesis, ASR/audio understanding, music generation, speech-to-speech, and even audio/video models.

The 2nd chart, Task × Technology Matrix, shows which build blocks power which types of audio models.

All figures here: https://github.com/0xShug0/audio.cpp/tree/main/assets/figure/

225 0 218 10/3 04:20 10/9 02:52 UTC
scorecomments32 sightings
first seen 2026-10-03 04:20 UTClast seen 2026-10-09 02:52 UTCscore then 214score now 218gained +4sightings 32
open on reddit ↗ 💬 24 (+1)