← r/LocalLLaMA
▲
206
-4
23👁
r/LocalLLaMA · u/Terminator857 · 18d ago

Deepseek training 2T and plans 8T model

Quote: DeepSeek is training a 2T-parameter model and plans to eventually build an 8T-parameter model.

https://x.com/wallstengine/status/2101982843656388644

Current DeepSeek models:

  1. Flash parameter count of 552 billion
  2. Pro: 1.6T (trillion) total parameters with 49B (billion) activated weights per token

Mythos / Fable is estimated to be 10T parameter count.

214 0 206 10/3 04:20 10/9 02:52 UTC
scorecomments23 sightings
first seen 2026-10-03 04:20 UTClast seen 2026-10-09 02:52 UTCscore then 210score now 206gained -4sightings 23
open on reddit ↗ 💬 140