Deepseek training 2T and plans 8T model
Quote: DeepSeek is training a 2T-parameter model and plans to eventually build an 8T-parameter model.
https://x.com/wallstengine/status/2101982843656388644
Current DeepSeek models:
- Flash parameter count of 552 billion
- Pro: 1.6T (trillion) total parameters with 49B (billion) activated weights per token
Mythos / Fable is estimated to be 10T parameter count.
scorecomments23 sightings
first seen 2026-10-03 04:20 UTClast seen 2026-10-09 02:52 UTCscore then 210score now 206gained -4sightings 23
open on reddit ↗
💬 140