← r/LocalLLaMA
▲
0
 
8👁
r/LocalLLaMA · u/power97992 · 9d ago

Next year, the pro models will have 8-10 T parameters, who will have enough vram to run them?

Deepseek said they will release an 8 T model later and qwen said they will have a 10 T model and kimi will probably follow suit. The flash models will probably be around 1 -2 T parameters. Then only companies and corporations And cloud providers and rich people will be able to afford to run these pro models and fairly rich people for the flash models . At this rate, you would need 9 512 gb m5 ultras or 48 rtx 6000 pros to run A 4.4 bit 8T model with full context ? That is probably 153k for the ultras or 768k for the rtx pro Gpus plus probably another 100k for the other parts. I guess either use the cloud or people will use smaller models like qwen 5 27b in the future but most people won‘t be able To run the biggest models locally. In fact, most people will struggle to run a 4.4 bit 1 t flash model locally. It will cost 100-120usd/h just to host the mod in the cloud

1 0 0 10/3 06:29 10/5 11:49 UTC
scorecomments8 sightings
first seen 2026-10-03 06:29 UTClast seen 2026-10-05 11:49 UTCscore then 0score now 0gained 0sightings 8
open on reddit ↗ 💬 53