Next year, the pro models will have 8-10 T parameters, who will have enough vram to run them?
Deepseek said they will release an 8 T model later and qwen said they will have a 10 T model and kimi will probably follow suit. The flash models will probably be around 1 -2 T parameters. Then only companies and corporations And cloud providers and rich people will be able to afford to run these pro models and fairly rich people for the flash models . At this rate, you would need 9 512 gb m5 ultras or 48 rtx 6000 pros to run A 4.4 bit 8T model with full context ? That is probably 153k for the ultras or 768k for the rtx pro Gpus plus probably another 100k for the other parts. I guess either use the cloud or people will use smaller models like qwen 5 27b in the future but most people won‘t be able To run the biggest models locally. In fact, most people will struggle to run a 4.4 bit 1 t flash model locally. It will cost 100-120usd/h just to host the mod in the cloud