← r/LocalLLaMA
▲
123
-3
23👁
r/LocalLLaMA · u/pmttyji · 26d ago

Hoping for Optimized Smarter Upcoming Models .... Like DeepSeek-V4.1-Flash( KVCache + Engram) in Small/Medium/Big sizes

post image

It's still a dream for many folks to run medium size(30B range) models @ Q8 with Unquantized KVCache (256K Context) on their GPUs.

It would be awesome to have DeepSeek-V4.1-Flash's KVCache + Engram for all Upcoming models. Even for big models.

Engram - Heard that approximately 1/3-1/2 of Model size. Might come in different size range too. So 10-15 GB for 30B models.

Here few models with approximate numbers. Current models in Odd rows & Future/Fictional models in even rows(Bold). I just put 1GB for 256K context below though DeepSeek-V4.1-Flash takes only same 1GB for 1 million context.

|Model|Model Size|256K KVCache F16|MTP|Vision|Total GB|
|:-|:-|:-|:-|:-|:-|
|Qwen3.8-27B-Q8|29|16|1|1|47|
|Qwen4.0-27B-Q8|29|1|1|1|32|
|Qwen3.8-27B-Q4\_K\_M|17|16|1|1|35|
|Qwen4.0-27B-Q4\_K\_M|17|1|1|1|20|
|Muse-Glimmer-30B-Q8|30|16|1|1|48|
|Muse-Glimmer-2-30B-Q8|30|1|1|1|33|
|Gemma-4-31B|33|16|1|1|51|
|Gemma-5-31B|33|1|1|1|36|
|Qwen3.6-35B-A3B-Q4\_K\_M|23|6|1|1|31|
|Qwen4.0-35B-A3B-Q4\_K\_M|23|1|1|1|26|
|Gemma-4-26B-A4B-Q8|27|6|1|1|35|
|Gemma-5-26B-A4B-Q8|27|1|1|1|30|

Possibly there might be few more things(Please share those) to keep these number down. So I think 32GB VRAM is more than good enough for Upcoming (Optimized Smarter) Models. RAM is enough for Engram.

By above logic(based on Qwen4.0-27B), people could run Q4 of 54B models with same 32GB VRAM.

Maybe next year onwards, inventions could make 24GB enough for similar size models.

130 0 123 10/3 04:44 10/8 00:59 UTC
scorecomments23 sightings
first seen 2026-10-03 04:44 UTClast seen 2026-10-08 00:59 UTCscore then 126score now 123gained -3sightings 23
open on reddit ↗ 💬 35