← r/LocalLLaMA
▲
63
-1
32👁
r/LocalLLaMA · u/My_Unbiased_Opinion · 15d ago

PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192)

Just a quick PSA. llama.cpp does have prompt caching. if you are running large context lengths and have long multiturn projects, increasing -cram can provide you with massive speedups. There is a point where context lengths can get so large that 8192mb is not enough and the whole context needs to be re processed again on every turn. personally, I have found 20480 to work well with Qwen 27B 3.8 at 262K context.

the main downside is this uses more ram. vram usage doesnt increase.

68 0 63 10/3 04:49 10/9 05:27 UTC
scorecomments32 sightings
first seen 2026-10-03 04:49 UTClast seen 2026-10-09 05:27 UTCscore then 64score now 63gained -1sightings 32
open on reddit ↗ 💬 33