← r/LocalLLaMA
▲
60
+1
25👁
r/LocalLLaMA · u/ThomasAger · 20d ago

I enjoyed the daily HF papers today

Top 3 papers on HF Daily Paper are all unusually delightful and interesting reads for anyone on the leading edge of local LLMs, agent harness optimization, etc, felt like sharing.

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

https://huggingface.co/papers/2609.19969

Cross-layer KV reuse plus FP4 KV caching brings the global KV cache to 890 bytes per token, about a quarter of V4-Flash.

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

https://huggingface.co/papers/2609.20519

Auto-research loops that improve the agent harness, cutting token traffic by 44.7 to 49.0% at comparable performance.

An Empirical Study of Harness Design for Coding Agents

https://huggingface.co/papers/2609.20804

Varies planning, action space, and context management across 176 settings to see what each component actually contributes.

I'm still reading through, feel free to discuss

62 0 60 10/3 04:49 10/9 01:28 UTC
scorecomments25 sightings
first seen 2026-10-03 04:49 UTClast seen 2026-10-09 01:28 UTCscore then 59score now 60gained +1sightings 25
open on reddit ↗ 💬 8