Splash 1.1.0 released, GGUF quants support, MLX import and more
On my M5 Pro 64GB I can comfortably work in an agentic setup with the Qwen3.8 27B model in good quality (Unsloth UD-Q4\_K\_XL) at a decent speed of 50 t/s. Splash combines optimized kernels, excellent speculative decoding, a well-implemented prefix cache, and mixed-weight support in a single program. To me, this is a breakthrough in local inference on Apple Silicon. https://github.com/incoai/splash/releases/tag/1.1.0
scorecomments11 sightings
first seen 2026-10-03 06:31 UTClast seen 2026-10-05 05:54 UTCscore then 35score now 35gained 0sightings 11
open on reddit ↗
💬 11