← r/LocalLLaMA
▲
1674
+7
34👁
r/LocalLLaMA · u/xenovatech · 22d ago

Ternary Bonsai 2 (27B) just released on Hugging Face. At <6GB in size, it can even run locally in-browser on WebGPU.

post image

The model is derived from Qwen3.8-27B, a 27B hybrid-attention causal language model (architecture unchanged), but uses ternary weights to shrink model size down to <6GB in size. According to the model card, it's 9x smaller than FP16 while retaining 98.2% of the intelligence.
\- Collection: https://huggingface.co/collections/prism-ml/bonsai-2
\- Demo: https://huggingface.co/spaces/webml-community/ternary-bonsai-2-webgpu-kernels

1681 0 1674 10/2 11:33 10/8 14:40 UTC
scorecomments34 sightings
first seen 2026-10-02 11:33 UTClast seen 2026-10-08 14:40 UTCscore then 1667score now 1674gained +7sightings 34
open on reddit ↗ 💬 344