← r/LocalLLaMA
▲
2
 
1👁
r/LocalLLaMA · u/mikelau2026 · 6d ago

CASIA open-sources ZDTaichu5.0-9B: a 9B multimodal model built for 3D spatial and embodied reasoning

I keep seeing bigger vision models that crush OCR and chart QA, then fall apart the moment you ask where the free space is after a 90 degree turn, or which grasp point is actually reachable. ZDTaichu5.0-9B from the CAS Institute of Automation is interesting because it is only 9B, but the release is framed around physical-world spatial understanding: occlusion, cross-view 3D relations, and turning that into action plans. The official note says it took 8 of 9 firsts in its size band on spatial benchmarks, and they open-sourced the spatial data pipeline too. Not claiming it is the best across the board. Curious how it holds up if you have tried other ~10B vision models for robot planning. Source: https://ia.cas.cn/xwzx/cgzh/202609/t20260928_8287436.html

posted Sat, 03 Oct 2026 10:32:58 GMTseen 1 time
open on reddit ↗ 💬 0