← r/LocalLLaMA
▲
452
-2
30👁
r/LocalLLaMA · u/FullstackSensei · 31d ago

Qwen/Qwen-Drive-1.0-4B · Hugging Face

I don't think anyone posted about this here, but Qwen released a finetuned version of 3.5 4 for driving. The full Bf16 checkpoint is 9B.

This is a very interesting development of Chinese AI labs tackle self driving next with open weight models.

Edit: the HF repo links to the github repo, which in the citation links to a 40 page technical report. Here's the abstract:

We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model
for autonomous driving. Qwen-Drive-1.0 retains the architecture of the pretrained
vision-language model (VLM) and integrates 3D perception, visual question answering,
and motion planning within a unified framework. An external bird’s-eye-view (BEV)
perception head jointly performs 3D object detection, semantic occupancy prediction,
and BEV map segmentation. It serves as a probe of the 3D information accessible from
the shared representations and provides an explicit, inspectable interface to 3D scene
structure. A Planning Expert conditions on shared VLM representations to generate
future ego trajectories. A staged training recipe combines driving supervision with
general-purpose vision-language data to acquire driving-specific competence while
helping preserve broad visual understanding and instruction-following capabilities.
Experiments demonstrate strong 3D perception and driving scene understanding while
largely preserving general vision-language capability. Comprehensive evaluations across
open-loop, pseudo-closed-loop, and closed-loop settings further show highly competitive motion-planning performance.

460 0 452 10/2 17:28 10/7 22:40 UTC
scorecomments30 sightings
first seen 2026-10-02 17:28 UTClast seen 2026-10-07 22:40 UTCscore then 454score now 452gained -2sightings 30
open on reddit ↗ 💬 133