Publication: VLM-based Captioning for Short BEV Driving Sequences
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
Access Restrictions
Abstract
Generative driving simulators have emerged as a promising alternative to real-world data collection for autonomous vehicle (AV) development, but they typically offer limited controllability at inference time. Recent work on language-conditioned scenario generation addresses this by allowing users to specify desired scene attributes through text. However, capturing how road structure and agent interactions evolve over time in driving sequences remains a limited area of research. This thesis develops a BEV scenario captioning approach tailored to the nuPlan dataset. We introduce a rule-based labeling pipeline that converts per-frame scene-graph data into structured target captions. We then use the resulting dataset of BEV sequences and temporally-grounded reference descriptions to fine-tune a VLM for scenario captioning.