Principal Investigator
- Seongjin Choi, Assistant Professor, Civil, Environmental and Geo-Engineering
Co-Investigators
-
Zirui Liu, Assistant Professor, Computer Science and Engineering
Summary
In this project, researchers aim to embed human-understandable physical and symbolic rules into the decision-making process of Vision-Language Models (VLMs) for autonomous driving, enabling the model to justify and articulate its decision-making process and consequent driving behavior in real-time. Conventional end-to-end autonomous driving models, i.e., models that learn a direct mapping from raw sensor input (camera, radar, LiDAR) to low-level control commands (steering, throttle, and braking), typically operate as highly complex black-box models. While such models have achieved strong empirical autonomous driving performance, their internal reasoning and decision-making processes are opaque, which makes it difficult for human users to interpret or build trust in the models. Moreover, safety constraints and traffic regulations are often incorporated into these models only as soft constraint penalties during training. As a result, there is no formal guarantee that the probability of generating an unsafe or illegal action is strictly zero at inference. As a result, without an additional safety verification layer, the model may still produce rare but high-risk behaviors, which also may be amplified under distributional shifts or adversarial conditions.