Spatial Reasoning and Risk Assessment for Autonomous Vehicles on Consumer Electronics Platforms Using a Customized Vision–Language Model With Data Augmentation
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
Institute of Electrical and Electronics Engineers Inc.
Abstract
Accurate spatial reasoning and risk assessment from monocular video on consumer electronics platforms are prerequisites for safe decision-making in autonomous vehicles, yet general-purpose vision–language model (VLM) remains unreliable at lane-level localization and temporally grounded risk estimation. We present a deployment-oriented, spatially enhanced VLM that learns implicit 3D reconstructive information from monocular sequences via a reconstructive 3D tokenization pipeline and a spatial–visual fusion encoder built on MobileVLM. Concretely, the model fuses 3D reconstructive tokens with CLIP visual tokens through cross-attention to produce 3D-aware visual tokens, enabling lane-level localization and depth-aware reasoning about object orientation and inter-object relations. The unified decoder jointly generates scenario descriptions, object identities and positions, and converts spatial estimates into interpretable multi-level risk scores using a Time-to-Collision (TTC) mapping. Experimental results show that our system outperforms strong VLM baselines for multi-level risk classification on the nuScenes dataset. The generalization in CARLA and CoVLA further corroborates robust spatial reasoning and risk assessment. The findings indicate that coupling a lightweight VLM with monocular 3D cues via cross-attentional spatial–visual fusion yields accurate spatial localization and interpretable, transferable risk estimation that is robust under shift and practical for edge deployment in downstream tasks.
keywords: autonomous vehicles, consumer electronics platform, risk assessment, safety-critical scenario, Vision–language model,3D modeling, Autonomous vehicles, Behavioral research, Computer vision, Consumer electronic platform, electronic platforms, Language model, Localisation,Risk estimation.
Keywords
Citation
Song, Z., Yang, J., Fang, L., Ali, M. U., Kumar, G., Bashir, A. K., ... & Lee, S. W. (2026). Spatial Reasoning and Risk Assessment for Autonomous Vehicles on Consumer Electronics Platforms Using a Customized Vision–Language Model with Data Augmentation. IEEE Transactions on Consumer Electronics.
