Spatial Reasoning and Risk Assessment for Autonomous Vehicles on Consumer Electronics Platforms Using a Customized Vision–Language Model With Data Augmentation

dc.contributor.authorBashir, Ali Kashif
dc.contributor.authorKumar, Gyanendra
dc.contributor.authorYang, Jing
dc.contributor.authoretal..
dc.date.accessioned2026-07-08T05:21:49Z
dc.date.available2026-07-08T05:21:49Z
dc.date.issued2026
dc.description.abstractAccurate spatial reasoning and risk assessment from monocular video on consumer electronics platforms are prerequisites for safe decision-making in autonomous vehicles, yet general-purpose vision–language model (VLM) remains unreliable at lane-level localization and temporally grounded risk estimation. We present a deployment-oriented, spatially enhanced VLM that learns implicit 3D reconstructive information from monocular sequences via a reconstructive 3D tokenization pipeline and a spatial–visual fusion encoder built on MobileVLM. Concretely, the model fuses 3D reconstructive tokens with CLIP visual tokens through cross-attention to produce 3D-aware visual tokens, enabling lane-level localization and depth-aware reasoning about object orientation and inter-object relations. The unified decoder jointly generates scenario descriptions, object identities and positions, and converts spatial estimates into interpretable multi-level risk scores using a Time-to-Collision (TTC) mapping. Experimental results show that our system outperforms strong VLM baselines for multi-level risk classification on the nuScenes dataset. The generalization in CARLA and CoVLA further corroborates robust spatial reasoning and risk assessment. The findings indicate that coupling a lightweight VLM with monocular 3D cues via cross-attentional spatial–visual fusion yields accurate spatial localization and interpretable, transferable risk estimation that is robust under shift and practical for edge deployment in downstream tasks. keywords: autonomous vehicles, consumer electronics platform, risk assessment, safety-critical scenario, Vision–language model,3D modeling, Autonomous vehicles, Behavioral research, Computer vision, Consumer electronic platform, electronic platforms, Language model, Localisation,Risk estimation.
dc.description.sponsorshipUnderstanding complex spatial information, such as object positions, and conducting risk assessment based on this information constitute fundamental challenges [1] for multimodal scene understanding in autonomous vehicles (AVs). Such spatial reasoning capabilities are crucial for various downstream tasks, including motion prediction [2], [3] and planning [4]. While large-scale datasets [5], [6], [7] have enabled significant progress in per-object recognition tasks, such as detection [8], [9] and semantic segmentation [10], [11], comprehensive 3D scene understanding that captures inter-object relations and risk from RGB images remains underexplored. Consequently, as AVs are rapidly evolving into sophisticated consumer electronics platforms with tightly integrated sensing, computing, and human–machine interfaces, robust spatial reasoning becomes a prerequisite for reliable hazard anticipation and decision-making in complex, interactive, high-risk scenarios [12], [13]. However, large-scale, safety-critical annotations that capture lane-level relations and interactive risks are scarce and quickly exhausted as models evolve. This motivates data augmentation to programmatically expand task-aligned supervision, such as spatial question-answering, risk narrations, and visual variations, while keeping the deployment constraints [14], [15], [16] of consumer electronics platforms in mind.
dc.identifier.citationSong, Z., Yang, J., Fang, L., Ali, M. U., Kumar, G., Bashir, A. K., ... & Lee, S. W. (2026). Spatial Reasoning and Risk Assessment for Autonomous Vehicles on Consumer Electronics Platforms Using a Customized Vision–Language Model with Data Augmentation. IEEE Transactions on Consumer Electronics.
dc.identifier.doihttps://doi.org/10.1109/TCE.2026.3655159
dc.identifier.urihttps://repository.adu.ac.ae/handle/1/8341
dc.language.isoen
dc.publisherInstitute of Electrical and Electronics Engineers Inc.
dc.titleSpatial Reasoning and Risk Assessment for Autonomous Vehicles on Consumer Electronics Platforms Using a Customized Vision–Language Model With Data Augmentation
dc.typeArticle

Files

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.71 KB
Format:
Item-specific license agreed to upon submission
Description: