works on

From the 1 of 22 linked papers with an AI index.

activity
20242026
collaborators

22 papers

cs.CV2026

VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

Ruiqi Xian, Yuehan Xian, Jing Liang +2

The paper introduces VISA, a training-time method that uses a visual‑language model to audit and correct semantic labels of 3D voxel occupancy maps, improving object and rare‑class…

cs.RO2026

VLM-Based Advanced Rider Assistance System for Motorcycle Safety

Mohamed Elnoor, Francesca Baldini, Ananya Trivedi +6

Motorcycles face disproportionately high crash risks compared to cars due to limited protection and heightened sensitivity to surface hazards, yet Advanced Rider Assistance Systems…

cs.RO2026

AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning

Yangzhe Kong, Daeun Song, Jing Liang +3

We present a novel method, AutoSpatial, an efficient approach with structured spatial grounding to enhance VLMs' spatial reasoning. By combining minimal manual supervision with lar…

cs.CV2026

MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models

Xiyang Wu, Zongxia Li, Jihui Jin +7

Vision Language Models (VLMs) perform well on standard video tasks but struggle with physics-related reasoning involving motion dynamics and spatial interactions. We present a nove…

cs.RO2026

ViLAM: Distilling Vision-Language Reasoning into Attention Maps for Social Robot Navigation

Mohamed Elnoor, Kasun Weerakoon, Gershom Seneviratne +3

We introduce ViLAM, a novel method for distilling vision-language reasoning from large Vision-Language Models (VLMs) into spatial attention maps for socially compliant robot naviga…

cs.CV2026

FALCON: Future-Aware Learning with Contextual Object-Centric Pretraining for UAV Action Recognition

Ruiqi Xian, Xiyang Wu, Tianrui Guan +3

We introduce FALCON, a unified self-supervised video pretraining approach for UAV action recognition from raw RGB aerial footage, requiring no additional preprocessing at inference…