6 papers
SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
Peizheng Li, Zhenghao Zhang, David Holtz +6
End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning ca…
Glance-Say: Multimodal Human-Robot Collaboration and Intent Recognition via Sticky Glance
Yuzhi Lai, Shenghai Yuan, Peizheng Li +2
Gaze and speech are promising interaction modalities for individuals with motor impairments, yet robust intent recognition in multi-object environments remains challenging due to m…
Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding
Lina Zhang, Tonmoy Monsoor, Peizheng Li +23
While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporal…
FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech
Yuzhi Lai, Shenghai Yuan, Peizheng Li +4
ffective Human-Robot Interaction (HRI) is crucial for enhancing accessibility and usability in real-world robotics applications. However, existing solutions often rely on gesture-…
Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology
Lina Zhang, Tonmoy Monsoor, Mehmet Efe Lorasdagi +8
Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzing clinically significant in…
SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality
Yuzhi Lai, Shenghai Yuan, Peizheng Li +2
We present SEER-VAR, a novel framework for egocentric vehicle-based augmented reality (AR) that unifies semantic decomposition, Context-Aware SLAM Branches (CASB), and LLM-driven r…