collaborators

6 papers

cs.CV2026

SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving

Peizheng Li, Zhenghao Zhang, David Holtz +6

End-to-end autonomous driving methods built on vision language models (VLMs) have undergone rapid development driven by their universal visual understanding and strong reasoning ca…

cs.RO2026

Glance-Say: Multimodal Human-Robot Collaboration and Intent Recognition via Sticky Glance

Yuzhi Lai, Shenghai Yuan, Peizheng Li +2

Gaze and speech are promising interaction modalities for individuals with motor impairments, yet robust intent recognition in multi-object environments remains challenging due to m…

cs.CV2026

Seizure-Semiology-Suite (S3): A Clinically Multimodal Dataset, Benchmark, and Models for Seizure Semiology Understanding

Lina Zhang, Tonmoy Monsoor, Peizheng Li +23

While Multimodal Large Language Models (MLLMs) have demonstrated remarkable proficiency in general video understanding, their capacity to interpret involuntary, and spatio-temporal…

cs.HC2026

FAM-HRI: Foundation-Model Assisted Multi-Modal Human-Robot Interaction Combining Gaze and Speech

Yuzhi Lai, Shenghai Yuan, Peizheng Li +4

ffective Human-Robot Interaction (HRI) is crucial for enhancing accessibility and usability in real-world robotics applications. However, existing solutions often rely on gesture-…

cs.CV2026

Can Multimodal Large Language Models Understand Pathologic Movements? A Pilot Study on Seizure Semiology

Lina Zhang, Tonmoy Monsoor, Mehmet Efe Lorasdagi +8

Multimodal Large Language Models (MLLMs) have demonstrated robust capabilities in recognizing everyday human activities, yet their potential for analyzing clinically significant in…

cs.CV2025

SEER-VAR: Semantic Egocentric Environment Reasoner for Vehicle Augmented Reality

Yuzhi Lai, Shenghai Yuan, Peizheng Li +2

We present SEER-VAR, a novel framework for egocentric vehicle-based augmented reality (AR) that unifies semantic decomposition, Context-Aware SLAM Branches (CASB), and LLM-driven r…