2 citations · 7 across the 11 of their papers we have counts for
11 papers
Vision-Language Assistant for Emotional Reactions to Risky Driving
Harine Choi, Eun Hak Lee, Zhengzhong Tu
This study introduces a vision-language pipeline that detects risky driving behaviors and generates emotionally expressive responses to support driver awareness and comfort. Althou…
A Physics-Grounded Benchmark for Multi-Agent Dynamics in World Models
Nuo Chen, Lulin Liu, Zihao Li +12
Generative world models hold immense promise as scalable simulators for autonomous systems, particularly for synthesizing rare but safety-critical multi-agent interactions, such as…
Neuro-Symbolic Drive: Rule-Grounded Faithful Reasoning for Driving VLAs
Xiangbo Gao, Xiukun Huang, Boyu Lu +5
Driving VLA models incorporating Chain-of-Thought (CoT) reasoning are attractive because they leverage pretrained VLM representations and expose intermediate decisions in natural l…
RAPID: A Reproducible Multi-Agent Pipeline for Interpretable Disaster Damage Assessment from Satellite and Street-View Imagery
Yifan Yang, Wenjing Gong, Kaili Zhang +5
Due to the increasing frequency and intensity of extreme climate events, there is a clear demand for intelligent, scalable, and autonomous approaches to disaster damage assessment.…
Beyond Thinking: Imagining in 360 for Humanoid Visual Search
Jingdong Zhang, Yizhou Wang, Zhengzhong Tu +3
Humanoid Visual Search (HVS) requires agents to actively explore immersive 360 environments. While prior methods treat this as a monolithic task relying on cumulative, mult…
NavTrust: Benchmarking Trustworthiness for Embodied Navigation
Huaide Jiang, Yash Chaudhary, Yuping Wang +8
There are two major categories of embodied navigation: Vision-Language Navigation (VLN), where agents navigate by following natural language instructions; and Object-Goal Navigatio…