From the 1 of 17 linked papers with an AI index.
17 papers
Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics
Ran Chen, Jiaxing Ren, Zhikun Zhang +3
Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their…
RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
Jingxiang Fan, Junbao Zhuo, Bochao Zou
The paper proposes Reflective Retrieval Memory (RRM), a framework that adds a reflective experience memory to an entity‑centric multimodal memory graph, enabling agents to learn an…
Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images
Hongyuan Liu, Bochao Zou, Qiankun Liu +10
Creating realistic and simulation-ready 3D assets is crucial for autonomous driving research and virtual environment construction. However, existing 3D vehicle generation methods a…
Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models
Siqi Liu, Xinyang Li, Bochao Zou +3
As large language models (LLMs) continue to advance, there is increasing interest in their ability to infer human mental states and demonstrate a human-like Theory of Mind (ToM). M…
PHASE-Net: Physics-Grounded Harmonic Attention System for Efficient Remote Photoplethysmography Measurement
Bo Zhao, Dan Guo, Junzhe Cao +5
Remote photoplethysmography (rPPG) measurement enables non-contact physiological monitoring but suffers from accuracy degradation under head motion and illumination changes. Existi…
AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving Scenarios
Yunhao Hou, Bochao Zou, Min Zhang +7
By sharing information across multiple agents, collaborative perception helps autonomous vehicles mitigate occlusions and improve overall perception accuracy. While most previous w…