From the 1 of 6 linked papers with an AI index.
6 papers
Geo-VLA: Geometry-Aware Vision-Language-Action Planning via Internalization of Map Semantics
Ran Chen, Jiaxing Ren, Zhikun Zhang +3
Vision-language-action (VLA) models have advanced end-to-end autonomous driving by leveraging foundation models for semantic reasoning and long-tail generalization. However, their…
RRM: Experience-Driven Reflective Retrieval Memory for Long-Horizon Multimodal Reasoning
Jingxiang Fan, Junbao Zhuo, Bochao Zou
The paper proposes Reflective Retrieval Memory (RRM), a framework that adds a reflective experience memory to an entity‑centric multimodal memory graph, enabling agents to learn an…
Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models
Siqi Liu, Xinyang Li, Bochao Zou +3
As large language models (LLMs) continue to advance, there is increasing interest in their ability to infer human mental states and demonstrate a human-like Theory of Mind (ToM). M…
AGC-Drive: A Large-Scale Dataset for Real-World Aerial-Ground Collaboration in Driving Scenarios
Yunhao Hou, Bochao Zou, Min Zhang +7
By sharing information across multiple agents, collaborative perception helps autonomous vehicles mitigate occlusions and improve overall perception accuracy. While most previous w…
ME-TST+: Micro-expression Analysis via Temporal State Transition with ROI Relationship Awareness
Zizheng Guo, Bochao Zou, Junbao Zhuo +1
Micro-expressions (MEs) are regarded as important indicators of an individual's intrinsic emotions, preferences, and tendencies. ME analysis requires spotting of ME intervals withi…
RhythmFormer: Extracting Patterned rPPG Signals based on Periodic Sparse Attention
Bochao Zou, Zizheng Guo, Jiansheng Chen +3
Remote photoplethysmography (rPPG) is a non-contact method for detecting physiological signals based on facial videos, holding high potential in various applications. Due to the pe…