From the 1 of 8 linked papers with an AI index.
8 papers
WorkDrive: Roadwork Chain of Causation for Autonomous Driving
Tianyi Jiang, Wen Zhang, Sihan Yang +2
The paper introduces WorkDrive, a framework that adds perception‑grounded causal reasoning to vision‑language models for autonomous driving in roadwork zones, improving trajectory…
Concept-as-Tree: A Controllable Synthetic Data Framework Makes Stronger Personalized VLMs
Ruichuan An, Kai Zeng, Ming Lu +5
Vision-Language Models (VLMs) have demonstrated exceptional performance in various multi-modal tasks. Recently, there has been an increasing interest in improving the personalizati…
MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
Teng Hu, Mingchun Lu, Yating Wang +6
Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a sin…
AD-MIR: Bridging the Gap from Perception to Persuasion in Advertising Video Understanding via Structured Reasoning
Binxiao Xu, Junyu Feng, Xiaopeng Lin +7
Multimodal understanding of advertising videos is essential for interpreting the intricate relationship between visual storytelling and abstract persuasion strategies. However, des…
M2A: Multimodal Memory Agent with Dual-Layer Hybrid Memory for Long-Term Personalized Interactions
Junyu Feng, Binxiao Xu, Jiayi Chen +8
This work addresses the challenge of personalized question answering in long-term human-machine interactions: when conversational history spans weeks or months and exceeds the cont…
VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
Ming Zhong, Yuanlei Wang, Liuzhou Zhang +7
While Multimodal Large Language Models (MLLMs) excel on benchmarks, their processing paradigm differs from the human ability to integrate visual information. Unlike humans who natu…