12 papers
UNIVERSE: Unified Video Action Models for Autonomous Driving with Flexible Mask-Modulated Modality Generation
Mengmeng Liu, Diankun Zhang, Jiuming Liu +7
World Action Models (WAMs) have shown strong potential for improving action generalization in autonomous driving by using future video prediction as dense supervision for scene dyn…
Revisiting Chain-of-Thought Reasoning under Limited Supervision: Semi-supervised Chain-of-Thought Learning
Hongyang He, Jiuming Liu, Victor Sanchez
Chain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent reasoning capabilities in large language models. However, most existing CoT methods use…
DriveVA: Video Action Models are Zero-Shot Drivers
Mengmeng Liu, Diankun Zhang, Jiuming Liu +7
Generalization is a central challenge in autonomous driving, as real-world deployment requires robust performance under unseen scenarios, sensor domains, and environmental conditio…
Semi-Supervised Vision-Language-Action Model
Hongyang He, Jiuming Liu, Victor Sanchez
Vision-Language-Action (VLA) models enable robots to predict actions directly from visual observations and language instructions, but adapting them to new environments still depend…
Towards Interactive Video World Modeling: Frontiers, Challenges, Benchmarks, and Future Trends
Jiuming Liu, Chaojun Ni, Mengmeng Liu +7
With rapid development of large language models and diffusion-based content generation, world modeling has attracted increasing research attention, benefiting various downstream do…
Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory
Tianchen Deng, Zhenxiang Xiong, Nailin Wang +4
Visual Geometry Grounded Transformers (VGGT) have set new benchmarks in high-fidelity 3D scene reconstruction. However, as the sequence length increases, these models suffer from c…