8 papers
Walking in the Implicit: Interactive World Exploration via Neural Scene Representation
Zhiqi Li, Chengrui Dong, Zhenhua Du +6
Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent video frames, entangling state transition with high-frequency obse…
ChainFlow-VLA: Causal Flow Planning with Vision-Language Models
Xiyang Wang, Xinlin Wang, Tingguang Zhou +7
Current end-to-end autonomous driving systems are fundamentally limited by a mismatch between temporal causal reasoning and global trajectory consistency. Autoregressive (AR) model…
SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals
Zihang Lin, Huaiyuan Qin, Muli Yang +1
Assessing progress toward the Sustainable Development Goals (SDGs) requires multi-step reasoning over visual cues, contextual knowledge, and development indicators, where incomplet…
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
Minqing Huang, Yujiao Xiang, Zihan Liang +7
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide plannin…
Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale
Dongxu Wei, Qi Xu, Zhiqi Li +6
3D scene generation has long been dominated by 2D multi-view or video diffusion models. This is due not only to the lack of scene-level 3D latent representation, but also to the fa…
Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
Jianbo Zhao, Taiyu Ban, Xiyang Wang +6
Controllable trajectory generation guided by high-level semantic decisions, termed meta-actions, is crucial for autonomous driving systems. A significant limitation of existing fra…