10 papers
X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling
Baolu Li, Jingyu Qian, Rui Guo +17
Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive…
X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving
Bohao Zhao, Chengrui Wei, Guangfeng Jiang +17
Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive percept…
TuringViT: Making SOTA Vision Transformers Accessible to All
Qiman Wu, Hanlin Chen, Lyujie Chen +19
Modern VLMs and VLA systems commonly adopt off-the-shelf ViTs such as SigLIP2 as visual encoders, but diverse downstream requirements in latency, temporal modeling, and VLM integra…
EvoDriveVLA: Evolving Driving VLA Models via Collaborative Perception-Planning Distillation
Jiajun Cao, Xiaoan Zhang, Xiaobao Wei +10
Vision-Language-Action models have shown great promise for autonomous driving, yet they suffer from degraded perception after unfreezing the visual encoder and struggle with accumu…
X-Cache: Cross-Chunk Block Caching for Few-Step Autoregressive World Models Inference
Yixiao Zeng, Jianlei Zheng, Chaoda Zheng +10
Real-time world simulation is becoming a key infrastructure for scalable evaluation and online reinforcement learning of autonomous driving systems. Recent driving world models bui…
X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving
Chaoda Zheng, Sean Li, Jinhao Deng +9
Scalable and reliable evaluation is increasingly critical in the end-to-end era of autonomous driving, where vision--language--action (VLA) policies directly map raw sensor streams…