Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
X-Mind: Efficient Visual Chain-of-Thought via Predictive World Model for End-to-End Driving
Bohao Zhao, Chengrui Wei, Guangfeng Jiang +17
Predicting future states is essential for autonomous agents, yet current Vision-Language-Action (VLA) models fundamentally lack this capability, relying instead on reactive percept…
cs.CV2026
X-Foresight: A Joint Vision-Action Causal Forecasting Network via Predictive World Modeling
Baolu Li, Jingyu Qian, Rui Guo +17
Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive…
cs.CV2024
You Only Need Less Attention at Each Stage in Vision Transformers
Shuoxi Zhang, Hanpeng Liu, Stephen Lin +1
The advent of Vision Transformers (ViTs) marks a substantial paradigm shift in the realm of computer vision. ViTs capture the global information of images through self-attention mo…