3 papers
cs.CV2026
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
Zhijing Cheng, Xuancheng Zhang, Donglin Di +4
End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downs…
cs.CV2026
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Baorui Ma, Jiahui Yang, Donglin Di +5
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camer…
cs.CV2026
Chain of World: World Model Thinking in Latent Motion
Fuxiang Yang, Donglin Di, Lulu Tang +6
Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynami…