4 papers · 1 filter
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
Zhijing Cheng, Xuancheng Zhang, Donglin Di +4
End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downs…
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Baorui Ma, Jiahui Yang, Donglin Di +5
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camer…
Chain of World: World Model Thinking in Latent Motion
Fuxiang Yang, Donglin Di, Lulu Tang +6
Vision-Language-Action (VLA) models are a promising path toward embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynami…
TV-3DG: Mastering Text-to-3D Customized Generation with Visual Prompt
Jiahui Yang, Donglin Di, Baorui Ma +8
In recent years, advancements in generative models have significantly expanded the capabilities of text-to-3D generation. Many approaches rely on Score Distillation Sampling (SDS)…