43 papers
MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving
Zhijing Cheng, Xuancheng Zhang, Donglin Di +4
End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downs…
MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources
Baorui Ma, Jiahui Yang, Donglin Di +5
Scaling has powered recent advances in vision foundation models, yet extending this paradigm to metric depth estimation remains challenging due to heterogeneous sensor noise, camer…
ArcAD: Anomaly-Rectified Calibration for Cold-Start Supervised Anomaly Detection
Ningning Han, Lei Fan, Jia Guo +5
The deployment of Industrial Anomaly Detection (IAD) in real-world manufacturing frequently encounters a challenging cold-start bottleneck, in which limited normal samples fail to…
CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning
He Feng, Yongjia Ma, Donglin Di +2
Portrait animation methods have achieved substantial visual quality and lip synchronization, but fine-grained manipulation of the eye region still faces a trade-off between input g…
Can Vision-Language Models Think from the Sky? Unifying UAV Reasoning and Generation
Jintao Sun, Gangyi Ding, Donglin Di +2
Vision-Language Models have achieved strong progress in ground-view visual understanding, yet they remain brittle in high-altitude Unmanned Aerial Vehicle scenes, where objects are…
ScrollScape: Unlocking 32K Image Generation With Video Diffusion Priors
Haodong Yu, Yabo Zhang, Donglin Di +2
While diffusion models excel at generating images with conventional dimensions, pushing them to synthesize ultra-high-resolution imagery at extreme aspect ratios (EAR) often trigge…