6 papers
Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
Tianshuai Hu, Xiaolu Liu, Song Wang +17
Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-ta…
One Flight Over the Gap: A Survey from Perspective to Panoramic Vision
Xin Lin, Xian Ge, Dizhe Zhang +8
Driven by the demand for spatial intelligence and holistic scene perception, omnidirectional images (ODIs), which provide a complete 360\textdegree{} field of view, are receiving g…
UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
Zhucun Xue, Jiangning Zhang, Teng Hu +8
The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for vide…
AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
Zhucun Xue, Jiangning Zhang, Xurong Xie +4
Multimodal Large Language Models (MLLMs) perform well in video understanding but degrade on long videos due to fixed-length context and weak long-term dependency modeling. Retrieva…
Generative Classifier for Domain Generalization
Shaocong Long, Qianyu Zhou, Xiangtai Li +5
Domain generalization (DG) aims to improve the generalizability of computer vision models toward distribution shifts. The mainstream DG methods focus on learning domain invariance,…
EMOv2: Pushing 5M Vision Model Frontier
Jiangning Zhang, Teng Hu, Haoyang He +6
This work focuses on developing parameter-efficient and lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Our goal is to set up the new…