activity
20242026
collaborators

6 papers

cs.RO2026

Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future

Tianshuai Hu, Xiaolu Liu, Song Wang +17

Autonomous driving has long relied on modular "Perception-Decision-Action" pipelines, where hand-crafted interfaces and rule-based components often break down in complex or long-ta…

cs.CV2025

One Flight Over the Gap: A Survey from Perspective to Panoramic Vision

Xin Lin, Xian Ge, Dizhe Zhang +8

Driven by the demand for spatial intelligence and holistic scene perception, omnidirectional images (ODIs), which provide a complete 360\textdegree{} field of view, are receiving g…

cs.CV2025

UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions

Zhucun Xue, Jiangning Zhang, Teng Hu +8

The quality of the video dataset (image quality, resolution, and fine-grained caption) greatly influences the performance of the video generation model. The growing demand for vide…

cs.CV2025

AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding

Zhucun Xue, Jiangning Zhang, Xurong Xie +4

Multimodal Large Language Models (MLLMs) perform well in video understanding but degrade on long videos due to fixed-length context and weak long-term dependency modeling. Retrieva…

cs.CV2025

Generative Classifier for Domain Generalization

Shaocong Long, Qianyu Zhou, Xiangtai Li +5

Domain generalization (DG) aims to improve the generalizability of computer vision models toward distribution shifts. The mainstream DG methods focus on learning domain invariance,…

cs.CV2024

EMOv2: Pushing 5M Vision Model Frontier

Jiangning Zhang, Teng Hu, Haoyang He +6

This work focuses on developing parameter-efficient and lightweight models for dense predictions while trading off parameters, FLOPs, and performance. Our goal is to set up the new…