22 citations · 38 across the 30 of their papers we have counts for
12 papers · 2 filters
MindDrive: A Vision-Language-Action Model for Autonomous Driving via Online Reinforcement Learning
Haoyu Fu, Diankun Zhang, Zongchuang Zhao +7
Current Vision-Language-Action (VLA) paradigms in autonomous driving primarily rely on Imitation Learning (IL), which introduces inherent challenges such as distribution shift and…
Cook and Clean Together: Teaching Embodied Agents for Parallel Task Execution
Dingkang Liang, Cheng Zhang, Xiaopeng Xu +3
Task scheduling is critical for embodied AI, enabling agents to follow natural language instructions and execute actions efficiently in 3D physical worlds. However, existing datase…
NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
Wei Xu, Cheng Wang, Dingkang Liang +4
Underwater exploration offers critical insights into our planet and attracts increasing attention for its broader applications in resource exploration, national security, etc. We s…
More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
Hongkai Lin, Dingkang Liang, Mingyang Du +2
Generative depth estimation methods leverage the rich visual priors stored in pre-trained text-to-image diffusion models, demonstrating astonishing zero-shot capability. However, p…
SoccerNet 2025 Challenges Results
Silvio Giancola, Anthony Cioppa, Marc Gutiérrez-Pérez +115
The SoccerNet 2025 Challenges mark the fifth annual edition of the SoccerNet open benchmarking effort, dedicated to advancing computer vision research in football video understandi…
Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching
Xin Zhou, Dingkang Liang, Kaijin Chen +7
Video generation models have demonstrated remarkable performance, yet their broader adoption remains constrained by slow inference speeds and substantial computational costs, prima…