5 papers
Metis: A Generalizable and Efficient World-Action Model for Autonomous Driving and Urban Navigation
Jingyu Li, Zhe Liu, Dongnan Hu +10
World action models~(WAMs) have shown great promise for autonomous driving and urban navigation. Built upon Vision-Language-Action models or video generation models, existing appro…
StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering
Ming Xie, Zizheng Huang, Xudong Tan +6
While streaming omni-video understanding demands continuous perception and proactive, real-time interaction, this crucial area remains largely under-explored. Current omni-modal me…
MCNav: Memory-Aware Dynamic Cognitive Map for Zero-shot Goal-oriented Navigation
Jingyu Li, Zhe Liu, Wenxiao Wu +1
Navigating to instance-level targets in complex environments is a challenging problem. Many existing zero-shot methods achieve strong performance by modeling the entire environment…
GeoTeacher: Geometry-Guided Semi-Supervised 3D Object Detection
Jingyu Li, Xiaolong Zhao, Zhe Liu +2
Semi-supervised 3D object detection, aiming to explore unlabeled data for boosting 3D object detectors, has emerged as an active research area in recent years. Some previous method…
Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
Wenxiao Wu, Jing-Hao Xue, Chengming Xu +5
Visual In-Context Learning (VICL) has emerged as a prominent approach for adapting visual foundation models to novel tasks, by effectively exploiting contextual information embedde…