5 papers
WA-JEPA: Rethinking the Video JEPA Paradigm for World-Action Modeling in Autonomous Driving
Xinlin Wang, Yujiao Xiang, Yuheng Zhou +11
Video Joint Embedding Predictive Architecture (V-JEPA) learns powerful spatiotemporal representations from video through self-supervised latent feature prediction. However, V-JEPA…
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
Minqing Huang, Yujiao Xiang, Zihan Liang +8
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide plannin…
From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes
Qifan Zhang, Sai Haneesh Allu, Jikai Wang +2
Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of template images, a robot must l…
Multimodal Reference Visual Grounding
Yangxiao Lu, Ruosen Li, Liqiang Jing +5
Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding pe…
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
Yangxiao Lu, Jishnu Jaykumar P, Yunhui Guo +2
Novel Instance Detection and Segmentation (NIDS) aims at detecting and segmenting novel object instances given a few examples of each instance. We propose a unified, simple, yet ef…