4 papers
CoWorld-VLA: Thinking in a Multi-Expert World Model for Autonomous Driving
Minqing Huang, Yujiao Xiang, Zihan Liang +7
Vision-Language-Action (VLA) models have emerged as a promising paradigm for end-to-end autonomous driving. However, existing reasoning mechanisms still struggle to provide plannin…
From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes
Qifan Zhang, Sai Haneesh Allu, Jikai Wang +2
Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of template images, a robot must l…
Multimodal Reference Visual Grounding
Yangxiao Lu, Ruosen Li, Liqiang Jing +5
Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding pe…
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
Yangxiao Lu, Jishnu Jaykumar P, Yunhui Guo +2
Novel Instance Detection and Segmentation (NIDS) aims at detecting and segmenting novel object instances given a few examples of each instance. We propose a unified, simple, yet ef…