activity
20242026
collaborators

5 papers

cs.CV2026

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes

Qifan Zhang, Sai Haneesh Allu, Jikai Wang +2

Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of template images, a robot must l…

cs.RO2026

iTeach: In the Wild Interactive Teaching for Failure-Driven Adaptation of Robot Perception

Jishnu Jaykumar P, Cole Salvato, Vinaya Bomnale +2

Robotic perception models often fail when deployed in real-world environments due to out-of-distribution conditions such as clutter, occlusion, and novel object instances. Existing…

cs.CV2025

Multimodal Reference Visual Grounding

Yangxiao Lu, Ruosen Li, Liqiang Jing +5

Visual grounding focuses on detecting objects from images based on language expressions. Recent Large Vision-Language Models (LVLMs) have significantly advanced visual grounding pe…

cs.CV2025

HO-Cap: A Capture System and Dataset for 3D Reconstruction and Pose Tracking of Hand-Object Interaction

Jikai Wang, Qifan Zhang, Yu-Wei Chao +3

We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGBD cameras and…

cs.CV2024

CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities

Rohith Peddi, Shivvrat Arya, Bharath Challa +10

Following step-by-step procedures is an essential component of various activities carried out by individuals in their daily lives. These procedures serve as a guiding framework tha…