30 citations · 30 across the 2 of their papers we have counts for
5 papers · 1 filter
Panoptic Narrative Grounding
C. González, N. Ayobi, I. Hernández +3
This paper proposes Panoptic Narrative Grounding, a spatially fine and general formulation of the natural language visual grounding problem. We establish an experimental framework…
APES: Audiovisual Person Search in Untrimmed Video
Juan Leon Alcazar, Long Mai, Federico Perazzi +4
Humans are arguably one of the most important subjects in video streams, many real-world applications such as video summarization or video editing workflows often require the autom…
ISINet: An Instance-Based Approach for Surgical Instrument Segmentation
Cristina González, Laura Bravo-Sánchez, Pablo Arbelaez
We study the task of semantic segmentation of surgical instruments in robotic-assisted surgery scenes. We propose the Instance-based Surgical Instrument Segmentation Network (ISINe…
MAIN: Multi-Attention Instance Network for Video Segmentation
Juan Leon Alcazar, Maria A. Bravo, Ali K. Thabet +4
Instance-level video segmentation requires a solid integration of spatial and temporal information. However, current methods rely mostly on domain-specific information (online lear…
Inferring 3D Object Pose in RGB-D Images
Saurabh Gupta, Pablo Arbeláez, Ross Girshick +1
The goal of this work is to replace objects in an RGB-D scene with corresponding 3D models from a library. We approach this problem by first detecting and segmenting object instanc…