70 citations · 131 across the 12 of their papers we have counts for
10 papers · 1 filter
Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach
Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal +1
The emergence of attention-based transformer models has led to their extensive use in various tasks, due to their superior generalization and transfer properties. Recent research h…
Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)
Shih-Han Chou, Matthew Kowal, Yasmin Niknam +8
While progress has been made in the domain of video-language understanding, current state-of-the-art algorithms are still limited in their ability to understand videos at high leve…
Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching
Junpeng Jing, Jiankun Li, Pengfei Xiong +7
Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not…
INVE: Interactive Neural Video Editing
Jiahui Huang, Leonid Sigal, Kwang Moo Yi +2
We present Interactive Neural Video Editing (INVE), a real-time video editing solution, which can assist the video editing process by consistently propagating sparse frame edits to…
MINOTAUR: Multi-task Video Grounding From Multimodal Queries
Raghav Goyal, Effrosyni Mavroudi, Xitong Yang +5
Video understanding tasks take many forms, from action detection to visual query localization and spatio-temporal grounding of sentences. These tasks differ in the type of inputs (…
Frustratingly Simple but Effective Zero-shot Detection and Segmentation: Analysis and a Strong Baseline
Siddhesh Khandelwal, Anirudth Nambirajan, Behjat Siddiquie +2
Methods for object detection and segmentation often require abundant instance-level annotations for training, which are time-consuming and expensive to collect. To address this, th…