activity
20122024
most citedLearning Language-Visual Embedding for Movie Understanding with Natural-Language

70 citations · 131 across the 12 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV20241 cited

Visual Prompting for Generalized Few-shot Segmentation: A Multi-scale Approach

Mir Rayat Imtiaz Hossain, Mennatullah Siam, Leonid Sigal +1

The emergence of attention-based transformer models has led to their extensive use in various tasks, due to their superior generalization and transfer properties. Recent research h…

cs.CV20241 cited

Multi-modal News Understanding with Professionally Labelled Videos (ReutersViLNews)

Shih-Han Chou, Matthew Kowal, Yasmin Niknam +8

While progress has been made in the domain of video-language understanding, current state-of-the-art algorithms are still limited in their ability to understand videos at high leve…

cs.CV2023

Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching

Junpeng Jing, Jiankun Li, Pengfei Xiong +7

Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not…

cs.CV20232 cited

INVE: Interactive Neural Video Editing

Jiahui Huang, Leonid Sigal, Kwang Moo Yi +2

We present Interactive Neural Video Editing (INVE), a real-time video editing solution, which can assist the video editing process by consistently propagating sparse frame edits to…

cs.CV20231 cited

MINOTAUR: Multi-task Video Grounding From Multimodal Queries

Raghav Goyal, Effrosyni Mavroudi, Xitong Yang +5

Video understanding tasks take many forms, from action detection to visual query localization and spatio-temporal grounding of sentences. These tasks differ in the type of inputs (…

cs.CV20234 cited

Frustratingly Simple but Effective Zero-shot Detection and Segmentation: Analysis and a Strong Baseline

Siddhesh Khandelwal, Anirudth Nambirajan, Behjat Siddiquie +2

Methods for object detection and segmentation often require abundant instance-level annotations for training, which are time-consuming and expensive to collect. To address this, th…