activity
20152023
most citedReferring Transformer: A One-step Approach to Multi-task Visual Grounding

73 citations · 191 across the 24 of their papers we have counts for

collaborators
Showing cs.CVShow all

41 papers · 1 filter

cs.CV2023

Uncertainty Guided Adaptive Warping for Robust and Efficient Stereo Matching

Junpeng Jing, Jiankun Li, Pengfei Xiong +7

Correlation based stereo matching has achieved outstanding performance, which pursues cost volume between two feature maps. Unfortunately, current methods with a fixed model do not…

cs.CV20232 cited

INVE: Interactive Neural Video Editing

Jiahui Huang, Leonid Sigal, Kwang Moo Yi +2

We present Interactive Neural Video Editing (INVE), a real-time video editing solution, which can assist the video editing process by consistently propagating sparse frame edits to…

cs.CV20221 cited

VLC-BERT: Visual Question Answering with Contextualized Commonsense Knowledge

Sahithya Ravi, Aditya Chinchure, Leonid Sigal +2

There has been a growing interest in solving Visual Question Answering (VQA) tasks that require the model to reason beyond the content present in the image. In this work, we focus…

cs.CV20225 cited

Real-Time Monitoring of User Stress, Heart Rate and Heart Rate Variability on Mobile Devices

Peyman Bateni, Leonid Sigal

Stress is considered to be the epidemic of the 21st-century. Yet, mobile apps cannot directly evaluate the impact of their content and services on user stress. We introduce the Bea…

cs.CV2021

TriBERT: Full-body Human-centric Audio-visual Representation Learning for Visual Sound Separation

Tanzila Rahman, Mengyu Yang, Leonid Sigal

The recent success of transformer models in language, such as BERT, has motivated the use of such architectures for multi-modal feature learning and tasks. However, most multi-moda…

cs.CV202173 cited

Referring Transformer: A One-step Approach to Multi-task Visual Grounding

Muchen Li, Leonid Sigal

As an important step towards visual reasoning, visual grounding (e.g., phrase localization, referring expression comprehension/segmentation) has been widely explored Previous appro…