activity
20192021
most citedReferring Transformer: A One-step Approach to Multi-task Visual Grounding

73 citations · 173 across the 6 of their papers we have counts for

collaborators

6 papers

cs.CV202173 cited

Referring Transformer: A One-step Approach to Multi-task Visual Grounding

Muchen Li, Leonid Sigal

As an important step towards visual reasoning, visual grounding (e.g., phrase localization, referring expression comprehension/segmentation) has been widely explored Previous appro…

cs.CV202150 cited

Learning Spatial and Spatio-Temporal Pixel Aggregations for Image and Video Denoising

Xiangyu Xu, Muchen Li, Wenxiu Sun +1

Existing denoising methods typically restore clear results by aggregating pixels from the noisy input. Instead of relying on hand-crafted aggregation schemes, we propose to explici…

cs.CV20202 cited

TDAF: Top-Down Attention Framework for Vision Tasks

Bo Pang, Yizhuo Li, Jiefeng Li +3

Human attention mechanisms often work in a top-down manner, yet it is not well explored in vision research. Here, we propose the Top-Down Attention Framework (TDAF) to capture top-…

eess.IV2020

NTIRE 2020 Challenge on Video Quality Mapping: Methods and Results

Dario Fuoli, Zhiwu Huang, Martin Danelljan +18

This paper reviews the NTIRE 2020 challenge on video quality mapping (VQM), which addresses the issues of quality mapping from source video domain to target video domain. The chall…

cs.CV202010 cited

TubeTK: Adopting Tubes to Track Multi-Object in a One-Step Training Model

Bo Pang, Yizhuo Li, Yifan Zhang +2

Multi-object tracking is a fundamental vision problem that has been studied for a long time. As deep learning brings excellent performances to object detection algorithms, Tracking…

cs.CV201938 cited

Learning Deformable Kernels for Image and Video Denoising

Xiangyu Xu, Muchen Li, Wenxiu Sun

Most of the classical denoising methods restore clear results by selecting and averaging pixels in the noisy input. Instead of relying on hand-crafted selecting and averaging strat…