activity
20182020
most citedMultiple Sound Sources Localization from Coarse to Fine

8 citations · 24 across the 9 of their papers we have counts for

collaborators

12 papers

cs.CV20204 cited

Mask Guided Matting via Progressive Refinement Network

Qihang Yu, Jianming Zhang, He Zhang +5

We propose Mask Guided (MG) Matting, a robust matting framework that takes a general coarse mask as guidance. MG Matting leverages a network (PRN) design which encourages the matti…

cs.CV2020

Finding Action Tubes with a Sparse-to-Dense Framework

Yuxi Li, Weiyao Lin, Tao Wang +5

The task of spatial-temporal action detection has attracted increasing attention among researchers. Existing dominant methods solve this problem by relying on short-term informatio…

cs.CV2020

CFAD: Coarse-to-Fine Action Detector for Spatiotemporal Action Localization

Yuxi Li, Weiyao Lin, John See +4

Most current pipelines for spatio-temporal action localization connect frame-wise or clip-wise detection results to generate action proposals, where only local information is explo…

cs.CL20201 cited

Video Question Answering on Screencast Tutorials

Wentian Zhao, Seokhwan Kim, Ning Xu +1

This paper presents a new video question answering task on screencast tutorials. We introduce a dataset including question, answer and context triples from the tutorial videos for…

cs.CV20201 cited

Incorporating Reinforced Adversarial Learning in Autoregressive Image Generation

Kenan E. Ak, Ning Xu, Zhe Lin +1

Autoregressive models recently achieved comparable results versus state-of-the-art Generative Adversarial Networks (GANs) with the help of Vector Quantized Variational AutoEncoders…

cs.CV20208 cited

Multiple Sound Sources Localization from Coarse to Fine

Rui Qian, Di Hu, Heinrich Dinkel +3

How to visually localize multiple sound sources in unconstrained videos is a formidable problem, especially when lack of the pairwise sound-object annotations. To solve this proble…