55 citations · 217 across the 34 of their papers we have counts for
11 papers · 1 filter
TDAN: Temporally Deformable Alignment Network for Video Super-Resolution
Yapeng Tian, Yulun Zhang, Yun Fu +1
Video super-resolution (VSR) aims to restore a photo-realistic high-resolution (HR) video frame from both its corresponding low-resolution (LR) frame (reference frame) and multiple…
An Attempt towards Interpretable Audio-Visual Video Captioning
Yapeng Tian, Chenxiao Guan, Justin Goodman +2
Automatically generating a natural language sentence to describe the content of an input video is a very challenging problem. It is an essential multimodal task in which auditory a…
How to Make a BLT Sandwich? Learning to Reason towards Understanding Web Instructional Videos
Shaojie Wang, Wentian Zhao, Ziyi Kou +1
Understanding web instructional videos is an essential branch of video understanding in two aspects. First, most existing video methods focus on short-term actions for a-few-second…
GAN-EM: GAN based EM learning framework
Wentian Zhao, Shaojie Wang, Zhihuai Xie +2
Expectation maximization (EM) algorithm is to find maximum likelihood solution for models having latent variables. A typical example is Gaussian Mixture Model (GMM) which requires…
Dynamic Graph Modules for Modeling Object-Object Interactions in Activity Recognition
Hao Huang, Luowei Zhou, Wei Zhang +2
Video action recognition, a critical problem in video understanding, has been gaining increasing attention. To identify actions induced by complex object-object interactions, we ne…
Navigation by Imitation in a Pedestrian-Rich Environment
Jing Bi, Tianyou Xiao, Qiuyue Sun +1
Deep neural networks trained on demonstrations of human actions give robot the ability to perform self-driving on the road. However, navigation in a pedestrian-rich environment, su…