16 citations · 50 across the 6 of their papers we have counts for
6 papers
Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders
Haosen Yang, Deng Huang, Bin Wen +5
Masked autoencoders (MAEs) have emerged recently as art self-supervised spatiotemporal representation learners. Inheriting from the image counterparts, however, existing video MAEs…
ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency
Deng Huang, Wenhao Wu, Weiwen Hu +6
We study self-supervised video representation learning, which is a challenging task due to 1) lack of labels for explicit supervision; 2) unstructured and noisy visual information.…
RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning
Peihao Chen, Deng Huang, Dongliang He +5
We study unsupervised video representation learning that seeks to learn both motion and appearance features from unlabeled video only, which can be reused for downstream tasks such…
Location-aware Graph Convolutional Networks for Video Question Answering
Deng Huang, Peihao Chen, Runhao Zeng +3
We addressed the challenging task of video question answering, which requires machines to answer questions about videos in a natural language form. Previous state-of-the-art method…
Foley Music: Learning to Generate Music from Videos
Chuang Gan, Deng Huang, Peihao Chen +2
In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key in…
Music Gesture for Visual Sound Separation
Chuang Gan, Deng Huang, Hang Zhao +2
Recent deep learning approaches have achieved impressive performance on visual sound separation tasks. However, these approaches are mostly built on appearance and optical flow lik…