20 citations · 49 across the 4 of their papers we have counts for
4 papers
Location-aware Graph Convolutional Networks for Video Question Answering
Deng Huang, Peihao Chen, Runhao Zeng +3
We addressed the challenging task of video question answering, which requires machines to answer questions about videos in a natural language form. Previous state-of-the-art method…
Foley Music: Learning to Generate Music from Videos
Chuang Gan, Deng Huang, Peihao Chen +2
In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key in…
Dense Regression Network for Video Grounding
Runhao Zeng, Haoming Xu, Wenbing Huang +3
We address the problem of video grounding from natural language queries. The key challenge in this task is that one training video might only contain a few annotated starting/endin…
Self-supervised Moving Vehicle Tracking with Stereo Sound
Chuang Gan, Hang Zhao, Peihao Chen +2
Humans are able to localize objects in the environment using both visual and auditory cues, integrating information from multiple modalities into a common reference frame. We intro…