20 citations · 89 across the 7 of their papers we have counts for
7 papers · 1 filter
Weakly-Supervised Multi-Granularity Map Learning for Vision-and-Language Navigation
Peihao Chen, Dongyu Ji, Kunyang Lin +4
We address a practical yet challenging problem of training robot agents to navigate in an environment following a path described by some language instructions. The instructions oft…
Learning Active Camera for Multi-Object Navigation
Peihao Chen, Dongyu Ji, Kunyang Lin +5
Getting robots to navigate to multiple objects autonomously is essential yet difficult in robot applications. One of the key challenges is how to explore environments efficiently w…
RSPNet: Relative Speed Perception for Unsupervised Video Representation Learning
Peihao Chen, Deng Huang, Dongliang He +5
We study unsupervised video representation learning that seeks to learn both motion and appearance features from unlabeled video only, which can be reused for downstream tasks such…
Location-aware Graph Convolutional Networks for Video Question Answering
Deng Huang, Peihao Chen, Runhao Zeng +3
We addressed the challenging task of video question answering, which requires machines to answer questions about videos in a natural language form. Previous state-of-the-art method…
Foley Music: Learning to Generate Music from Videos
Chuang Gan, Deng Huang, Peihao Chen +2
In this paper, we introduce Foley Music, a system that can synthesize plausible music for a silent video clip about people playing musical instruments. We first identify two key in…
Dense Regression Network for Video Grounding
Runhao Zeng, Haoming Xu, Wenbing Huang +3
We address the problem of video grounding from natural language queries. The key challenge in this task is that one training video might only contain a few annotated starting/endin…