66 citations · 180 across the 17 of their papers we have counts for
32 papers
Self-supervised 3D Semantic Representation Learning for Vision-and-Language Navigation
Sinan Tan, Mengmeng Ge, Di Guo +2
In the Vision-and-Language Navigation task, the embodied agent follows linguistic instructions and navigates to a specific goal. It is important in many practical scenarios and has…
Audio-Visual Grounding Referring Expression for Robotic Manipulation
Yefei Wang, Kaili Wang, Yi Wang +3
Referring expressions are commonly used when referring to a specific target in people's daily dialogue. In this paper, we develop a novel task of audio-visual grounding referring e…
Multi-Agent Embodied Visual Semantic Navigation with Scene Prior Knowledge
Xinzhu Liu, Di Guo, Huaping Liu +1
In visual semantic navigation, the robot navigates to a target object with egocentric visual observations and the class label of the target is given. It is a meaningful task inspir…
Knowledge-based Embodied Question Answering
Sinan Tan, Mengmeng Ge, Di Guo +2
In this paper, we propose a novel Knowledge-based Embodied Question Answering (K-EQA) task, in which the agent intelligently explores the environment to answer various questions wi…
Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer Fusion
Yikai Wang, Fuchun Sun, Ming Lu +1
We propose a compact and effective framework to fuse multimodal features at multiple layers in a single network. The framework consists of two innovative fusion schemes. Firstly, u…
Adversarial Option-Aware Hierarchical Imitation Learning
Mingxuan Jing, Wenbing Huang, Fuchun Sun +4
It has been a challenge to learning skills for an agent from long-horizon unannotated demonstrations. Existing approaches like Hierarchical Imitation Learning(HIL) are prone to com…