90 citations · 272 across the 19 of their papers we have counts for
15 papers · 1 filter
Self-supervised 3D Semantic Representation Learning for Vision-and-Language Navigation
Sinan Tan, Mengmeng Ge, Di Guo +2
In the Vision-and-Language Navigation task, the embodied agent follows linguistic instructions and navigates to a specific goal. It is important in many practical scenarios and has…
Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer Fusion
Yikai Wang, Fuchun Sun, Ming Lu +1
We propose a compact and effective framework to fuse multimodal features at multiple layers in a single network. The framework consists of two innovative fusion schemes. Firstly, u…
Unsupervised Representation Learning by InvariancePropagation
Feng Wang, Huaping Liu, Di Guo +1
Unsupervised learning methods based on contrastive learning have drawn increasing attention and achieved promising results. Most of them aim to learn representations invariant to i…
Resolution Switchable Networks for Runtime Efficient Image Recognition
Yikai Wang, Fuchun Sun, Duo Li +1
We propose a general method to train a single convolutional neural network which is capable of switching image resolutions at inference. Thus the running speed can be selected to m…
Reusing Discriminators for Encoding: Towards Unsupervised Image-to-Image Translation
Runfa Chen, Wenbing Huang, Binghui Huang +2
Unsupervised image-to-image translation is a central task in computer vision. Current translation frameworks will abandon the discriminator once the training process is completed.…
Deep Point-wise Prediction for Action Temporal Proposal
Luxuan Li, Tao Kong, Fuchun Sun +1
Detecting actions in videos is an important yet challenging task. Previous works usually utilize (a) sliding window paradigms, or (b) per-frame action scoring and grouping to enume…