most citedControllable and Lossless Non-Autoregressive End-to-End Text-to-Speech

5 citations · 10 across the 6 of their papers we have counts for

collaborators

6 papers

cs.LG2023

Imitator Learning: Achieve Out-of-the-Box Imitation Ability in Variable Environments

Xiong-Hui Chen, Junyin Ye, Hang Zhao +9

Imitation learning (IL) enables agents to mimic expert behaviors. Most previous IL techniques focus on precisely imitating one policy through mass demonstrations. However, in many…

cs.CV2023

Improving Discriminative Multi-Modal Learning with Large-Scale Pre-Trained Models

Chenzhuang Du, Yue Zhao, Chonghua Liao +3

This paper investigates how to better leverage large-scale pre-trained uni-modal models to further enhance discriminative multi-modal learning. Even when fine-tuned with only uni-m…

cs.CV20233 cited

GeoMAE: Masked Geometric Target Prediction for Self-supervised Point Cloud Pre-Training

Xiaoyu Tian, Haoxi Ran, Yue Wang +1

This paper tries to address a fundamental question in point cloud self-supervised learning: what is a good signal we should leverage to learn features from point clouds without ann…

cs.CV2023

What Happened 3 Seconds Ago? Inferring the Past with Thermal Imaging

Zitian Tang, Wenjie Ye, Wei-Chiu Ma +1

Inferring past human motion from RGB images is challenging due to the inherent uncertainty of the prediction problem. Thermal images, on the other hand, encode traces of past human…

cs.SD20225 cited

Controllable and Lossless Non-Autoregressive End-to-End Text-to-Speech

Zhengxi Liu, Qiao Tian, Chenxu Hu +5

Some recent studies have demonstrated the feasibility of single-stage neural text-to-speech, which does not need to generate mel-spectrograms but generates the raw waveforms direct…

cs.CV20222 cited

AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection

Zehui Chen, Zhenyu Li, Shiquan Zhang +5

Object detection through either RGB images or the LiDAR point clouds has been extensively explored in autonomous driving. However, it remains challenging to make these two data sou…