16 citations · 43 across the 6 of their papers we have counts for
17 papers
InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
Zehua Chen, Xu Tan, Ke Wang +4
Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses…
AdaSpeech 3: Adaptive Text to Speech for Spontaneous Style
Yuzi Yan, Xu Tan, Bohan Li +6
While recent text to speech (TTS) models perform very well in synthesizing reading-style (e.g., audiobook) speech, it is still challenging to synthesize spontaneous-style speech (e…
Libri-adhoc40: A dataset collected from synchronized ad-hoc microphone arrays
Shanzheng Guan, Shupei Liu, Junqi Chen +8
Recently, there is a research trend on ad-hoc microphone arrays. However, most research was conducted on simulated data. Although some data sets were collected with a small number…
Cross-domain Speech Recognition with Unsupervised Character-level Distribution Matching
Wenxin Hou, Jindong Wang, Xu Tan +2
End-to-end automatic speech recognition (ASR) can achieve promising performance with large-scale training data. However, it is known that domain mismatch between training and testi…
Adaptive Logit Adjustment Loss for Long-Tailed Visual Recognition
Yan Zhao, Weicong Chen, Xu Tan +2
Data in the real world tends to exhibit a long-tailed label distribution, which poses great challenges for the training of neural networks in visual recognition. Existing methods t…
BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction
Yuhang Li, Ruihao Gong, Xu Tan +6
We study the challenging task of neural network quantization without end-to-end retraining, called Post-training Quantization (PTQ). PTQ usually requires a small subset of training…