78 citations · 121 across the 7 of their papers we have counts for
9 papers
AlignVE: Visual Entailment Recognition Based on Alignment Relations
Biwei Cao, Jiuxin Cao, Jie Gui +5
Visual entailment (VE) is to recognize whether the semantics of a hypothesis text can be inferred from the given premise image, which is one special task among recent emerged visio…
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
Xu Tan, Jiawei Chen, Haohe Liu +11
Text to speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality…
AdaSpeech 4: Adaptive Text to Speech in Zero-Shot Scenarios
Yihan Wu, Xu Tan, Bohan Li +5
Adaptive text to speech (TTS) can synthesize new voices in zero-shot scenarios efficiently, by using a well-trained source TTS model without adapting it on the speech data of new s…
InferGrad: Improving Diffusion Models for Vocoder by Considering Inference in Training
Zehua Chen, Xu Tan, Ke Wang +4
Denoising diffusion probabilistic models (diffusion models for short) require a large number of iterations in inference to achieve the generation quality that matches or surpasses…
Diff-Net: Image Feature Difference based High-Definition Map Change Detection for Autonomous Driving
Lei He, Shengjie Jiang, Xiaoqing Liang +2
Up-to-date High-Definition (HD) maps are essential for self-driving cars. To achieve constantly updated HD maps, we present a deep neural network (DNN), Diff-Net, to detect changes…
NPE: An FPGA-based Overlay Processor for Natural Language Processing
Hamza Khan, Asma Khan, Zainab Khan +3
In recent years, transformer-based models have shown state-of-the-art results for Natural Language Processing (NLP). In particular, the introduction of the BERT language model brou…