activity
20202023
most citedTowards Multi-Scale Style Control for Expressive Speech Synthesis

4 citations · 6 across the 8 of their papers we have counts for

collaborators

8 papers

cs.SD2023

Text-Only Domain Adaptation for End-to-End Speech Recognition through Down-Sampling Acoustic Representation

Jiaxu Zhu, Weinan Tong, Yaoxun Xu +6

Mapping two modalities, speech and text, into a shared representation space, is a research topic of using text-only data to improve end-to-end automatic speech recognition (ASR) pe…

cs.SD2023

SememeASR: Boosting Performance of End-to-End Speech Recognition against Domain and Long-Tailed Data Shift with Sememe Semantic Knowledge

Jiaxu Zhu, Changhe Song, Zhiyong Wu +1

Recently, excellent progress has been made in speech recognition. However, pure data-driven approaches have struggled to solve the problem in domain-mismatch and long-tailed data.…

cs.SD2023

Improving Mandarin Prosodic Structure Prediction with Multi-level Contextual Information

Jie Chen, Changhe Song, Deyi Tuo +4

For text-to-speech (TTS) synthesis, prosodic structure prediction (PSP) plays an important role in producing natural and intelligible speech. Although inter-utterance linguistic in…

cs.SD2022

Towards Cross-speaker Reading Style Transfer on Audiobook Dataset

Xiang Li, Changhe Song, Xianhao Wei +3

Cross-speaker style transfer aims to extract the speech style of the given reference speech, which can be reproduced in the timbre of arbitrary target speakers. Existing methods on…

cs.CL2022

An End-to-end Chinese Text Normalization Model based on Rule-guided Flat-Lattice Transformer

Wenlin Dai, Changhe Song, Xiang Li +4

Text normalization, defined as a procedure transforming non standard words to spoken-form words, is crucial to the intelligibility of synthesized speech in text-to-speech system. R…

cs.SD20221 cited

Disentangleing Content and Fine-grained Prosody Information via Hybrid ASR Bottleneck Features for Voice Conversion

Xintao Zhao, Feng Liu, Changhe Song +4

Non-parallel data voice conversion (VC) have achieved considerable breakthroughs recently through introducing bottleneck features (BNFs) extracted by the automatic speech recogniti…