activity
20192022
most citedAudio-visual Recognition of Overlapped speech for the LRS2 dataset

10 citations · 16 across the 7 of their papers we have counts for

collaborators

9 papers

cs.SD2022

Towards Expressive Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis

Shun Lei, Yixuan Zhou, Liyang Chen +3

Previous works on expressive speech synthesis mainly focus on current sentence. The context in adjacent sentences is neglected, resulting in inflexible speaking style for the same…

cs.SD2022

FullSubNet+: Channel Attention FullSubNet with Complex Spectrograms for Speech Enhancement

Jun Chen, Zilin Wang, Deyi Tuo +3

Previously proposed FullSubNet has achieved outstanding performance in Deep Noise Suppression (DNS) Challenge and attracted much attention. However, it still encounters issues such…

cs.SD20221 cited

Disentangleing Content and Fine-grained Prosody Information via Hybrid ASR Bottleneck Features for Voice Conversion

Xintao Zhao, Feng Liu, Changhe Song +4

Non-parallel data voice conversion (VC) have achieved considerable breakthroughs recently through introducing bottleneck features (BNFs) extracted by the automatic speech recogniti…

cs.SD2021

VAENAR-TTS: Variational Auto-Encoder based Non-AutoRegressive Text-to-Speech Synthesis

Hui Lu, Zhiyong Wu, Xixin Wu +4

This paper describes a variational auto-encoder based non-autoregressive text-to-speech (VAENAR-TTS) model. The autoregressive TTS (AR-TTS) models based on the sequence-to-sequence…

eess.AS20203 cited

Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams

Huirong Huang, Zhiyong Wu, Shiyin Kang +9

Generating 3D speech-driven talking head has received more and more attention in recent years. Recent approaches mainly have following limitations: 1) most speaker-independent meth…

eess.AS20202 cited

Transferring Source Style in Non-Parallel Voice Conversion

Songxiang Liu, Yuewen Cao, Shiyin Kang +5

Voice conversion (VC) techniques aim to modify speaker identity of an utterance while preserving the underlying linguistic information. Most VC approaches ignore modeling of the sp…