activity
20182022
collaborators

5 papers

eess.AS2022

Self-Supervised Audio-Visual Speech Representations Learning By Multimodal Self-Distillation

Jing-Xuan Zhang, Genshun Wan, Zhen-Hua Ling +3

In this work, we present a novel method, named AV2vec, for learning audio-visual speech representations by multimodal self-distillation. AV2vec has a student and a teacher module,…

eess.AS2019

Non-Parallel Sequence-to-Sequence Voice Conversion with Disentangled Linguistic and Speaker Representations

Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai

This paper presents a method of sequence-to-sequence (seq2seq) voice conversion using non-parallel training data. In this method, disentangled linguistic and speaker representation…

cs.SD2018

Improving Sequence-to-Sequence Acoustic Modeling by Adding Text-Supervision

Jing-Xuan Zhang, Zhen-Hua Ling, Yuan Jiang +3

This paper presents methods of making using of text supervision to improve the performance of sequence-to-sequence (seq2seq) voice conversion. Compared with conventional frame-to-f…

cs.SD2018

Sequence-to-Sequence Acoustic Modeling for Voice Conversion

Jing-Xuan Zhang, Zhen-Hua Ling, Li-Juan Liu +2

In this paper, a neural network named Sequence-to-sequence ConvErsion NeTwork (SCENT) is presented for acoustic modeling in voice conversion. At training stage, a SCENT model is es…

cs.CL2018

Forward Attention in Sequence-to-sequence Acoustic Modelling for Speech Synthesis

Jing-Xuan Zhang, Zhen-Hua Ling, Li-Rong Dai

This paper proposes a forward attention method for the sequenceto- sequence acoustic modeling of speech synthesis. This method is motivated by the nature of the monotonic alignment…