activity
20182023
most citedAn Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning

20 citations · 45 across the 6 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL20235 cited

PolyVoice: Language Models for Speech to Speech Translation

Qianqian Dong, Zhiying Huang, Qiao Tian +15

We propose PolyVoice, a language model-based framework for speech-to-speech translation (S2ST) system. Our framework consists of two language models: a translation language model a…

cs.CL202214 cited

M3ST: Mix at Three Levels for Speech Translation

Xuxin Cheng, Qianqian Dong, Fengpeng Yue +3

How to solve the data scarcity problem for end-to-end speech-to-text translation (ST)? It's well known that data augmentation is an efficient method to improve performance for many…

cs.CL20222 cited

Leveraging Pseudo-labeled Data to Improve Direct Speech-to-Speech Translation

Qianqian Dong, Fengpeng Yue, Tom Ko +3

Direct Speech-to-speech translation (S2ST) has drawn more and more attention recently. The task is very challenging due to data scarcity and complex speech-to-speech mapping. In th…

cs.CL20214 cited

Exploring Machine Speech Chain for Domain Adaptation and Few-Shot Speaker Adaptation

Fengpeng Yue, Yan Deng, Lei He +1

Machine Speech Chain, which integrates both end-to-end (E2E) automatic speech recognition (ASR) and text-to-speech (TTS) into one circle for joint training, has been proven to be e…

cs.CL2018

An Investigation of Few-Shot Learning in Spoken Term Classification

Yangbin Chen, Tom Ko, Lifeng Shang +3

In this paper, we investigate the feasibility of applying few-shot learning algorithms to a speech task. We formulate a user-defined scenario of spoken term classification as a few…