activity
20182026
most citedAn Encoder-Decoder Based Audio Captioning System With Transfer and Reinforcement Learning

20 citations · 47 across the 9 of their papers we have counts for

collaborators
Showing 2023Show all

6 papers · 1 filter

cs.CL2023

Recent Advances in Direct Speech-to-text Translation

Chen Xu, Rong Ye, Qianqian Dong +5

Recently, speech-to-text translation has attracted more and more attention and many studies have emerged rapidly. In this paper, we present a comprehensive survey on direct speech…

cs.SD2023

MOSPC: MOS Prediction Based on Pairwise Comparison

Kexin Wang, Yunlong Zhao, Qianqian Dong +2

As a subjective metric to evaluate the quality of synthesized speech, Mean opinion score~(MOS) usually requires multiple annotators to score the same speech. Such an annotation app…

cs.CL20235 cited

PolyVoice: Language Models for Speech to Speech Translation

Qianqian Dong, Zhiying Huang, Qiao Tian +15

We propose PolyVoice, a language model-based framework for speech-to-speech translation (S2ST) system. Our framework consists of two language models: a translation language model a…

cs.CL2023

CTC-based Non-autoregressive Speech Translation

Chen Xu, Xiaoqian Liu, Xiaowen Liu +9

Combining end-to-end speech translation (ST) and non-autoregressive (NAR) generation is promising in language and speech processing for their advantages of less error propagation a…

cs.CL20232 cited

DUB: Discrete Unit Back-translation for Speech Translation

Dong Zhang, Rong Ye, Tom Ko +2

How can speech-to-text translation (ST) perform as well as machine translation (MT)? The key point is to bridge the modality gap between speech and text so that useful MT technique…

eess.AS2023

WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research

Xinhao Mei, Chutong Meng, Haohe Liu +6

The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years. However, researchers face challenges due to the costly and time-consuming col…