activity
20192022
most citedTowards adversarial learning of speaker-invariant representation for speech emotion recognition

17 citations · 22 across the 4 of their papers we have counts for

collaborators

6 papers

cs.CL2022

Improving Speech-to-Speech Translation Through Unlabeled Text

Xuan-Phi Nguyen, Sravya Popuri, Changhan Wang +3

Direct speech-to-speech translation (S2ST) is among the most challenging problems in the translation paradigm due to the significant scarcity of S2ST data. While effort has been ma…

cs.CL20222 cited

Simple and Effective Unsupervised Speech Translation

Changhan Wang, Hirofumi Inaguma, Peng-Jen Chen +5

The amount of labeled data to train models for speech tasks is limited for most languages, however, the data scarcity is exacerbated for speech translation which requires labeled d…

cs.CL2020

Multilingual Speech Translation with Efficient Finetuning of Pretrained Models

Xian Li, Changhan Wang, Yun Tang +6

We present a simple yet effective approach to build multilingual speech-to-text (ST) translation by efficient transfer learning from pretrained speech encoder and text decoder. Our…

cs.CL2019

Zero-shot Text-to-SQL Learning with Auxiliary Task

Shuaichen Chang, Pengfei Liu, Yun Tang +3

Recent years have seen great success in the use of neural seq2seq models on the text-to-SQL task. However, little work has paid attention to how these models generalize to realisti…

eess.AS20193 cited

I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared Experiences

Kong Aik Lee, Ville Hautamaki, Tomi Kinnunen +43

The I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which…

eess.AS201917 cited

Towards adversarial learning of speaker-invariant representation for speech emotion recognition

Ming Tu, Yun Tang, Jing Huang +2

Speech emotion recognition (SER) has attracted great attention in recent years due to the high demand for emotionally intelligent speech interfaces. Deriving speaker-invariant repr…