activity
20182022
most citedAdversarial Music: Real World Audio Adversary Against Wake-word Detection System

23 citations · 44 across the 15 of their papers we have counts for

collaborators

20 papers

cs.CL20221 cited

Textless Direct Speech-to-Speech Translation with Discrete Speech Representation

Xinjian Li, Ye Jia, Chung-Cheng Chiu

Research on speech-to-speech translation (S2ST) has progressed rapidly in recent years. Many end-to-end systems have been proposed and show advantages over conventional cascade sys…

cs.CL2022

ASR2K: Speech Recognition for Around 2000 Languages without Audio

Xinjian Li, Florian Metze, David R Mortensen +2

Most recent speech recognition models rely on large supervised datasets, which are unavailable for many low-resource languages. In this work, we present a speech recognition pipeli…

cs.SD20221 cited

On Adversarial Robustness of Large-scale Audio Visual Learning

Juncheng B Li, Shuhui Qu, Xinjian Li +2

As audio-visual systems are being deployed for safety-critical tasks such as surveillance and malicious content filtering, their robustness remains an under-studied area. Existing…

cs.LG2021

Multi-Faceted Hierarchical Multi-Task Learning for a Large Number of Tasks with Multi-dimensional Relations

Junning Liu, Zijie Xia, Yu Lei +2

There has been many studies on improving the efficiency of shared learning in Multi-Task Learning(MTL). Previous work focused on the "micro" sharing perspective for a small number…

cs.SD20211 cited

On Prosody Modeling for ASR+TTS based Voice Conversion

Wen-Chin Huang, Tomoki Hayashi, Xinjian Li +2

In voice conversion (VC), an approach showing promising results in the latest voice conversion challenge (VCC) 2020 is to first use an automatic speech recognition (ASR) model to t…

cs.CL2021

Phoneme Recognition through Fine Tuning of Phonetic Representations: a Case Study on Luhya Language Varieties

Kathleen Siminyu, Xinjian Li, Antonios Anastasopoulos +3

Models pre-trained on multiple languages have shown significant promise for improving speech recognition, particularly for low-resource languages. In this work, we focus on phoneme…