activity
20182022
most citedQuantifying Bias in Automatic Speech Recognition

55 citations · 68 across the 11 of their papers we have counts for

collaborators

20 papers

cs.SD20223 cited

Manipulation of oral cancer speech using neural articulatory synthesis

Bence Mark Halpern, Teja Rebernik, Thomas Tienkamp +5

We present an articulatory synthesis framework for the synthesis and manipulation of oral cancer speech for clinical decision making and alleviation of patient stress. Objective an…

cs.CL2022

Modelling word learning and recognition using visually grounded speech

Danny Merkx, Sebastiaan Scholten, Stefan L. Frank +2

Background: Computational models of speech recognition often assume that the set of target words is already given. This implies that these models do not learn to recognise speech f…

cs.SD2022

Discovering Phonetic Inventories with Crosslingual Automatic Speech Recognition

Piotr Żelasko, Siyuan Feng, Laureano Moro Velazquez +5

The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a writt…

cs.SD2022

The Effectiveness of Time Stretching for Enhancing Dysarthric Speech for Improved Dysarthric Speech Recognition

Luke Prananta, Bence Mark Halpern, Siyuan Feng +1

In this paper, we investigate several existing and a new state-of-the-art generative adversarial network-based (GAN) voice conversion method for enhancing dysarthric speech for imp…

cs.SD20213 cited

Towards Identity Preserving Normal to Dysarthric Voice Conversion

Wen-Chin Huang, Bence Mark Halpern, Lester Phillip Violeta +2

We present a voice conversion framework that converts normal speech into dysarthric speech while preserving the speaker identity. Such a framework is essential for (1) clinical dec…

cs.CV2021

AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person

Xinsheng Wang, Qicong Xie, Jihua Zhu +2

Automatically generating videos in which synthesized speech is synchronized with lip movements in a talking head has great potential in many human-computer interaction scenarios. I…