activity
20182025
most citedAn Objective Evaluation Framework for Pathological Speech Synthesis

2 citations · 2 across the 4 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2026

Data Augmentation for Pathological Speech Enhancement

Mingchi Hou, Enno Hermann, Ina Kodrasi

The performance of state-of-the-art speech enhancement (SE) models considerably degrades for pathological speech due to atypical acoustic characteristics and limited data availabil…

eess.AS2025

Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech

Karl El Hajal, Enno Hermann, Sevada Hovsepyan +1

Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-…

eess.AS2025

Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR

Karl El Hajal, Enno Hermann, Ajinkya Kulkarni +1

Automatic speech recognition (ASR) systems are well known to perform poorly on dysarthric speech. Previous works have addressed this by speaking rate modification to reduce the mis…

eess.AS2024

kNN Retrieval for Simple and Effective Zero-Shot Multi-speaker Text-to-Speech

Karl El Hajal, Ajinkya Kulkarni, Enno Hermann +1

While recent zero-shot multi-speaker text-to-speech (TTS) models achieve impressive results, they typically rely on extensive transcribed speech datasets from numerous speakers and…

eess.AS20246 cited

Towards interfacing large language models with ASR systems using confidence measures and prompting

Maryam Naderi, Enno Hermann, Alexandre Nanchen +2

As large language models (LLMs) grow in parameter size and capabilities, such as interaction through prompting, they open up new ways of interfacing with automatic speech recogniti…

eess.AS2018

Multilingual and Unsupervised Subword Modeling for Zero-Resource Languages

Enno Hermann, Herman Kamper, Sharon Goldwater

Subword modeling for zero-resource languages aims to learn low-level representations of speech audio without using transcriptions or other resources from the target language (such…