activity
20242026
collaborators

7 papers

eess.AS2026

An Evaluation Framework for Text-to-Speech Voice Reconstruction

Ariadna Sanchez, Christoph Minixhofer, Korin Richmond +3

Voice reconstruction using Text-to-Speech (TTS) offers a communication method for people with speech disorders, which aims to retain their speaker identity while improving intellig…

cs.SD2026

LISE : Listenable Interpretable Speaker Embeddings

Xiaoliang Wu, Chongxin Gan, Ke Liu +2

Deep neural network-based automatic speaker verification (ASV) systems achieve impressive performance but their embedding representations remain opaque, lacking a structured and pe…

eess.AS2026

The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs

Shree Harsha Bokkahalli Satish, Christoph Minixhofer, Maria Teleki +5

Speech Large Language Models (SpeechLLMs) process spoken input directly, retaining cues such as accent and perceived gender that were previously removed in cascaded pipelines. This…

cs.HC2026

From Seeing it to Experiencing it: Interactive Evaluation of Intersectional Voice Bias in Human-AI Speech Interaction

Shree Harsha Bokkahalli Satish, Maria Teleki, Christoph Minixhofer +3

SpeechLLMs process spoken language directly from audio, but accent and vocal identity cues can lead to biased behaviour. Current bias evaluations often miss how such bias manifests…

cs.SD2026

TTSDS2: Resources and Benchmark for Evaluating Human-Quality Text to Speech Systems

Christoph Minixhofer, Ondrej Klejch, Peter Bell

Evaluation of Text to Speech (TTS) systems is challenging and resource-intensive. Subjective metrics such as Mean Opinion Score (MOS) are not easily comparable between works. Objec…

cs.CL2025

Prosodic Structure Beyond Lexical Content: A Study of Self-Supervised Learning

Sarenne Wallbridge, Christoph Minixhofer, Catherine Lai +1

People exploit the predictability of lexical structures during text comprehension. Though predictable structure is also present in speech, the degree to which prosody, e.g. intonat…