12 papers
Time vs. Layer: Locating Predictive Cues for Dysarthric Speech Descriptors in wav2vec 2.0
Natalie Engert, Dominik Wagner, Korbinian Riedhammer +1
Wav2vec 2.0 (W2V2) has shown strong performance in pathological speech analysis by effectively capturing the characteristics of atypical speech. Despite its success, it remains unc…
The PARLO Dementia Corpus: A German Multi-Center Resource for Alzheimer's Disease
Franziska Braun, Christopher Witzl, Florian Hönig +3
Early and accessible detection of Alzheimer's disease (AD) remains a major challenge, as current diagnostic methods often rely on costly and invasive biomarkers. Speech and languag…
Bias and Fairness in Self-Supervised Acoustic Representations for Cognitive Impairment Detection
Kashaf Gulzar, Korbinian Riedhammer, Elmar Nöth +2
Speech-based detection of cognitive impairment (CI) offers a promising non-invasive approach for early diagnosis, yet performance disparities across demographic and clinical subgro…
Reading Between the Waves: Robust Topic Segmentation Using Inter-Sentence Audio Features
Steffen Freisinger, Philipp Seeberger, Tobias Bocklet +1
Spoken content, such as online videos and podcasts, often spans multiple topics, which makes automatic topic segmentation essential for user navigation and downstream applications.…
Towards Multi-Level Transcript Segmentation: LoRA Fine-Tuning for Table-of-Contents Generation
Steffen Freisinger, Philipp Seeberger, Thomas Ranzenberger +2
Segmenting speech transcripts into thematic sections benefits both downstream processing and users who depend on written text for accessibility. We introduce a novel approach to hi…
Shared Multi-modal Embedding Space for Face-Voice Association
Christopher Simic, Korbinian Riedhammer, Tobias Bocklet
The FAME 2026 challenge comprises two demanding tasks: training face-voice associations combined with a multilingual setting that includes testing on languages on which the model w…