collaborators

10 papers

cs.SD2026

An Interpretable, Controllable Time-Varying IIR Denoiser for On-Device Assistive Hearing

Riccardo Rota, Kiril Ratmanski, Jozef Coldenhoff +1

We present TVF (Time-Varying Filtering), an interpretable, low-latency speech enhancement model for real-time, on-device assistive hearing. A lightweight neural controller predicts…

eess.AS2026

StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection

Zhentao Liu, Milos Cernak

The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive…

eess.AS2026

Calibration-Reasoning Framework for Descriptive Speech Quality Assessment

Elizaveta Kostenok, Mathieu Salzmann, Milos Cernak

Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-train…

eess.AS2025

Zimtohrli: An Efficient Psychoacoustic Audio Similarity Metric

Jyrki Alakuijala, Martin Bruse, Sami Boukortt +2

This paper introduces Zimtohrli, a novel, full-reference audio similarity metric designed for efficient and perceptually accurate quality assessment. In an era dominated by computa…

cs.SD2025

Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training

Naisong Zhou, Saisamarth Rajesh Phaye, Milos Cernak +4

Diffusion-based generative models have achieved state-of-the-art performance for perceptual quality in speech enhancement (SE). However, their iterative nature requires numerous Ne…

eess.AS2025

Low-latency Assistive Audio Enhancement for Neurodivergent People

Alexander Popescu, Rosie Frost, Milos Cernak

Neurodivergent people frequently experience decreased sound tolerance, with estimates suggesting it affects 50-70% of this population. This heightened sensitivity can provoke react…