10 papers
An Interpretable, Controllable Time-Varying IIR Denoiser for On-Device Assistive Hearing
Riccardo Rota, Kiril Ratmanski, Jozef Coldenhoff +1
We present TVF (Time-Varying Filtering), an interpretable, low-latency speech enhancement model for real-time, on-device assistive hearing. A lightweight neural controller predicts…
StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection
Zhentao Liu, Milos Cernak
The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive…
Calibration-Reasoning Framework for Descriptive Speech Quality Assessment
Elizaveta Kostenok, Mathieu Salzmann, Milos Cernak
Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-train…
Zimtohrli: An Efficient Psychoacoustic Audio Similarity Metric
Jyrki Alakuijala, Martin Bruse, Sami Boukortt +2
This paper introduces Zimtohrli, a novel, full-reference audio similarity metric designed for efficient and perceptually accurate quality assessment. In an era dominated by computa…
Shortcut Flow Matching for Speech Enhancement: Step-Invariant flows via single stage training
Naisong Zhou, Saisamarth Rajesh Phaye, Milos Cernak +4
Diffusion-based generative models have achieved state-of-the-art performance for perceptual quality in speech enhancement (SE). However, their iterative nature requires numerous Ne…
Low-latency Assistive Audio Enhancement for Neurodivergent People
Alexander Popescu, Rosie Frost, Milos Cernak
Neurodivergent people frequently experience decreased sound tolerance, with estimates suggesting it affects 50-70% of this population. This heightened sensitivity can provoke react…