8 papers · 1 filter
StreamMark: A Deep Learning-Based Semi-Fragile Audio Watermarking for Proactive Deepfake Detection
Zhentao Liu, Milos Cernak
The rapid advancement of generative AI has made it increasingly challenging to distinguish between deepfake audio and authentic human speech. To overcome the limitations of passive…
Calibration-Reasoning Framework for Descriptive Speech Quality Assessment
Elizaveta Kostenok, Mathieu Salzmann, Milos Cernak
Explainable speech quality assessment requires moving beyond Mean Opinion Scores (MOS) to analyze underlying perceptual dimensions. To address this, we introduce a novel post-train…
Zimtohrli: An Efficient Psychoacoustic Audio Similarity Metric
Jyrki Alakuijala, Martin Bruse, Sami Boukortt +2
This paper introduces Zimtohrli, a novel, full-reference audio similarity metric designed for efficient and perceptually accurate quality assessment. In an era dominated by computa…
Low-latency Assistive Audio Enhancement for Neurodivergent People
Alexander Popescu, Rosie Frost, Milos Cernak
Neurodivergent people frequently experience decreased sound tolerance, with estimates suggesting it affects 50-70% of this population. This heightened sensitivity can provoke react…
OpenACE: An Open Benchmark for Evaluating Audio Coding Performance
Jozef Coldenhoff, Niclas Granqvist, Milos Cernak
Audio and speech coding lack unified evaluation and open-source testing. Many candidate systems were evaluated on proprietary, non-reproducible, or small data, and machine learning…
DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic Regeneration
Sanberk Serbest, Tijana Stojkovic, Milos Cernak +1
In this work, we propose a full-band real-time speech enhancement system with GAN-based stochastic regeneration. Predictive models focus on estimating the mean of the target distri…