From the 1 of 26 linked papers with an AI index.
17 papers · 1 filter
Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation
Joonyong Park, David M. Chan, Yuki Saito +1
The paper investigates how large audio-language models used as automatic judges for speech evaluation can exploit protocol-level shortcuts—relying on provided labels or reference d…
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
Wataru Nakata, Yuki Saito, Kazuki Yamauchi +2
Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most…
Geneses: Unified Generative Speech Enhancement and Separation
Kohei Asai, Wataru Nakata, Yuki Saito +1
Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-a…
Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing
Wataru Nakata, Yuki Saito, Yota Ueda +1
Large-scale text-to-speech (TTS) systems are limited by the scarcity of clean, multilingual recordings. We introduce Sidon, a fast, open-source speech restoration model that conver…
Dissecting Performance Degradation in Audio Source Separation under Sampling Frequency Mismatch
Kanami Imamura, Tomohiko Nakamura, Kohei Yatabe +1
Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, b…
DistilMOS: Layer-Wise Self-Distillation For Self-Supervised Learning Model-Based MOS Prediction
Jianing Yang, Wataru Nakata, Yuki Saito +1
With the advancement of self-supervised learning (SSL), fine-tuning pretrained SSL models for mean opinion score (MOS) prediction has achieved state-of-the-art performance. However…