works on

From the 1 of 26 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.SDShow all

17 papers · 1 filter

cs.SD2026

Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation

Joonyong Park, David M. Chan, Yuki Saito +1

The paper investigates how large audio-language models used as automatic judges for speech evaluation can exploit protocol-level shortcuts—relying on provided labels or reference d…

cs.SD2026

DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio

Wataru Nakata, Yuki Saito, Kazuki Yamauchi +2

Full-duplex dialogue audio, in which each speaker is recorded on a separate track, is an important resource for spoken dialogue research, but is difficult to collect at scale. Most…

cs.SD2026

Geneses: Unified Generative Speech Enhancement and Separation

Kohei Asai, Wataru Nakata, Yuki Saito +1

Real-world audio recordings often contain multiple speakers and various degradations, which limit both the quantity and quality of speech data available for building state-of-the-a…

cs.SD2026

Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing

Wataru Nakata, Yuki Saito, Yota Ueda +1

Large-scale text-to-speech (TTS) systems are limited by the scarcity of clean, multilingual recordings. We introduce Sidon, a fast, open-source speech restoration model that conver…

cs.SD2026

Dissecting Performance Degradation in Audio Source Separation under Sampling Frequency Mismatch

Kanami Imamura, Tomohiko Nakamura, Kohei Yatabe +1

Audio processing methods based on deep neural networks are typically trained at a single sampling frequency (SF). To handle untrained SFs, signal resampling is commonly employed, b…

cs.SD2026

DistilMOS: Layer-Wise Self-Distillation For Self-Supervised Learning Model-Based MOS Prediction

Jianing Yang, Wataru Nakata, Yuki Saito +1

With the advancement of self-supervised learning (SSL), fine-tuning pretrained SSL models for mean opinion score (MOS) prediction has achieved state-of-the-art performance. However…