From the 1 of 7 linked papers with an AI index.
7 papers
Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation
Joonyong Park, David M. Chan, Yuki Saito +1
The paper investigates how large audio-language models used as automatic judges for speech evaluation can exploit protocol-level shortcuts—relying on provided labels or reference d…
CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech
Dong Yang, Yuki Saito, Wataru Nakata +1
This paper introduces CraBERT, a pre-trained phoneme encoder (PPEnc) designed for efficient pre-training in text-to-speech (TTS). CraBERT employs a cascade-fusion architecture and…
Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech
Dong Yang, Yiyi Cai, Haoyu Zhang +2
Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues: heuristic schedulers requirin…
Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis
Dong Yang, Yiyi Cai, Yuki Saito +2
We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike con…
Analysing the Language of Neural Audio Codecs
Joonyong Park, Shinnosuke Takamichi, David M. Chan +3
This study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC…
Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model
Dong Yang, Yuki Saito, Takaaki Saeki +4
This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker emb…