works on

From the 1 of 24 linked papers with an AI index.

activity
20242026
collaborators
Showing eess.ASShow all

7 papers · 1 filter

eess.AS2026

CraBERT: Efficient Phoneme Encoder Pre-Training via Cascade Fusion of Subword Representations for Text-to-Speech

Dong Yang, Yuki Saito, Wataru Nakata +1

This paper introduces CraBERT, a pre-trained phoneme encoder (PPEnc) designed for efficient pre-training in text-to-speech (TTS). CraBERT employs a cascade-fusion architecture and…

eess.AS2026

Fast Multichannel NMF with Block-Diagonal Spatial Covariance Matrices for Efficient Blind Source Separation Using Distributed Microphone Arrays

Hirotaka Nishikori, Nobutaka Ito, Kouei Yamaoka +2

Distributed microphone arrays composed of multiple subarrays enable blind source separation over a wide spatial area. Directly applying fast multichannel nonnegative matrix factori…

eess.AS2026

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech

Dong Yang, Yiyi Cai, Haoyu Zhang +2

Metric-induced discrete flow matching (MI-DFM) exploits token-latent geometry for discrete generation, but its practical use is limited by two issues: heuristic schedulers requirin…

eess.AS2026

Human-CLAP: Human-perception-based contrastive language-audio pretraining

Taisei Takano, Yuki Okamoto, Yusuke Kanamori +3

Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, h…

eess.AS2025

Shallow Flow Matching for Coarse-to-Fine Text-to-Speech Synthesis

Dong Yang, Yiyi Cai, Yuki Saito +2

We propose Shallow Flow Matching (SFM), a novel mechanism that enhances flow matching (FM)-based text-to-speech (TTS) models within a coarse-to-fine generation paradigm. Unlike con…

eess.AS2025

Speaker-Conditioned Phrase Break Prediction for Text-to-Speech with Phoneme-Level Pre-trained Language Model

Dong Yang, Yuki Saito, Takaaki Saeki +4

This paper advances phrase break prediction (also known as phrasing) in multi-speaker text-to-speech (TTS) systems. We integrate speaker-specific features by leveraging speaker emb…