4 papers
Training-Free Pronunciation Transcription via Text-Constrained Acoustic Rescoring
Hikaru Asano, Yotaro Kubo, So Kuroki
Accurate and efficient pronunciation transcription is essential for preparing text-to-speech training data at scale. Existing approaches have different limitations: grapheme-to-pro…
Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance
Manato Yaguchi, Yotaro Kubo, Hikaru Asano +1
Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying…
Where-to-Unmask: Ground-Truth-Guided Unmasking Order Learning for Masked Diffusion Language Models
Hikaru Asano, Tadashi Kozuno, Kuniaki Saito +1
Masked Diffusion Language Models (MDLMs) generate text by iteratively filling masked tokens, requiring two coupled decisions at each step: which positions to unmask (where-to-unmas…
Self Iterative Label Refinement via Robust Unlabeled Learning
Hikaru Asano, Tadashi Kozuno, Yukino Baba
Recent advances in large language models (LLMs) have yielded impressive performance on various tasks, yet they often depend on high-quality feedback that can be costly. Self-refine…