20 papers
Do Music Foundation Models Embed Pitch in Helical Structure?
Hayato Yagi, Shinnosuke Takamichi, Rin Sato +2
This study analyzes the intermediate representations of music foundation models (MFMs) and reports the geometric structures used to represent pitch information. By inputting isolat…
On the Effect of Segmentation Width and Cluster Size on Speech Resynthesis and Continuation in Generative Spoken Language Models
Shunsuke Kando, Wataru Nakata, Shinnosuke Takamichi +1
Generative Spoken Language Modeling (GSLM) enables text-free speech modeling by training language models (LMs) using discrete speech representations instead of textual transcriptio…
Do speech foundation models perceive speaker similarity as humans do?
Minoru Kishi, Hayato Yagi, Shinnosuke Takamichi +1
This study presents a comparative analysis between the speaker embeddings of speech foundation models and human subjective perception of speaker similarity. Human listeners have th…
ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation
Shuntaro Suzuki, Kento Tokura, Daichi Yashima +3
Text-to-audio (TTA) generation, synthesizing audio from natural language, has been widely studied for its ability to capture precise user intent. To effectively advance TTA models,…
Low-Latency Real-Time Audio Game Commentary System via LLM-Based Parallel Text Generation
Ryota Kawamatsu, Anum Afzal, Yuki Saito +5
We present a low-latency real-time audio game commentary system that generates spoken commentary directly from live gameplay video. In this end-to-end setting, a key bottleneck is…
Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN
Toranosuke Manabe, Yuto Shibata, Shinnosuke Takamichi +1
Deep learning models have improved sign language-to-text translation and made it easier for non-signers to understand signed messages. When the goal is spoken communication, a naiv…