9 papers
Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3
Neural audio codecs (NACs) have become popular for obtaining speech representations as discrete tokens. Beyond compression, discrete tokens can be used to train self-supervised lea…
UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation
Shunsuke Yoshida, Yu-Hua Chen, Satoru Fukayama
This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describe…
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encode…
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
Hitoshi Suda, Shinnosuke Takamichi, Satoru Fukayama
Perceived voice likability plays a crucial role in various social interactions, such as partner selection and advertising. A system that provides reference likable voice samples ta…
IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
Hitoshi Suda, Junya Koguchi, Shunsuke Yoshida +3
Japanese idol groups, comprising performers known as "idols," are an indispensable part of Japanese pop culture. They frequently appear in live concerts and television programs, en…
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
Masked token prediction has emerged as a powerful pre-training objective across language, vision, and speech, offering the potential to unify these diverse modalities through a sin…