From the 2 of 7 linked papers with an AI index.
7 papers
Dissecting Sensitivity to Training Language in Self-Supervised Speech Learning Using Neural Audio Codec Tokens
Daigo Takizawa, Tomohiko Nakamura, Samuele Cornell +3
The paper investigates how the language used to train neural audio codecs and self‑supervised speech models affects performance, finding that codec training language has little imp…
OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
OpenBEATs is an open-source framework that extends the BEATs audio encoder with multi-domain masked-token pretraining, achieving state-of-the-art results on a wide range of audio t…
UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation
Shunsuke Yoshida, Yu-Hua Chen, Satoru Fukayama
This work investigates the effect of batch sampling strategies during training for text-to-audio music generation under low-data and small-scale model settings. This paper describe…
The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge
Shikhar Bharadwaj, Samuele Cornell, Kwanghee Choi +4
This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encode…
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
Hitoshi Suda, Shinnosuke Takamichi, Satoru Fukayama
Perceived voice likability plays a crucial role in various social interactions, such as partner selection and advertising. A system that provides reference likable voice samples ta…
IdolSongsJp Corpus: A Multi-Singer Song Corpus in the Style of Japanese Idol Groups
Hitoshi Suda, Junya Koguchi, Shunsuke Yoshida +3
Japanese idol groups, comprising performers known as "idols," are an indispensable part of Japanese pop culture. They frequently appear in live concerts and television programs, en…