activity
20242026
collaborators
Showing eess.ASShow all

8 papers · 1 filter

eess.AS2026

ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation

Shuntaro Suzuki, Kento Tokura, Daichi Yashima +3

Text-to-audio (TTA) generation, synthesizing audio from natural language, has been widely studied for its ability to capture precise user intent. To effectively advance TTA models,…

eess.AS2025

Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora

Hitoshi Suda, Shinnosuke Takamichi, Satoru Fukayama

Perceived voice likability plays a crucial role in various social interactions, such as partner selection and advertising. A system that provides reference likable voice samples ta…

eess.AS2025

Text-To-Speech Synthesis In The Wild

Jee-weon Jung, Wangyou Zhang, Soumi Maiti +11

Traditional Text-to-Speech (TTS) systems rely on studio-quality speech recorded in controlled settings.a Recently, an effort known as noisy-TTS training has emerged, aiming to util…

eess.AS2024

DNN-based ensemble singing voice synthesis with interactions between singers

Hiroaki Hyodo, Shinnosuke Takamichi, Tomohiko Nakamura +2

We propose a singing voice synthesis (SVS) method for a more unified ensemble singing voice by modeling interactions between singers. Most existing SVS methods aim to synthesize a…

eess.AS2024

BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Detai Xin, Xu Tan, Shinnosuke Takamichi +1

We present BigCodec, a low-bitrate neural speech codec. While recent neural speech codecs have shown impressive progress, their performance significantly deteriorates at low bitrat…

eess.AS2024

SaSLaW: Dialogue Speech Corpus with Audio-visual Egocentric Information Toward Environment-adaptive Dialogue Speech Synthesis

Osamu Take, Shinnosuke Takamichi, Kentaro Seki +2

This paper presents SaSLaW, a spontaneous dialogue speech corpus containing synchronous recordings of what speakers speak, listen to, and watch. Humans consider the diverse environ…