activity
20242026
collaborators

9 papers

eess.AS2026

Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces

Wangzixi Zhou, Bagus Tris Atmaja, Sakriani Sakti

The rise of conversational AI has increased interest in emotional Text-to-Speech (TTS). Most systems rely on discrete emotion labels, which fail to capture the nuanced nature of hu…

eess.AS2026

WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models

Wangzixi Zhou, Takuma Okamoto, Yamato Ohtani +2

Most neural vocoders are limited to one type: either GAN or diffusion-based. While state-of-the-art models like Vocos and WaveNeXt use powerful ConvNeXt-based generators, they have…

eess.AS2026

Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control

Wangzixi Zhou, Bagus Tris Atmaja, Sakriani Sakti

While current emotional Text-to-Speech (TTS) models have successfully controlled verbal prosody, they often ignore non-verbal vocalizations (NVs), which are essential for authentic…

cs.CL2026

SimulSense: Sense-Driven Interpreting for Efficient Simultaneous Speech Translation

Haotian Tan, Hiroki Ouchi, Sakriani Sakti

How to make human-interpreter-like read/write decisions for simultaneous speech translation (SimulST) systems? Current state-of-the-art systems formulate SimulST as a multi-turn di…

cs.SD2025

Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller Identification

Bin Wu, Shinnosuke Takamichi, Sakriani Sakti +1

The marmoset, a highly vocal primate, is a key model for studying social-communicative behavior. Unlike human speech, marmoset vocalizations are less structured, highly variable, a…

cs.SD2025

SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition

Yuta Hirano, Sakriani Sakti

We propose Speaker-Conditioned Serialized Output Training (SC-SOT), an enhanced SOT-based training for E2E multi-talker ASR. We first probe how SOT handles overlapped speech, and w…