2 papers
cs.SD2026
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Dongseong Hwang, Prasanth Yadla, Kaan Elgin +8
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. T…
eess.AS2025
Switchboard-Affect: Emotion Perception Labels from Conversational Speech
Amrit Romana, Jaya Narain, Tien Dung Tran +4
Understanding the nuances of speech emotion dataset curation and labeling is essential for assessing speech emotion recognition (SER) model potential in real-world applications. Mo…