2 citations · 2 across the 2 of their papers we have counts for
Showing cs.SDShow all
2 papers · 1 filter
cs.SD2026
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Dongseong Hwang, Prasanth Yadla, Kaan Elgin +8
Siri Expressive Voices synthesize rich, configurable speech in real time and entirely on device, powered by AFM 3 Core Advanced, Apple's most powerful on-device foundation model. T…
cs.SD2016
Statistical Parametric Speech Synthesis Using Bottleneck Representation From Sequence Auto-encoder
Sivanand Achanta, KNRK Raju Alluri, Suryakanth V Gangashetty
In this paper, we describe a statistical parametric speech synthesis approach with unit-level acoustic representation. In conventional deep neural network based speech synthesis, t…