2 papers
cs.HC2025
SpeakEasy: Enhancing Text-to-Speech Interactions for Expressive Content Creation
Stephen Brade, Sam Anderson, Rithesh Kumar +2
Novice content creators often invest significant time recording expressive speech for social media videos. While recent advancements in text-to-speech (TTS) technology can generate…
eess.AS2025
DMOSpeech: Direct Metric Optimization via Distilled Diffusion Model in Zero-Shot Speech Synthesis
Yingahao Aaron Li, Rithesh Kumar, Zeyu Jin
Diffusion models have demonstrated significant potential in speech synthesis tasks, including text-to-speech (TTS) and voice cloning. However, their iterative denoising processes a…