2 papers
cs.SD2025
Room Impulse Response Generation Conditioned on Acoustic Parameters
Silvia Arellano, Chunghsin Yeh, Gautam Bhattacharya +1
The generation of room impulse responses (RIRs) using deep neural networks has attracted growing research interest due to its applications in virtual and augmented reality, audio p…
cs.SD2024
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
Gerard I. Gállego, Roy Fejgin, Chunghsin Yeh +2
Audio token modeling has become a powerful framework for speech synthesis, with two-stage approaches employing semantic tokens remaining prevalent. In this paper, we aim to simplif…