2 papers
eess.AS2024
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
John Janiczek, Dading Chong, Dongyang Dai +4
A text-to-speech (TTS) model trained to reconstruct speech given text tends towards predictions that are close to the average characteristics of a dataset, failing to model the var…
eess.AS2023
Fixed-point quantization aware training for on-device keyword-spotting
Sashank Macha, Om Oza, Alex Escott +5
Fixed-point (FXP) inference has proven suitable for embedded devices with limited computational resources, and yet model training is continually performed in floating-point (FLP).…