6 papers · 1 filter
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
Tan Dat Nguyen, Sangmin Bae, Joon Son Chung +1
Despite the remarkable quality of LLM-based text-to-speech systems, their reliance on autoregressive Transformers leads to quadratic computational complexity, which severely limits…
Plug-and-Steer: Decoupling Separation and Selection in Audio-Visual Target Speaker Extraction
Doyeop Kwak, Suyeon Lee, Joon Son Chung
The goal of this paper is to provide a new perspective on audio-visual target speaker extraction (AV-TSE) by decoupling separation and target selection. Conventional AV-TSE systems…
LRS-VoxMM: A benchmark for in-the-wild audio-visual speech recognition
Doyeop Kwak, Jeongsoo Choi, Suyeon Lee +1
We introduce LRS-VoxMM, an in-the-wild benchmark for audio-visual speech recognition (AVSR). The benchmark is derived from VoxMM, a dataset of diverse real-world spoken conversatio…
MAGE: A Coarse-to-Fine Speech Enhancer with Masked Generative Model
The Hieu Pham, Tan Dat Nguyen, Phuong Thanh Tran +2
Speech enhancement remains challenging due to the trade-off between efficiency and perceptual quality. In this paper, we introduce MAGE, a Masked Audio Generative Enhancer that adv…
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
Tan Dat Nguyen, Jaehun Kim, Ji-Hoon Kim +3
The goal of this paper is to introduce SPADE, a framework for Structured Pruning and Adaptive Distillation for Efficient Large Language Model-based text-to-speech (LLM-TTS). Recent…
SCORE: Scaling audio generation using Standardized COmposite REwards
Jaemin Jung, Jaehun Kim, Inkyu Shin +1
The goal of this paper is to enhance Text-to-Audio generation at inference, focusing on generating realistic audio that precisely aligns with text prompts. Despite the rapid advanc…