17 papers
ProsoCodec: Prosody-Oriented Speech Codec for Voice Conversion
Jeongsoo Choi, Ji-Hoon Kim, Shujie Hu +1
Neural speech codecs efficiently compress speech and have become a foundation for speech generation, but they are typically learned as holistic representations that intertwine ling…
MamTra: A Hybrid Mamba-Transformer Backbone for Speech Synthesis
Tan Dat Nguyen, Sangmin Bae, Joon Son Chung +1
Despite the remarkable quality of LLM-based text-to-speech systems, their reliance on autoregressive Transformers leads to quadratic computational complexity, which severely limits…
Probing Cross-modal Information Hubs in Audio-Visual LLMs
Jihoo Jung, Chaeyoung Jung, Ji-Hoon Kim +1
Audio-visual large language models (AVLLMs) have recently emerged as a powerful architecture capable of jointly reasoning over audio, visual, and textual modalities. In AVLLMs, the…
Sommelier: Scalable Open Multi-turn Audio Pre-processing for Full-duplex Speech Language Models
Kyudan Jung, Jihwan Kim, Soyoon Kim +3
As the paradigm of AI shifts from text-based LLMs to Speech Language Models (SLMs), there is a growing demand for full-duplex systems capable of real-time, natural human-computer i…
SNAP: Speaker Nulling for Artifact Projection in Speech Deepfake Detection
Kyudan Jung, Jihwan Kim, Minwoo Lee +4
Recent advancements in text-to-speech technologies enable generating high-fidelity synthetic speech nearly indistinguishable from real human voices. While recent studies show the e…
SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
Tan Dat Nguyen, Jaehun Kim, Ji-Hoon Kim +3
The goal of this paper is to introduce SPADE, a framework for Structured Pruning and Adaptive Distillation for Efficient Large Language Model-based text-to-speech (LLM-TTS). Recent…