4 papers
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
Junwon Moon, Seungbeom Kim, Yejin Lee +4
Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to…
Mask2Flow-TSE: Two-Stage Target Speaker Extraction with Masking and Flow Matching
Junwon Moon, Seungbeom Kim, Hansol Park +4
Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech given a reference utterance. Existing masking-based approaches are lightweight and effec…
TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech
Yejin Lee, Junwon Moon, Hyoeun Kim +3
Codec-based autoregressive (AR) speech language models have achieved strong text-to-speech (TTS) quality by modeling speech as sequences of discrete audio tokens with large pretrai…
Evaluating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries under Diverse Acoustic Conditions
Hansol Park, Hoseong Ahn, Junwon Moon +2
Hallucinations in multimodal models have been extensively studied using benchmarks that probe reliability in image-text query settings. However, the effect of spoken queries on mul…