activity
20242026
collaborators

14 papers

eess.AS2026

DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech

Junwon Moon, Seungbeom Kim, Yejin Lee +4

Autoregressive (AR) text-to-speech (TTS) models generate discrete speech tokens sequentially, which makes inference slow and can degrade robustness, since local errors propagate to…

cs.SD2026

Mask2Flow-TSE: Two-Stage Target Speaker Extraction with Masking and Flow Matching

Junwon Moon, Seungbeom Kim, Hansol Park +4

Target speaker extraction (TSE) extracts the target speaker's voice from overlapping speech given a reference utterance. Existing masking-based approaches are lightweight and effec…

cs.CL2026

From Awareness to Adherence: Bridging the Context Gap in Spoken Dialogue Systems via Context-Aware Decoding

Che Hyun Lee, Heeseung Kim, Sungroh Yoon

Despite the success of end-to-end (E2E) spoken dialogue systems, maintaining strict context adherence in multi-round conversations remains a challenge. While prior works attribute…

cs.CL2026

NaturalFlow: Reducing Disruptive Pauses for Natural Speech Flow in Simultaneous Speech-to-Speech Translation

Dongwook Lee, Youngho Cho, Sangkwon Park +2

Simultaneous speech-to-speech translation aims to enable near-real-time communication by minimizing latency, offering a compelling, real-time alternative to the high latency of con…

cs.SD2026

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech

Yejin Lee, Junwon Moon, Hyoeun Kim +3

Codec-based autoregressive (AR) speech language models have achieved strong text-to-speech (TTS) quality by modeling speech as sequences of discrete audio tokens with large pretrai…

cs.CV2026

Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization

Yeongtak Oh, Dongwook Lee, Sangkwon Park +2

While multimodal large language models have advanced across text, image, and audio, personalization research has remained primarily vision-language, with unified omnimodal benchmar…