activity
20242026
collaborators

6 papers

cs.SD2026

LILAC: An Idempotent Neural Speech Codec

June Young Yi, Dongwook Lee, Jiheum Yeom +1

Neural Audio Codecs are widely adopted in speech generation and editing. However, existing neural audio codecs are not idempotent: across the paper's twelve baseline systems, every…

cs.CL2026

Learning When to Reason for Text-to-SQL via SFT and DPO

Soohyuk Jang, Jiheum Yeom, Nohil Park +4

Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inferen…

cs.CL2025

EdiText: Controllable Coarse-to-Fine Text Editing with Diffusion Language Models

Che Hyun Lee, Heeseung Kim, Jiheum Yeom +1

We propose EdiText, a controllable text editing method that modifies the reference text to desired attributes at various scales. We integrate an SDEdit-based editing technique that…

cs.SD2025

Does Your Voice Assistant Remember? Analyzing Conversational Context Recall and Utilization in Voice Interaction Models

Heeseung Kim, Che Hyun Lee, Sangkwon Park +4

Recent advancements in multi-turn voice interaction models have improved user-model communication. However, while closed-source models effectively retain and recall past utterances…

cs.SD2024

NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers

Nohil Park, Heeseung Kim, Che Hyun Lee +3

We present NanoVoice, a personalized text-to-speech model that efficiently constructs voice adapters for multiple speakers simultaneously. NanoVoice introduces a batch-wise speaker…

cs.SD2024

VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance

Jiheum Yeom, Heeseung Kim, Jooyoung Choi +3

When applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, espec…