activity
20242026
collaborators

12 papers

cs.SD2026

LILAC: An Idempotent Neural Speech Codec

June Young Yi, Dongwook Lee, Jiheum Yeom +1

Neural Audio Codecs are widely adopted in speech generation and editing. However, existing neural audio codecs are not idempotent: across the paper's twelve baseline systems, every…

cs.CL2026

Learning When to Reason for Text-to-SQL via SFT and DPO

Soohyuk Jang, Jiheum Yeom, Nohil Park +4

Recent Text-to-SQL methods rely heavily on reasoning-centric paradigms such as Chain-of-Thought (CoT), achieving substantial gains on complex benchmarks at the cost of high inferen…

cs.CV2025

TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment

Kanghyun Baek, Sangyub Lee, Jin Young Choi +6

Despite recent advances, diffusion-based text-to-image models still struggle with accurate text rendering. Several studies have proposed fine-tuning or training-free refinement met…

cs.CV2025

DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy

Jaewoo Song, Jooyoung Choi, Kanghyun Baek +3

Despite recent text-to-image models achieving highfidelity text rendering, they still struggle with long or multiple texts due to diluted global attention. We propose DCText, a tra…

cs.CV2025

Negative-Guided Subject Fidelity Optimization for Zero-Shot Subject-Driven Generation

Chaehun Shin, Jooyoung Choi, Johan Barthelemy +2

We present Subject Fidelity Optimization (SFO), a novel comparative learning framework for zero-shot subject-driven generation that enhances subject fidelity. Existing supervised f…

cs.CV2025

Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator

Chaehun Shin, Jooyoung Choi, Heeseung Kim +1

Subject-driven text-to-image generation aims to produce images of a new subject within a desired context by accurately capturing both the visual characteristics of the subject and…