8 papers
Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding
Dohyun Kim, Sungjun Han, Hyungguk Kim +4
Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting in…
Compositional SVG Generation via VLM-Driven Hierarchical Semantic Parsing
Sehwan Park, Taehoon Kim, Geonhee Han +3
While Vision-Language Models (VLMs) excel at visual reasoning, generating structured, editable Scalable Vector Graphics (SVG) remains a fundamental challenge. Existing pipelines pr…
Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision
Dohyun Kim, Seungwoo Lyu, Seung Wook Kim +1
Diffusion models have achieved impressive results in generative tasks such as text-to-image synthesis, yet they often struggle to fully align outputs with nuanced user intent and m…
Image Diffusion Models Exhibit Emergent Temporal Propagation in Videos
Youngseo Kim, Dohyun Kim, Geonhee Han +1
Image diffusion models, though originally developed for image generation, implicitly capture rich semantic structures that enable various recognition and localization tasks beyond…
DGMO: Training-Free Audio Source Separation through Diffusion-Guided Mask Optimization
Geonyoung Lee, Geonhee Han, Paul Hongsuck Seo
Language-queried Audio Source Separation (LASS) enables open-vocabulary sound separation via natural language queries. While existing methods rely on task-specific training, we exp…
Random Conditioning with Distillation for Data-Efficient Diffusion Model Compression
Dohyun Kim, Sehwan Park, Geonhee Han +2
Diffusion models generate high-quality images through progressive denoising but are computationally intensive due to large model sizes and repeated sampling. Knowledge distillation…