3 papers
cs.CV2025
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
Kanghyun Baek, Sangyub Lee, Jin Young Choi +6
Despite recent advances, diffusion-based text-to-image models still struggle with accurate text rendering. Several studies have proposed fine-tuning or training-free refinement met…
cs.LG2025
Enhanced Diffusion Sampling via Extrapolation with Multiple ODE Solutions
Jinyoung Choi, Junoh Kang, Bohyung Han
Diffusion probabilistic models (DPMs), while effective in generating high-quality samples, often suffer from high computational costs due to their iterative sampling process. To ad…
cs.CV2024
Hierarchical Visual Feature Aggregation for OCR-Free Document Understanding
Jaeyoo Park, Jin Young Choi, Jeonghyung Park +1
We present a novel OCR-free document understanding framework based on pretrained Multimodal Large Language Models (MLLMs). Our approach employs multi-scale visual features to effec…