4 papers
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
Kanghyun Baek, Jaihyun Lew, Chaehun Shin +2
Multimodal Diffusion Transformers (MM-DiTs) have achieved remarkable progress in text-to-image generation, yet they frequently suffer from concept omission, where specified objects…
TextGuider: Training-Free Guidance for Text Rendering via Attention Alignment
Kanghyun Baek, Sangyub Lee, Jin Young Choi +6
Despite recent advances, diffusion-based text-to-image models still struggle with accurate text rendering. Several studies have proposed fine-tuning or training-free refinement met…
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
Jaewoo Song, Jooyoung Choi, Kanghyun Baek +3
Despite recent text-to-image models achieving highfidelity text rendering, they still struggle with long or multiple texts due to diluted global attention. We propose DCText, a tra…
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
Jaewoo Song, Daemin Park, Kanghyun Baek +4
Developing effective visual inspection models remains challenging due to the scarcity of defect data. While image generation models have been used to synthesize defect images, prod…