4 papers · 1 filter
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
Sangha Park, Eunji Kim, Yeongtak Oh +2
Despite substantial progress in text-to-image generation, achieving precise text-image alignment remains challenging, particularly for prompts with rich compositional structure or…
DefectFill: Realistic Defect Generation with Inpainting Diffusion Model for Visual Inspection
Jaewoo Song, Daemin Park, Kanghyun Baek +4
Developing effective visual inspection models remains challenging due to the scarcity of defect data. While image generation models have been used to synthesize defect images, prod…
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
Mingi Jung, Saehyung Lee, Eunji Kim +1
Detailed image captioning is essential for tasks like data generation and aiding visually impaired individuals. High-quality captions require a balance between precision and recall…
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
Jaihyun Lew, Soohyuk Jang, Jaehoon Lee +6
Transformers, a groundbreaking architecture proposed for Natural Language Processing (NLP), have also achieved remarkable success in Computer Vision. A cornerstone of their success…