activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation

Jungmin Ko, Jungwon Park, Jimyeong Kim +3

Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object or attribute remains a fund…

cs.CV2026

Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering

Changin Choi, Wonseok Lee, Jungmin Ko +1

Knowledge-intensive visual question answering (VQA) requires external knowledge beyond image content, demanding precise visual grounding and coherent integration of visual and text…

cs.CV2026

Selective Aggregation of Attention Maps Improves Diffusion-Based Visual Interpretation

Jungwon Park, Jungmin Ko, Dongnam Byun +1

Numerous studies on text-to-image (T2I) generative models have utilized cross-attention maps to boost application performance and interpret model behavior. However, the distinct ch…

cs.CV2026

DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation

Dongnam Byun, Jungwon Park, Jungmin Ko +2

Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models stil…

cs.CV2025

ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation

Jimyeong Kim, Jungwon Park, Yeji Song +2

Rectified Flow text-to-image models surpass diffusion models in image quality and text alignment, but adapting ReFlow for real-image editing remains challenging. We propose a new r…

cs.CV2025

Cross-Attention Head Position Patterns Can Align with Human Visual Concepts in Text-to-Image Generative Models

Jungwon Park, Jungmin Ko, Dongnam Byun +2

Recent text-to-image diffusion models leverage cross-attention layers, which have been effectively utilized to enhance a range of visual generative tasks. However, our understandin…