5 papers
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
Jungmin Ko, Jungwon Park, Jimyeong Kim +3
Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object or attribute remains a fund…
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
Changin Choi, Wonseok Lee, Jungmin Ko +1
Knowledge-intensive visual question answering (VQA) requires external knowledge beyond image content, demanding precise visual grounding and coherent integration of visual and text…
Soft Head Selection for Injecting ICL-Derived Task Embeddings
Jungwon Park, Jimyeong Kim, Changin Choi +1
Large language models (LLMs) are commonly adapted to downstream tasks using parameter-efficient fine-tuning (PEFT) or in-context learning (ICL). Recently, ICL-driven embedding-base…
DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation
Dongnam Byun, Jungwon Park, Jungmin Ko +2
Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models stil…
Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning
Choi Changin, Lim Sungjun, Rhee Wonjong
Retrieval-augmented generation can improve audio captioning by incorporating relevant audio-text pairs from a knowledge base. Existing methods typically rely solely on the input au…