1 citations · 1 across the 4 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Orthogonal Negative Guidance in Attention Feature Space for Text-to-Image Generation
Jungmin Ko, Jungwon Park, Jimyeong Kim +3
Text-to-image (T2I) models have become increasingly capable of generating high-quality images. Yet, enforcing the explicit absence of a specified object or attribute remains a fund…
cs.CV2025
Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
Changin Choi, Wonseok Lee, Jungmin Ko +1
Knowledge-intensive visual question answering (VQA) requires external knowledge beyond image content, demanding precise visual grounding and coherent integration of visual and text…
cs.CV2024★ 1 cited
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
Wonkyun Kim, Changin Choi, Wonseok Lee +1
Stimulated by the sophisticated reasoning capabilities of recent Large Language Models (LLMs), a variety of strategies for bridging video modality have been devised. A prominent st…