3 papers
cs.CV2026
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
Hyesoo Hong, Minsoo Kim, Wonje Jeung +3
Vision-language models (VLMs) achieve strong performance on multimodal benchmarks, but may still lack robust control over basic visual operations. We study \textit{line tracing}, w…
cs.CV2025
Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment
Namu Kim, Wonbin Kweon, Minsoo Kim +1
We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic map…
cs.IR2024
Controlling Diversity at Inference: Guiding Diffusion Recommender Models with Targeted Category Preferences
Gwangseok Han, Wonbin Kweon, Minsoo Kim +1
Diversity control is an important task to alleviate bias amplification and filter bubble problems. The desired degree of diversity may fluctuate based on users' daily moods or busi…