Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
VLMs Trace Without Tracking: Diagnosing Failures in Visual Path Following
Hyesoo Hong, Minsoo Kim, Wonje Jeung +3
Vision-language models (VLMs) achieve strong performance on multimodal benchmarks, but may still lack robust control over basic visual operations. We study \textit{line tracing}, w…
cs.CV2025
Q-Align: Alleviating Attention Leakage in Zero-Shot Appearance Transfer via Query-Query Alignment
Namu Kim, Wonbin Kweon, Minsoo Kim +1
We observe that zero-shot appearance transfer with large-scale image generation models faces a significant challenge: Attention Leakage. This challenge arises when the semantic map…