collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them

Woojung Han, Seil Kang, Youngjun Jun +3

Image-to-Video diffusion models leverage input images to generate visually stunning content, yet frequently produce motion that violates physical laws. We reveal a surprising findi…

cs.CV2026

Real-Time Visual Attribution Streaming in Thinking Model

Seil Kang, Woojung Han, Junhyeok Kim +3

We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems…

cs.CV2026

ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting

Yeonkyung Lee, Dayun Ju, Youngmin Kim +2

Recent advancements in Video Large Language Models (VideoLLMs) have enabled strong performance across diverse multimodal video tasks. To reduce the high computational cost of proce…

cs.CV2026

Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers

Youngjun Jun, Seil Kang, Woojung Han +1

Video Diffusion Transformers (DiTs) have been synthesizing high-quality video with high fidelity from given text descriptions involving motion. However, understanding how Video DiT…

cs.CV2025

Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models

Jinyeong Kim, Seil Kang, Jiwoo Park +2

Large Vision-Language Models (LVLMs) answer visual questions by transferring information from images to text through a series of attention heads. While this image-to-text informati…

cs.CV2025

Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding

Seil Kang, Jinyeong Kim, Junhyeok Kim +1

Visual grounding seeks to localize the image region corresponding to a free-form text description. Recently, the strong multimodal capabilities of Large Vision-Language Models (LVL…