From the 1 of 22 linked papers with an AI index.
22 papers
Controlling Motion Transfer in Diffusion Transformers via Attention Heads
Sunyoung Jung, Jiwoo Park, Yoonseok Choi +3
The paper studies how individual attention heads in diffusion transformer models handle motion and spatial structure, and introduces a head-aware method to control motion transfer…
QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers
Kyobin Choo, Youngmin Kim, Hyunkyung Han +4
Video diffusion transformers (DiTs) generate high-fidelity and temporally coherent videos, yet motion control remains implicit, primarily relying on text prompts. As a result, achi…
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
Sujung Hong, Chanyong Yoon, Seong Jae Hwang
Large diffusion vision-language models (LDVLMs) have recently emerged as a promising alternative to autoregressive models, enabling parallel decoding for efficient inference and le…
Towards Continuous Sign Language Conversation from Isolated Signs
Youngmin Kim, Kyobin Choo, Jiwoo Park +4
Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction through spoken or written langua…
Rethinking Graph Convolution for 2D-to-3D Hand Pose Lifting
Chanyoung Kim, Donghyun Kim, Dong-Hyun Sim +2
Graph convolutional networks (GCNs) are widely used for 3D hand pose estimation, where the hand skeleton is encoded as a fixed adjacency graph. We revisit whether this is the most…
Real-Time Visual Attribution Streaming in Thinking Model
Seil Kang, Woojung Han, Junhyeok Kim +3
We present an amortized framework for real-time visual attribution streaming in multimodal thinking models. When these models generate code from a screenshot or solve math problems…