4 papers
QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers
Kyobin Choo, Youngmin Kim, Hyunkyung Han +4
Video diffusion transformers (DiTs) generate high-fidelity and temporally coherent videos, yet motion control remains implicit, primarily relying on text prompts. As a result, achi…
Towards Continuous Sign Language Conversation from Isolated Signs
Youngmin Kim, Kyobin Choo, Jiwoo Park +4
Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction through spoken or written langua…
CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
Seonghyun Jin, Youngmin Kim, Sunwoo Park +1
Camera-conditioned video generation requires positional encoding that remains reliable under changes in camera motion, lens configuration, and scene structure. However, existing at…
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
Yeonkyung Lee, Dayun Ju, Youngmin Kim +2
Recent advancements in Video Large Language Models (VideoLLMs) have enabled strong performance across diverse multimodal video tasks. To reduce the high computational cost of proce…