12 papers
QWERTY: Training-Free Motion Control via Query-Warped Video Diffusion Transformers
Kyobin Choo, Youngmin Kim, Hyunkyung Han +4
Video diffusion transformers (DiTs) generate high-fidelity and temporally coherent videos, yet motion control remains implicit, primarily relying on text prompts. As a result, achi…
Anatomically Consistent TMJ Disc Segmentation via Semantic Anchoring and Clinical Priors
Dayun Ju, Chanyoung Kim, Sunyoung Jung +4
Segmenting the temporomandibular joint (TMJ) disc from MRI is essential for accurate diagnosis of internal derangement, yet it remains unreliable in practice due to its small size,…
How Noisy Poses Break Inverse Dynamics: Analysis and Mitigation for Video-Based Joint Torque Estimation
Donghyun Kim, Chanyoung Kim, Eunseo Jeong +2
Recent advances in monocular 3D human pose estimation enable accurate body tracking from video. However, translating these kinematic estimates into physical quantities, such as joi…
Towards Continuous Sign Language Conversation from Isolated Signs
Youngmin Kim, Kyobin Choo, Jiwoo Park +4
Sign language is the primary language for many Deaf and Hard-of-Hearing (DHH) signers, yet most conversational AI systems still mediate interaction through spoken or written langua…
WarmPrior: Straightening Flow-Matching Policies with Temporal Priors
Sinjae Kang, Chanyoung Kim, Kaixin Wang +2
Generative policies based on diffusion and flow matching have become a dominant paradigm for visuomotor robotic control. We show that replacing the standard Gaussian source distrib…
Rethinking Graph Convolution for 2D-to-3D Hand Pose Lifting
Chanyoung Kim, Donghyun Kim, Dong-Hyun Sim +2
Graph convolutional networks (GCNs) are widely used for 3D hand pose estimation, where the hand skeleton is encoded as a fixed adjacency graph. We revisit whether this is the most…