3 papers
cs.CV2026
Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation
Rui Hong, Jana Kosecka
Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a…
cs.CV2026
Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis
Rui Hong, Jana Kosecka
Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative…
cs.CV2026
Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion
Rui Hong, Shuxue Quan
We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content…