5 papers
Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument
Rui Hong, Jana Košecká
Sign Language Production (SLP) is the task of generating avatar sign language motion from natural language text. The quality of the generated motion is typically evaluated by a mot…
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
Rui Hong, Shuxue Quan
When VLMs answer correctly, do they genuinely rely on visual information? We introduce a Tri-Layer Diagnostic Framework with three per-sample metrics: Latent Anomaly Detection, Vis…
Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation
Rui Hong, Jana Kosecka
Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a…
Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis
Rui Hong, Jana Kosecka
Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative…
Motion-Adaptive Temporal Attention for Lightweight Video Generation with Stable Diffusion
Rui Hong, Shuxue Quan
We present a motion-adaptive temporal attention mechanism for parameter-efficient video generation built upon frozen Stable Diffusion models. Rather than treating all video content…