6 papers
Conditional Collapse in Sign Language Production: A Diagnostic and a Scaling Argument
Rui Hong, Jana Košecká
Sign Language Production (SLP) is the task of generating avatar sign language motion from natural language text. The quality of the generated motion is typically evaluated by a mot…
Gesture-Aware Pretraining and Token Fusion for 3D Hand Pose Estimation
Rui Hong, Jana Kosecka
Estimating 3D hand pose from monocular RGB images is fundamental for applications in AR/VR, human-computer interaction, and sign language understanding. In this work we focus on a…
Toward Phonology-Guided Sign Language Motion Generation: A Diffusion Baseline and Conditioning Analysis
Rui Hong, Jana Kosecka
Generating natural, correct, and visually smooth 3D avatar sign language motion conditioned on the text inputs continues to be very challenging. In this work, we train a generative…
Compositional Image-Text Matching and Retrieval by Grounding Entities
Madhukar Reddy Vongala, Saurabh Srivastava, Jana Košecká
Vision-language pretraining on large datasets of images-text pairs is one of the main building blocks of current Vision-Language Models. While with additional training, these model…
Structured Spatial Reasoning with Open Vocabulary Object Detectors
Negar Nejatishahidin, Madhukar Reddy Vongala, Jana Kosecka
Reasoning about spatial relationships between objects is essential for many real-world robotic tasks, such as fetch-and-delivery, object rearrangement, and object search. The abili…
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
Pooya Fayyazsanavi, Antonios Anastasopoulos, Jana Košecká
Sign language translation from video to spoken text presents unique challenges owing to the distinct grammar, expression nuances, and high variation of visual appearance across dif…