4 papers
Teaching an Agent to Sketch One Part at a Time
Xiaodan Du, Ruize Xu, David Yunis +2
We develop a method for producing vector sketches one part at a time. To do this, we train a multi-modal language model-based agent using a novel multi-turn process-reward reinforc…
SHuBERT: Self-Supervised Sign Language Representation Learning via Multi-Stream Cluster Prediction
Shester Gueuwou, Xiaodan Du, Greg Shakhnarovich +2
Sign language processing has traditionally relied on task-specific models, limiting the potential for transfer learning across tasks. Pre-training methods for sign language have ty…
SignMusketeers: An Efficient Multi-Stream Approach for Sign Language Translation at Scale
Shester Gueuwou, Xiaodan Du, Greg Shakhnarovich +1
A persistent challenge in sign language video processing, including the task of sign to written language translation, is how we learn representations of sign language in an effecti…
Generative Models: What Do They Know? Do They Know Things? Let's Find Out!
Xiaodan Du, Nicholas Kolkin, Greg Shakhnarovich +1
Generative models excel at mimicking real scenes, suggesting they might inherently encode important intrinsic scene properties. In this paper, we aim to explore the following key q…