14 papers
VTaMo: Video-Text Alignment Model for Sign Language Translation
Junyi Hu, Zhewen He, Haomian Huang +2
Sign language translation (SLT) converts continuous sign videos into spoken language text. Gloss-free approaches leverage pre-trained visual encoders and language models but rely o…
SignNet-1M: Large-Scale Multilingual Sign Language Video Dataset with Downstream Benchmarks
Zhewen He, Junyi Hu, Haomian Huang +3
Sign language models are typically trained on datasets captured under constrained conditions, with limited viewpoint, background, and signer-identity diversity, leading to poor rob…
One-shot Adaptation of Humanoid Whole-body Motion with Walking Priors
Hao Huang, Geeta Chandra Raju Bethala, Shuaihang Yuan +4
Whole-body humanoid motion represents a fundamental challenge in robotics, requiring balance, coordination, and adaptability to enable human-like behaviors. However, existing metho…
On Demographic Group Fairness Guarantees in Deep Learning
Yan Luo, Congcong Wen, Min Shi +3
We present a theoretical framework analyzing the relationship between data distributions and fairness guarantees in equitable deep learning. We establish novel bounds that account…
Humanoid Agent via Embodied Chain-of-Action Reasoning with Multimodal Foundation Models for Zero-Shot Loco-Manipulation
Congcong Wen, Geeta Chandra Raju Bethala, Yu Hao +8
Humanoid loco-manipulation, which integrates whole-body locomotion with dexterous manipulation, remains a fundamental challenge in robotics. Beyond whole-body coordination and bala…
CurveFlow: Curvature-Guided Flow Matching for Image Generation
Yan Luo, Drake Du, Hao Huang +2
Existing rectified flow models are based on linear trajectories between data and noise distributions. This linearity enforces zero curvature, which can inadvertently force the imag…