4 papers
From Sign Language Generation to Humanoid Execution: Vision-Language Guided Retargeting with Collision Mitigation
Nabeela Khan, Bowen Wu, Runwu Shi +5
Recent sign language generation (SLG) systems increasingly output dense 3D body representations, which better preserve full-body kinematics and geometry for downstream embodiment o…
Unsupervised Single-Channel Speech Separation with Diffusion under Speaker-Embedding Guidance
Runwu Shi, Kai Li, Chang Li +5
Speech separation is a fundamental task in audio processing, typically addressed with fully supervised systems trained on paired mixtures. While effective, such systems typically r…
Multilingual Gloss-free Sign Language Translation: Towards Building a Sign Language Foundation Model
Sihan Tan, Taro Miyazaki, Kazuhiro Nakadai
Sign Language Translation (SLT) aims to convert sign language (SL) videos into spoken language text, thereby bridging the communication gap between the sign and the spoken communit…
Improvement in Sign Language Translation Using Text CTC Alignment
Sihan Tan, Taro Miyazaki, Nabeela Khan +1
Current sign language translation (SLT) approaches often rely on gloss-based supervision with Connectionist Temporal Classification (CTC), limiting their ability to handle non-mono…