Showing 2024Show all
3 papers · 1 filter
cs.CV2024
GEXIA: Granularity Expansion and Iterative Approximation for Scalable Multi-grained Video-language Learning
Yicheng Wang, Zhikang Zhang, Jue Wang +6
In various video-language learning tasks, the challenge of achieving cross-modality alignment with multi-grained data persists. We propose a method to tackle this challenge from tw…
cs.CV2024
DiffSign: AI-Assisted Generation of Customizable Sign Language Videos With Enhanced Realism
Sudha Krishnamurthy, Vimal Bhat, Abhinav Jain
The proliferation of several streaming services in recent years has now made it possible for a diverse audience across the world to view the same media content, such as movies or T…
cs.CV2024
Text-Guided Video Masked Autoencoder
David Fan, Jue Wang, Shuai Liao +3
Recent video masked autoencoder (MAE) works have designed improved masking algorithms focused on saliency. These works leverage visual cues such as motion to mask the most salient…