Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Peiyuan Zhu, Shaoan Xie, Zijian Li +5
Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled re…
cs.CV2025
Controllable Video Generation with Provable Disentanglement
Yifan Shen, Peiyuan Zhu, Zijian Li +6
Controllable video generation remains a significant challenge, despite recent advances in generating high-quality and consistent videos. Most existing methods for controlling video…
cs.CV2024
NurtureNet: A Multi-task Video-based Approach for Newborn Anthropometry
Yash Khandelwal, Mayur Arvind, Sriram Kumar +17
Malnutrition among newborns is a top public health concern in developing countries. Identification and subsequent growth monitoring are key to successful interventions. However, th…