5 papers
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Peiyuan Zhu, Shaoan Xie, Zijian Li +5
Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled re…
What Must Generalist Agents Remember?
Khurram Yamin, Namrata Deka, Maitreyi Swaroop +3
This paper develops a formal account of what generalist agents must store in memory in order to act near-optimally across multiple environments and goals. It shows that when two do…
SEDGE: Structural Extrapolated Data Generation
Kun Zhang, Jiaqi Sun, Yiqing Li +3
This paper aims to address the challenge of data generation beyond the training data and proposes a framework for Structural Extrapolated Data GEneration (SEDGE) based on suitable…
On the Hardness of Conditional Independence Testing In Practice
Zheng He, Roman Pogodin, Yazhe Li +3
Tests of conditional independence (CI) underpin a number of important problems in machine learning and statistics, from causal discovery to evaluation of predictor fairness and out…
Controllable Video Generation with Provable Disentanglement
Yifan Shen, Peiyuan Zhu, Zijian Li +6
Controllable video generation remains a significant challenge, despite recent advances in generating high-quality and consistent videos. Most existing methods for controlling video…