5 papers
Building a Precise Video Language with Human-AI Oversight
Zhiqiu Lin, Chancharik Mitra, Siyuan Cen +13
Video-language models (VLMs) learn to reason about the dynamic visual world through natural language. We introduce a suite of open datasets, benchmarks, and recipes for scalable ov…
Physion-Eval: Evaluating Physical Realism in Generated Video via Human Reasoning
Qin Zhang, Peiyu Jing, Hong-Xing Yu +7
Video generation models are increasingly used as world simulators for storytelling, simulation, and embodied AI. As these models advance, a key question arises: do generated videos…
Ctrl-VI: Controllable Video Synthesis via Variational Inference
Haoyi Duan, Yunzhi Zhang, Yilun Du +1
Many video workflows benefit from a mixture of user controls with varying granularity, from exact 4D object trajectories and camera paths to coarse text prompts, while existing vid…
Product of Experts for Visual Generation
Yunzhi Zhang, Carson Murtuza-Lanier, Zizhang Li +2
Modern neural models capture rich priors and have complementary knowledge over shared data domains, e.g., images and videos. Integrating diverse knowledge from multiple sources --…
Few-Shot Task Learning through Inverse Generative Modeling
Aviv Netanyahu, Yilun Du, Antonia Bronars +4
Learning the intents of an agent, defined by its goals or motion style, is often extremely challenging from just a few examples. We refer to this problem as task concept learning a…