From the 2 of 61 linked papers with an AI index.
61 papers
Amortized Moment Matching for Visual Generation
Wenze Liu, Xintao Wang, Pengfei Wan +1
The paper introduces amortized moment matching, using neural networks to learn data moments as training signals, and proposes the Amortized Fréchet Distance loss to improve one-ste…
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
Kaiqi Liu, Yunyao Mao, Ziqi Cai +8
MAVIN is a framework for generating multi-shot audio‑visual content with fine‑grained narrative control, using boundary‑aware attention to align temporal segments and ID‑aware prop…
UniVideo: Unified Understanding, Generation, and Editing for Videos
Cong Wei, Quande Liu, Zixuan Ye +5
Unified multimodal models have shown promising results in multimodal content generation and editing but remain largely limited to the image domain. In this work, we present UniVide…
HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control
Yushuo Chen, Xiaoyu Shi, Xiaoshi Wu +3
We present HandsOnWorld, a framework for hand-controlled egocentric video generation that learns directly from unconstrained monocular video. Prior generators depend on 3D hand ann…
AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
Yu Li, Menghan Xia, Gongye Liu +8
Despite being a pivotal frontier, interactive world modeling remains underexplored in terms of the versatile controllability required by practical scenarios. To bridge this gap, we…
VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
Minghong Cai, Qiulin Wang, Zongli Ye +7
Existing controllable video generation methods are typically designed for rigid, task-specific settings, such as first-frame image-to-video, inpainting, or interpolation, treating…