audio-visual generation 1identity consistency 1multi-shot video synthesis 1narrative control 1temporal alignment 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.CV2026
MAVIN: Multi-Shot Audio-Visual Generation with Customized Narrative Control
Kaiqi Liu, Yunyao Mao, Ziqi Cai +8
MAVIN is a framework for generating multi-shot audio‑visual content with fine‑grained narrative control, using boundary‑aware attention to align temporal segments and ID‑aware prop…
cs.CV2026
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
Yiyan Xu, Qiulin Wang, Wenjie Wang +5
Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enh…
cs.LG2025
Hyper-Connections
Defa Zhu, Hongzhi Huang, Zihao Huang +5
We present hyper-connections, a simple yet effective method that can serve as an alternative to residual connections. This approach specifically addresses common drawbacks observed…