From the 1 of 16 linked papers with an AI index.
16 papers
VicEdit: Learning to Edit Videos from Visual In-Context Examples
Yuji Wang, Teng Hu, Yuheng Chen +6
Despite progress in instruction-based video editing, unimodal textual instructions inherently struggle to convey fine-grained textures and complex dynamics. To bridge this perceptu…
PersonaShot: Benchmarking Person-Centric Narrative Continuity in Multi-Shot Video Generation
Yuji Wang, Yuheng Chen, Teng Hu +7
Video generation is rapidly evolving from single-shot clips to multi-shot narratives, where the human character serves as the core narrative anchor. However, existing benchmarks ma…
Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention
Shengchuan Gao, Teng Hu, Bohao Feng +4
Recent audio-visual generation models can synthesize synchronized video and sound in a unified diffusion process, but their inference cost remains high because long video token seq…
Cycle-World: Mitigating Error Accumulation in Long-term Video World Models via Reverse-Prediction Cycle Consistency
Zihan Su, Teng Hu, Jiangning Zhang +4
The paper introduces Cycle-World, a framework that uses reverse‑prediction cycle consistency to reduce error accumulation in long‑horizon video generation, improving temporal consi…
MetaWorld: Scaling Multi-Agent Video World Model from Single-view Video Data
Teng Hu, Mingchun Lu, Yating Wang +6
Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a sin…
Evolution of Optimization Methods: Algorithms, Scenarios, and Evaluations
Tong Zhang, Jiangning Zhang, Zhucun Xue +9
Balancing convergence speed, generalization capability, and computational efficiency remains a core challenge in deep learning optimization. First-order gradient descent methods, e…