3 papers
cs.CV2026
Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video Continuation
Yingmao Miao, Pengfei Zhang, Chaoran Xu +5
Video generators build long videos by composing shorter parts, either by generating segments one after another or by autoregressively extending chunks. Each new part usually depend…
cs.CV2026
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
Yingmao Miao, Pengfei Zhang, Xiaochen Lv +5
While recent image editing models have made rapid progress, multi-reference editing remains challenging, particularly in maintaining visual consistency across references and ensuri…
cs.SD2025
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation
Tianxin Xie, Wentao Lei, Kai Jiang +27
Text-to-audio-video (T2AV) generation is central to applications such as filmmaking and world modeling. However, current models often fail to produce physically plausible sounds. P…