From the 1 of 7 linked papers with an AI index.
7 papers
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Yuyang Yin, Zixiang Li, Longxuan Deng +11
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras,…
UniVR: Thinking in Visual Space for Unified Visual Reasoning
Zhongwei Ren, Yunchao Wei, Yao Zhao +5
The paper presents UniVR, a system that learns complex visual reasoning, fine-grained physical dynamics, and long-term planning directly from raw video demonstrations using a novel…
VideoWorld 2: Learning Transferable Knowledge from Real-world Videos
Zhongwei Ren, Yunchao Wei, Xiao Yu +5
Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, wh…
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
Ke Xing, Xiaojie Jin, Longfei Li +8
The growing adoption of XR devices has fueled strong demand for high-quality stereo video, yet its production remains costly and artifact-prone. To address this challenge, we prese…
A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
Dongheng Lin, Mengxue Qu, Kunyang Han +3
Most video-anomaly research stops at frame-wise detection, offering little insight into why an event is abnormal, typically outputting only frame-wise anomaly scores without spatia…
PreFM: Online Audio-Visual Event Parsing via Predictive Future Modeling
Xiao Yu, Yan Fang, Xiaojie Jin +2
Audio-visual event parsing plays a crucial role in understanding multimodal video content, but existing methods typically rely on offline processing of entire videos with huge mode…