From the 1 of 12 linked papers with an AI index.
12 papers
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
Haoyu Chen, Kaichen Zhou, Hang Hua +11
The paper introduces MemoBench, a benchmark that tests video generation models' ability to remember and correctly update objects that disappear and later reappear in dynamically ch…
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
Kaichen Zhou, Zeyang Bai, Xinhai Chang +3
View-conditioned 3D generators such as SAM 3D, TRELLIS, and Hunyuan3D produce high-quality object reconstructions from a single view, but real-world visual observation often arrive…
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8
Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the genera…
PAGE-4D: Disentangled pose and geometry estimation for vggt-4d perception
Kaichen Zhou, Yuhan Wang, Grace Chen +5
Recent 3D feed-forward models, such as the Visual Geometry Grounded Transformer (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since the…
EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation
Yue Ma, Xu Ye, Qinghe Wang +9
Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of spatial textures and temporal…
PaintCopilot: Modeling Painting as Autonomous Artistic Continuation
Yunge Wen, Yuancheng Shen, Paul Pu Liang
We present PaintCopilot, a co-creative neural painting assistant that models painting as an open-ended autoregressive artistic behavior conditioned on evolving canvas states and pr…