8 papers · 1 filter
MemoBench: Benchmarking World Modeling in Dynamically Changing Environments
Haoyu Chen, Kaichen Zhou, Hang Hua +11
Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while t…
EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX Generation
Yue Ma, Xu Ye, Qinghe Wang +9
Generating high-fidelity visual effects (VFX) typically demands massive datasets and prohibitive computational power due to the intricate coupling of spatial textures and temporal…
PaintCopilot: Modeling Painting as Autonomous Artistic Continuation
Yunge Wen, Yaluo Wang, Yuancheng Shen +2
Existing neural painting methods are target-driven: given a reference image, strokes are optimized to reconstruct it, fixing the outcome before painting begins. We instead ask whet…
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8
Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the genera…
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
Kaichen Zhou, Zeyang Bai, Xinhai Chang +3
View-conditioned 3D generators such as SAM 3D, TRELLIS, and Hunyuan3D produce high-quality object reconstructions from a single view, but real-world visual observation often arrive…
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
Yifan Liu, Fangneng Zhan, Kaichen Zhou +3
Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and e…