collaborators

34 papers

cs.CV2026

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

Yuyang Yin, Zixiang Li, Longxuan Deng +11

Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras,…

cs.CV2026

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling

Jianing Peng, Mengyu Wang, Henghui Ding +6

Multi-reference image generation aims to synthesize images by integrating attributes from multiple reference images under textual instructions. As the number of references increase…

cs.CV2026

Rethinking One-Step Image Editing through ChordEdit: Reproduction, Simplification, and New Insights

Minghan Li, Jeremy Moebel, Mengyu Wang

One-step image editing is important for making text-guided editing fast, practical, and easy to deploy, but its underlying mechanism is still not fully understood. We revisit Chord…

cs.CV2026

Stream3D: Sequential Multi-View 3D Generation via Evidential Memory

Kaichen Zhou, Zeyang Bai, Xinhai Chang +3

View-conditioned 3D generators such as SAM 3D, TRELLIS, and Hunyuan3D produce high-quality object reconstructions from a single view, but real-world visual observation often arrive…

cs.CV2026

GeoWorld-VLM: Geometry from World Models for Vision-Language Models

Renjie Gu, Kaichen Zhou, Yan Luo +1

Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between. One cause of…

cs.CV2026

GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation

Kaichen Zhou, Yuzhen Chen, Fangneng Zhan +8

Video world models can generate realistic futures from a single instruction, but they often fail to track the same physical points consistently across time. As a result, the genera…