39 papers
StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization
Yuyang Yin, Zixiang Li, Longxuan Deng +11
Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras,…
StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
Jianing Peng, Mengyu Wang, Henghui Ding +6
Multi-reference image generation aims to synthesize images by integrating attributes from multiple reference images under textual instructions. As the number of references increase…
Let ViT Speak: Generative Language-Image Pre-training
Yan Fang, Mengcheng Lan, Zilong Huang +7
In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers…
CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer
Wenbo Nie, Zixiang Li, Renshuai Tao +3
Transferring visual style between images while preserving semantic correspondence between similar objects remains a central challenge in computer vision. While existing methods hav…
On Exact Editing of Flow-Based Diffusion Models
Zixiang Li, Yue Song, Jianing Peng +7
Recent methods in flow-based diffusion editing have enabled direct transformations between source and target image distribution without explicit inversion. However, the latent traj…
StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
Ke Xing, Xiaojie Jin, Longfei Li +8
The growing adoption of XR devices has fueled strong demand for high-quality stereo video, yet its production remains costly and artifact-prone. To address this challenge, we prese…