activity
20232026
most citedLearning Mask-aware CLIP Representations for Zero-Shot Segmentation

13 citations · 27 across the 27 of their papers we have counts for

collaborators
Showing cs.CVShow all

50 papers · 1 filter

cs.CV2026

StateFlow: Building, Evolving, and Accessing 3D World States for Previsualization

Yuyang Yin, Zixiang Li, Longxuan Deng +11

Previsualization is an intermediate layer between ideas and production in film, games, architecture, and urban design. It lets creators iteratively refine scenes, actions, cameras,…

cs.CV2026

StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling

Jianing Peng, Mengyu Wang, Henghui Ding +6

Multi-reference image generation aims to synthesize images by integrating attributes from multiple reference images under textual instructions. As the number of references increase…

cs.CV2026

Orca: The World is in Your Mind

Yihao Wang, Yuheng Ji, Mingyu Cao +54

We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals and exposes it through multi…

cs.CV2026

Let ViT Speak: Generative Language-Image Pre-training

Yan Fang, Mengcheng Lan, Zilong Huang +7

In this paper, we present \textbf{Gen}erative \textbf{L}anguage-\textbf{I}mage \textbf{P}re-training (GenLIP), a minimalist generative pretraining framework for Vision Transformers…

cs.CV2026

CoCoDiff: Correspondence-Consistent Diffusion Model for Fine-grained Style Transfer

Wenbo Nie, Zixiang Li, Renshuai Tao +3

Transferring visual style between images while preserving semantic correspondence between similar objects remains a central challenge in computer vision. While existing methods hav…

cs.CV2026

On Exact Editing of Flow-Based Diffusion Models

Zixiang Li, Yue Song, Jianing Peng +7

Recent methods in flow-based diffusion editing have enabled direct transformations between source and target image distribution without explicit inversion. However, the latent traj…