7 papers
DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable
Pu Cao, Qingye Kong, Xuedan Yin +5
Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their raster outputs remain difficult…
A Tilted Seesaw: Revisiting Autoencoder Trade-off for Controllable Diffusion
Pu Cao, Yiyang Ma, Feng Zhou +3
In latent diffusion models, the autoencoder (AE) is typically expected to balance two capabilities: faithful reconstruction and a generation-friendly latent space (e.g., low gFID).…
Controllable Generation with Text-to-Image Diffusion Models: A Survey
Pu Cao, Feng Zhou, Qing Song +1
In the rapidly advancing realm of visual generation, diffusion models have revolutionized the landscape, marking a significant shift in capabilities with their impressive text-guid…
LSAP: Rethinking Inversion Fidelity, Perception and Editability in GAN Latent Space
Xuekun Zhao, Pu Cao, Xiaoya Yang +3
As research on image inversion advances, the process is generally divided into two stages. The first step is Image Embedding, involves using an encoder or optimization procedure to…
Preliminary Explorations with GPT-4o(mni) Native Image Generation
Pu Cao, Feng Zhou, Junyi Ji +8
Recently, the visual generation ability by GPT-4o(mni) has been unlocked by OpenAI. It demonstrates a very remarkable generation capability with excellent multimodal condition unde…
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
Pu Cao, Feng Zhou, Lu Yang +2
In-domain generation aims to perform a variety of tasks within a specific domain, such as unconditional generation, text-to-image, image editing, 3D generation, and more. Early res…