4 papers
DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable
Pu Cao, Qingye Kong, Xuedan Yin +5
Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their raster outputs remain difficult…
A Tilted Seesaw: Revisiting Autoencoder Trade-off for Controllable Diffusion
Pu Cao, Yiyang Ma, Feng Zhou +3
In latent diffusion models, the autoencoder (AE) is typically expected to balance two capabilities: faithful reconstruction and a generation-friendly latent space (e.g., low gFID).…
Preliminary Explorations with GPT-4o(mni) Native Image Generation
Pu Cao, Feng Zhou, Junyi Ji +8
Recently, the visual generation ability by GPT-4o(mni) has been unlocked by OpenAI. It demonstrates a very remarkable generation capability with excellent multimodal condition unde…
Exploring Position Encoding in Diffusion U-Net for Training-free High-resolution Image Generation
Feng Zhou, Pu Cao, Yiyang Ma +2
Denoising higher-resolution latents via a pre-trained U-Net leads to repetitive and disordered image patterns. Although recent studies make efforts to improve generative quality by…