3 papers
cs.CV2026
DrawAI: Agentic Benchmark and Workflow for Making Raster Images Editable
Pu Cao, Qingye Kong, Xuedan Yin +5
Recent image-generation models and multimodal agents can produce high-quality visuals for increasingly complex visual communication tasks. Yet their raster outputs remain difficult…
cs.CV2026
A Tilted Seesaw: Revisiting Autoencoder Trade-off for Controllable Diffusion
Pu Cao, Yiyang Ma, Feng Zhou +3
In latent diffusion models, the autoencoder (AE) is typically expected to balance two capabilities: faithful reconstruction and a generation-friendly latent space (e.g., low gFID).…
cs.CV2025
ResDiT: Evoking the Intrinsic Resolution Scalability in Diffusion Transformers
Yiyang Ma, Feng Zhou, Xuedan Yin +3
Leveraging pre-trained Diffusion Transformers (DiTs) for high-resolution (HR) image synthesis often leads to spatial layout collapse and degraded texture fidelity. Prior work mitig…