6 papers
VIRAL: Visual In-Context Reasoning via Analogy in Diffusion Transformers
Zhiwen Li, Zhongjie Duan, Jinyan Ye +4
Replicating In-Context Learning (ICL) in computer vision remains challenging due to task heterogeneity. We propose \textbf{VIRAL}, a framework that elicits visual reasoning from a…
Spectral Evolution Search: Efficient Inference-Time Scaling for Reward-Aligned Image Generation
Jinyan Ye, Zhongjie Duan, Zhiwen Li +4
Inference-time scaling offers a versatile paradigm for aligning visual generative models with downstream objectives without parameter updates. However, existing approaches that opt…
AutoLoRA: Automatic LoRA Retrieval and Fine-Grained Gated Fusion for Text-to-Image Generation
Zhiwen Li, Zhongjie Duan, Die Chen +4
Despite recent advances in photorealistic image generation through large-scale models like FLUX and Stable Diffusion v3, the practical deployment of these architectures remains con…
AttriCtrl: Fine-Grained Control of Aesthetic Attribute Intensity in Diffusion Models
Die Chen, Zhongjie Duan, Zhiwen Li +4
Diffusion models have recently become the dominant paradigm for image generation, yet existing systems struggle to interpret and follow numeric instructions for adjusting semantic…
Responsible Diffusion Models via Constraining Text Embeddings within Safe Regions
Zhiwen Li, Die Chen, Mingyuan Fan +4
The remarkable ability of diffusion models to generate high-fidelity images has led to their widespread adoption. However, concerns have also arisen regarding their potential to pr…
ArtAug: Enhancing Text-to-Image Generation through Synthesis-Understanding Interaction
Zhongjie Duan, Qianyi Zhao, Cen Chen +4
The emergence of diffusion models has significantly advanced image synthesis. The recent studies of model interaction and self-corrective reasoning approach in large language model…