4 papers
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
Zijing Hu, Yunze Tong, Fengda Zhang +3
Diffusion models have achieved impressive results in generating high-quality images. Yet, they often struggle to faithfully align the generated images with the input prompts. This…
TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning
Junhua Liu, Zhangcheng Wang, Zhike Han +3
Visual Chain-of-Thought (VCoT) has emerged as a promising paradigm for enhancing multimodal reasoning by integrating visual perception into intermediate reasoning steps. However, e…
CDLM: Causal Concept-Guided Diffusion Large Language Models
Kairong Han, Nuanqiao Shan, Ziyu Zhao +6
Autoregressive (AR) language models and Diffusion Language Models (DLMs) constitute the two principal paradigms of large language models. However, both paradigms suffer from insuff…
D-Fusion: Direct Preference Optimization for Aligning Diffusion Models with Visually Consistent Samples
Zijing Hu, Fengda Zhang, Kun Kuang
The practical applications of diffusion models have been limited by the misalignment between generated images and corresponding text prompts. Recent studies have introduced direct…