From the 1 of 6 linked papers with an AI index.
6 papers
Temporal Concentration from Rollout Errors: Implicit Preference Optimization for Text-to-Video Diffusion
Henglin Liu, Fangyuan Kong, Jing Wang +7
The paper introduces concentrated Implicit Preference Optimization (cIPO), a post‑training method for text‑to‑video diffusion models that derives preference signals from reconstruc…
Diffusing in the Right Space: A Systematic Study of Latent Diffusability
Tianxiong Zhong, Xingye Tian, Xuebo Wang +2
Latent diffusion models leverage visual tokenizers to compress images into latent spaces for efficient generative modeling. However, better reconstruction quality of a tokenizer do…
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation
Yiyan Xu, Qiulin Wang, Wenjie Wang +5
Multi-reference image generation aims to synthesize images from textual instructions while faithfully preserving subject identities from multiple reference images. Existing VLM-enh…
Efficient Video Diffusion Models: Advancements and Challenges
Shitong Shao, Lichen Bai, Pengfei Wan +2
Video diffusion models have rapidly become the dominant paradigm for high-fidelity generative video synthesis, but their practical deployment remains constrained by severe inferenc…
Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment
Henglin Liu, Nisha Huang, Chang Liu +6
The aesthetic quality assessment task is crucial for developing a human-aligned quantitative evaluation system for AIGC. However, its inherently complex nature, spanning visual per…
Free Lunch Alignment of Text-to-Image Diffusion Models without Preference Image Pairs
Jia Jun Cheng Xian, Muchen Li, Haotian Yang +4
Recent advances in diffusion-based text-to-image (T2I) models have led to remarkable success in generating high-quality images from textual prompts. However, ensuring accurate alig…