2 papers
cs.LG2026
PG-MAP: Joint MAP Optimization for Inference-Time Alignment of Diffusion and Flow-Matching Models
Ruolan Sun, Pawel Polak
Inference-time alignment of pretrained text-to-image models is typically performed along a single control axis, such as classifier-free guidance, attention editing, or reward-based…
cs.CV2025
Every Image Listens, Every Image Dances: Music-Driven Image Animation
Zhikang Dong, Weituo Hao, Ju-Chiang Wang +2
Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video g…