Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
Koen P. de Vries, Xavier Alameda-Pineda, Estefanía Talavera +1
Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become t…
cs.CV2026
Diffusion Reinforcement Learning via Centered Reward Distillation
Yuanzhi Zhu, Xi Wang, Stéphane Lathuilière +1
Diffusion and flow models achieve State-Of-The-Art (SOTA) generative performance, yet many practically important behaviors such as fine-grained prompt fidelity, compositional corre…
cs.CV2025
Don't Forget your Inverse DDIM for Image Editing
Guillermo Gomez-Trenado, Pablo Mesejo, Oscar Cordón +1
The field of text-to-image generation has undergone significant advancements with the introduction of diffusion models. Nevertheless, the challenge of editing real images persists,…