5 papers
Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
Koen P. de Vries, Xavier Alameda-Pineda, Estefanía Talavera +1
Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become t…
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
Massimiliano Pappa, Luca Romani, Valentino Sacco +5
Deploying safety-critical agents requires anticipating the consequences of actions before they are executed. While world models offer a paradigm for this proactive foresight, curre…
Diffusion Reinforcement Learning via Centered Reward Distillation
Yuanzhi Zhu, Xi Wang, Stéphane Lathuilière +1
Diffusion and flow models achieve State-Of-The-Art (SOTA) generative performance, yet many practically important behaviors such as fine-grained prompt fidelity, compositional corre…
Residual Tokens Enhance Masked Autoencoders for Speech Modeling
Samir Sadok, Stéphane Lathuilière, Xavier Alameda-Pineda
Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce…
Don't Forget your Inverse DDIM for Image Editing
Guillermo Gomez-Trenado, Pablo Mesejo, Oscar Cordón +1
The field of text-to-image generation has undergone significant advancements with the introduction of diffusion models. Nevertheless, the challenge of editing real images persists,…