3 papers
cs.CV2026
Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
Koen P. de Vries, Xavier Alameda-Pineda, Estefanía Talavera +1
Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become t…
cs.AI2026
Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
Massimiliano Pappa, Luca Romani, Valentino Sacco +5
Deploying safety-critical agents requires anticipating the consequences of actions before they are executed. While world models offer a paradigm for this proactive foresight, curre…
cs.CV2025
Layover or Direct Flight: Rethinking Audio-Guided Image Segmentation
Joel Alberto Santos, Zongwei Wu, Xavier Alameda-Pineda +1
Understanding human instructions is essential for enabling smooth human-robot interaction. In this work, we focus on object grounding, i.e., localizing an object of interest in a v…