2 papers
cs.CV2026
Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
Koen P. de Vries, Xavier Alameda-Pineda, Estefanía Talavera +1
Training Multimodal Large Language Models for audio-visual social understanding is a crucial step toward embodied social intelligence. Chain-of-thought (CoT) reasoning has become t…
cs.CV2025
DADO: A Depth-Attention framework for Object Discovery
Federico Gonzalez, Estefania Talavera, Petia Radeva
Unsupervised object discovery, the task of identifying and localizing objects in images without human-annotated labels, remains a significant challenge and a growing focus in compu…