4 papers
Understanding Multimodal Complementarity for Single-Frame Action Anticipation
Manuel Benavent-Lledo, Konstantinos Bacharidis, Konstantinos Papoutsakis +2
Human action anticipation is commonly treated as a video understanding problem, implicitly assuming that dense temporal information is required to reason about future actions. In t…
CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection
David Ortiz-Perez, Manuel Benavent-Lledo, Javier Rodriguez-Juan +2
Early detection of cognitive disorders such as Alzheimer's disease is critical for enabling timely clinical intervention and improving patient outcomes. In this work, we introduce…
Detecting Facial Image Manipulations with Multi-Layer CNN Models
Alejandro Marco Montejano, Angela Sanchez Perez, Javier Barrachina +3
The rapid evolution of digital image manipulation techniques poses significant challenges for content verification, with models such as stable diffusion and mid-journey producing h…
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
Manuel Benavent-Lledo, David Mulero-Pérez, David Ortiz-Perez +2
We propose a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including…