activity
20242026
collaborators

5 papers

cs.CV2026

Understanding Multimodal Complementarity for Single-Frame Action Anticipation

Manuel Benavent-Lledo, Konstantinos Bacharidis, Konstantinos Papoutsakis +2

Human action anticipation is commonly treated as a video understanding problem, implicitly assuming that dense temporal information is required to reason about future actions. In t…

cs.LG2025

CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection

David Ortiz-Perez, Manuel Benavent-Lledo, Javier Rodriguez-Juan +2

Early detection of cognitive disorders such as Alzheimer's disease is critical for enabling timely clinical intervention and improving patient outcomes. In this work, we introduce…

cs.CV2025

Text-driven Online Action Detection

Manuel Benavent-Lledo, David Mulero-Pérez, David Ortiz-Perez +1

Detecting actions as they occur is essential for applications like video surveillance, autonomous driving, and human-robot interaction. Known as online action detection, this task…

cs.CV2025

Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos

Javier Rodriguez-Juan, David Ortiz-Perez, Manuel Benavent-Lledo +5

The current biodiversity loss crisis makes animal monitoring a relevant field of study. In light of this, data collected through monitoring can provide essential insights, and info…

cs.CV2024

Detecting Facial Image Manipulations with Multi-Layer CNN Models

Alejandro Marco Montejano, Angela Sanchez Perez, Javier Barrachina +3

The rapid evolution of digital image manipulation techniques poses significant challenges for content verification, with models such as stable diffusion and mid-journey producing h…