activity
20172026
collaborators

7 papers

cs.CV2026

From Frames to Temporal Graphs: In-Context Egocentric Action Recognition with Vision-Language Models

Bessie Dominguez-Dager, Francisco Gomez-Donoso, Miguel Cazorla +3

Action reasoning in egocentric video requires capturing fine-grained transitions of hand-object interactions, a task where general-purpose Vision-Language Models (VLMs) often strug…

cs.RO2026

VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator

Bessie Dominguez-Dager, Sergio Suescun-Ferrandiz, Felix Escalona +2

This paper introduces VLN-Pilot, a novel framework in which a large Vision-and-Language Model (VLLM) assumes the role of a human pilot for indoor drone navigation. By leveraging th…

cs.CV2025

AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired

Luis Marquez-Carpintero, Francisco Gomez-Donoso, Zuria Bauer +6

This paper presents AIDEN, an artificial intelligence-based assistant designed to enhance the autonomy and daily quality of life of visually impaired individuals, who often struggl…

cs.CV2025

CADDI: An in-Class Activity Detection Dataset using IMU data from low-cost sensors

Luis Marquez-Carpintero, Sergio Suescun-Ferrandiz, Monica Pina-Navarro +2

The monitoring and prediction of in-class student activities is of paramount importance for the comprehension of engagement and the enhancement of pedagogical efficacy. The accurat…

cs.CV2025

CHIRLA: Comprehensive High-resolution Identification and Re-identification for Large-scale Analysis

Bessie Dominguez-Dager, Felix Escalona, Francisco Gomez-Donoso +1

Person re-identification (Re-ID) is a key challenge in computer vision, requiring the matching of individuals across cameras, locations, and time. While most research focuses on sh…

cs.CV2020

Attentional-GCNN: Adaptive Pedestrian Trajectory Prediction towards Generic Autonomous Vehicle Use Cases

Kunming Li, Stuart Eiffert, Mao Shan +3

Autonomous vehicle navigation in shared pedestrian environments requires the ability to predict future crowd motion both accurately and with minimal delay. Understanding the uncert…