1 citations · 1 across the 3 of their papers we have counts for
8 papers
From Frames to Temporal Graphs: In-Context Egocentric Action Recognition with Vision-Language Models
Bessie Dominguez-Dager, Francisco Gomez-Donoso, Miguel Cazorla +3
Action reasoning in egocentric video requires capturing fine-grained transitions of hand-object interactions, a task where general-purpose Vision-Language Models (VLMs) often strug…
AIDEN: Design and Pilot Study of an AI Assistant for the Visually Impaired
Luis Marquez-Carpintero, Francisco Gomez-Donoso, Zuria Bauer +6
This paper presents AIDEN, an artificial intelligence-based assistant designed to enhance the autonomy and daily quality of life of visually impaired individuals, who often struggl…
DIPSER: A Dataset for In-Person Student Engagement Recognition in the Wild
Luis Marquez-Carpintero, Sergio Suescun-Ferrandiz, Carolina Lorenzo Ãlvarez +4
In this paper, a novel dataset is introduced, designed to assess student attention within in-person classroom settings. This dataset encompasses RGB camera data, featuring multiple…
VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator
Bessie Dominguez-Dager, Sergio Suescun-Ferrandiz, Felix Escalona +2
This paper introduces VLN-Pilot, a novel framework in which a large Vision-and-Language Model (VLLM) assumes the role of a human pilot for indoor drone navigation. By leveraging th…
Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI
Luis Marquez-Carpintero, Alberto Lopez-Sellers, Miguel Cazorla
Simulated Students offer a valuable methodological framework for evaluating pedagogical approaches and modelling diverse learner profiles, tasks which are otherwise challenging to…
CHIRLA: Comprehensive High-resolution Identification and Re-identification for Large-scale Analysis
Bessie Dominguez-Dager, Felix Escalona, Francisco Gomez-Donoso +1
Person re-identification (Re-ID) is a key challenge in computer vision, requiring the matching of individuals across cameras, locations, and time. While most research focuses on sh…