8 papers
Integrating Affordances and Attention models for Short-Term Object Interaction Anticipation
Lorenzo Mur Labadia, Ruben Martinez-Cantin, Jose J. Guerrero +2
Short Term object-interaction Anticipation consists in detecting the location of the next active objects, the noun and verb categories of the interaction, as well as the time to co…
FALCONEye: Finding Answers and Localizing Content in ONE-hour-long videos with multi-modal LLMs
Carlos Plou, Cesar Borja, Ruben Martinez-Cantin +1
Finding information in hour-long videos is a challenging task even for top-performing Vision Language Models (VLMs), as encoding visual content quickly exceeds available context wi…
O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views
Lorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo +3
Understanding the world from multiple perspectives is essential for intelligent systems operating together, where segmenting common objects across different views remains an open p…
Gen-Swarms: Adapting Deep Generative Models to Swarms of Drones
Carlos Plou, Pablo Pueyo, Ruben Martinez-Cantin +3
Gen-Swarms is an innovative method that leverages and combines the capabilities of deep generative models with reactive navigation algorithms to automate the creation of drone show…
DIV-FF: Dynamic Image-Video Feature Fields For Environment Understanding in Egocentric Videos
Lorenzo Mur-Labadia, Josechu Guerrero, Ruben Martinez-Cantin
Environment understanding in egocentric videos is an important step for applications like robotics, augmented reality and assistive technologies. These videos are characterized by…
Influence of field of view in visual prostheses design: Analysis with a VR system
Melani Sanchez-Garcia, Ruben Martinez-Cantin, Jesus Bermudez-Cameo +1
Visual prostheses are designed to restore partial functional vision in patients with total vision loss. Retinal visual prostheses provide limited capabilities as a result of low re…