4 papers
ART-VS: Adaptive Resolution Tiling for Vision Transformer Visual Servoing
Alessandro Scherl, Bernhard Neuberger, Simon Schwaiger +3
Visual servoing with self-supervised Vision Transformer (ViT) features enables training-free robotic positioning with strong generalization, but faces a fundamental trade-off betwe…
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
Manuel Benavent-Lledo, David Mulero-Pérez, David Ortiz-Perez +2
We propose a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including…
Text-driven Online Action Detection
Manuel Benavent-Lledo, David Mulero-Pérez, David Ortiz-Perez +1
Detecting actions as they occur is essential for applications like video surveillance, autonomous driving, and human-robot interaction. Known as online action detection, this task…
Visual WetlandBirds Dataset: Bird Species Identification and Behavior Recognition in Videos
Javier Rodriguez-Juan, David Ortiz-Perez, Manuel Benavent-Lledo +5
The current biodiversity loss crisis makes animal monitoring a relevant field of study. In light of this, data collected through monitoring can provide essential insights, and info…