Showing 2026Show all
2 papers · 1 filter
cs.CV2026
From Frames to Temporal Graphs: In-Context Egocentric Action Recognition with Vision-Language Models
Bessie Dominguez-Dager, Francisco Gomez-Donoso, Miguel Cazorla +3
Action reasoning in egocentric video requires capturing fine-grained transitions of hand-object interactions, a task where general-purpose Vision-Language Models (VLMs) often strug…
cs.RO2026
VLN-Pilot: Large Vision-Language Model as an Autonomous Indoor Drone Operator
Bessie Dominguez-Dager, Sergio Suescun-Ferrandiz, Felix Escalona +2
This paper introduces VLN-Pilot, a novel framework in which a large Vision-and-Language Model (VLLM) assumes the role of a human pilot for indoor drone navigation. By leveraging th…