24 papers
SLU-2K: A Question-Based Benchmark for Semantic Evaluation of Sign Language Translation
Zeno Testa, Antonino Furnari, Lorenzo Baraldi +1
Sign Language Translation (SLT) is typically evaluated with surface-form metrics such as BLEU and ROUGE, which reward lexical overlap but do not directly measure whether a translat…
EGOSTREAM: A Diagnostic Benchmark for Streaming Episodic Memory in Egocentric Vision
Rosario Forte, Giuseppe Lando, Antonino Furnari
Continuous episodic memory is a core capability for autonomous agents operating in dynamic, real-world environments, yet current streaming video benchmarks provide limited tools fo…
Ego-METAS: Egocentric online Multimodal Energy-efficient Temporal Action Segmentation benchmark
Maria Santos-Villafranca, Jesus Bermudez-cameo, Alejandro Perez-Yus +2
To operate in the physical world, embodied agents must perceive their environment in an "always-on" fashion, selectively accessing the most informative sensors to balance energy co…
RECIPE: Procedural Planning via Grounding in Instructional Video
Luigi Seminara, Antonino Furnari, Lorenzo Torresani
Visual planning asks a model to generate the remaining steps of a procedure in natural language given a partial video context and a goal. Progress on this task is bottlenecked by a…
Exploring Multimodal LMMs for Online Episodic Memory Question Answering on the Edge
Giuseppe Lando, Rosario Forte, Antonino Furnari
We investigate the feasibility of using Multimodal Large Language Models (MLLMs) for real-time online episodic memory question answering. While cloud offloading is common, it raise…
Leveraging Synthetic Data for Enhancing Egocentric Hand-Object Interaction Detection
Rosario Leonardi, Antonino Furnari, Francesco Ragusa +1
In this work, we explore the role of synthetic data in improving the detection of Hand-Object Interactions from egocentric images. Through extensive experimentation and comparative…