4 papers
Memory-Augmented Vision-Language Agents for Persistent and Semantically Consistent Object Captioning
Tommaso Galliena, Stefano Rosa, Tommaso Apicella +3
Vision-Language Models (VLMs) often yield inconsistent descriptions of the same object across viewpoints, hindering the ability of embodied agents to construct consistent semantic…
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
Fanqi Yu, Matteo Tiezzi, Tommaso Apicella +2
We introduce a lifelong imitation learning framework that enables continual policy refinement across sequential tasks under realistic memory and data constraints. Our approach depa…
Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions
Tommaso Galliena, Tommaso Apicella, Stefano Rosa +3
We present a self-supervised method to improve an agent's abilities in describing arbitrary objects while actively exploring a generic environment. This is a challenging problem, a…
Learning to Evaluate Autonomous Behaviour in Human-Robot Interaction
Matteo Tiezzi, Tommaso Apicella, Carlos Cardenas-Perez +5
Evaluating and comparing the performance of autonomous Humanoid Robots is challenging, as success rate metrics are difficult to reproduce and fail to capture the complexity of robo…