5 papers
IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
Jelin Raphael Akkara, Filippo Ziliotto, Luciano Serafini +2
Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing a…
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
Filippo Ziliotto, Luciano Serafini, Lamberto Ballan +1
A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios…
Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?
Filippo Ziliotto, Ciro Beneduce, Bruno Lepri +3
In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogous functional capability emer…
PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents
Filippo Ziliotto, Jelin Raphael Akkara, Alessandro Daniele +3
Recent advances in Embodied AI have enabled agents to perform increasingly complex tasks and adapt to diverse environments. However, deploying such agents in realistic human-center…
TANGO: Training-free Embodied AI Agents for Open-world Tasks
Filippo Ziliotto, Tommaso Campari, Luciano Serafini +1
Large Language Models (LLMs) have demonstrated excellent capabilities in composing various modules together to create programs that can perform complex reasoning tasks on images. I…