6 papers
IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
Jelin Raphael Akkara, Filippo Ziliotto, Luciano Serafini +2
Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing a…
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
Filippo Ziliotto, Luciano Serafini, Lamberto Ballan +1
A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios…
Mirror, Mirror on the Wall: Can VLM Agents Tell Who They Are at All?
Filippo Ziliotto, Ciro Beneduce, Bruno Lepri +3
In the animal kingdom, mirror self-recognition is a canonical probe of higher-order cognition, emerging only in some species. We ask whether an analogous functional capability emer…
Discrete World Models via Regularization
Davide Bizzaro, Luciano Serafini
World models aim to capture the states and dynamics of an environment in a compact latent space. Moreover, using Boolean state representations is particularly useful for search heu…
PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents
Filippo Ziliotto, Jelin Raphael Akkara, Alessandro Daniele +3
Recent advances in Embodied AI have enabled agents to perform increasingly complex tasks and adapt to diverse environments. However, deploying such agents in realistic human-center…
TANGO: Training-free Embodied AI Agents for Open-world Tasks
Filippo Ziliotto, Tommaso Campari, Luciano Serafini +1
Large Language Models (LLMs) have demonstrated excellent capabilities in composing various modules together to create programs that can perform complex reasoning tasks on images. I…