14 papers
XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation
Elena Izzo, Riccardo Toniolo, Lamberto Ballan
Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computationally constrained devices due…
IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
Jelin Raphael Akkara, Filippo Ziliotto, Luciano Serafini +2
Embodied AI increasingly relies on queryable semantic maps built from pre-trained vision-language models to enable zero-shot Object Goal Navigation (ObjectNav). However, existing a…
What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility
Filippo Ziliotto, Luciano Serafini, Lamberto Ballan +1
A fundamental challenge in 3D reconstruction and robotic localization is co-visibility: determining which image pairs share overlapping visible surfaces, particularly in scenarios…
You Only Landmark Once: Lightweight U-Net Face Super Resolution with YOLO-World Landmark Heatmaps
Riccardo Carraro, Anna Briotto, Endi Hysa +2
Face image super-resolution aims to recover high-resolution facial images from severely degraded inputs. Under extreme upscaling factors, fine facial details are often lost, making…
Contrastive Learning under Noisy Temporal Self-Supervision for Colonoscopy Videos
Luca Parolari, Pietro Gori, Lamberto Ballan +2
Learning robust representations of polyp tracklets is key to enabling multiple AI-assisted colonoscopy applications, from polyp characterization to automated reporting and retrieva…
Benchmarking Layout-Guided Diffusion Models through Unified Semantic-Spatial Evaluation in Closed and Open Settings
Luca Parolari, Nicla Faccioli, Lamberto Ballan
Evaluating layout-guided text-to-image generative models requires assessing both semantic alignment with textual prompts and spatial fidelity to prescribed layouts. Assessing layou…