4 papers · 1 filter
XiDepth: a Lightweight and Efficient Network for Self-supervised Monocular Depth Estimation
Elena Izzo, Riccardo Toniolo, Lamberto Ballan
Self-supervised monocular depth estimation has emerged as an appealing solution to design lightweight and effective models for deployment on computationally constrained devices due…
7Bench: a Comprehensive Benchmark for Layout-guided Text-to-image Models
Elena Izzo, Luca Parolari, Davide Vezzaro +1
Layout-guided text-to-image models offer greater control over the generation process by explicitly conditioning image synthesis on the spatial arrangement of elements. As a result,…
Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension
Luca Parolari, Elena Izzo, Lamberto Ballan
Referring Expression Comprehension (REC) aims to identify a particular object in a scene by a natural language expression, and is an important topic in visual language understandin…
Where are my Neighbors? Exploiting Patches Relations in Self-Supervised Vision Transformer
Guglielmo Camporese, Elena Izzo, Lamberto Ballan
Vision Transformers (ViTs) enabled the use of the transformer architecture on vision tasks showing impressive performances when trained on big datasets. However, on relatively smal…