7 papers
DitHub: A Modular Framework for Incremental Open-Vocabulary Object Detection
Chiara Cappellino, Gianluca Mancusi, Matteo Mosconi +3
Open-Vocabulary object detectors can generalize to an unrestricted set of categories through simple textual prompting. However, adapting these models to rare classes or reinforcing…
BRUM: Robust 3D Vehicle Reconstruction from 360 Sparse Images
Davide Di Nucci, Matteo Tomei, Guido Borghi +3
Accurate 3D reconstruction of vehicles is vital for applications such as vehicle inspection, predictive maintenance, and urban planning. Existing methods like Neural Radiance Field…
One Transformer for All Time Series: Representing and Training with Time-Dependent Heterogeneous Tabular Data
Simone Luetto, Fabrizio Garuti, Enver Sangineto +2
There is a recent growing interest in applying Deep Learning techniques to tabular data, in order to replicate the success of other Artificial Intelligence areas in this structured…
A Second-Order Perspective on Model Compositionality and Incremental Learning
Angelo Porrello, Lorenzo Bonicelli, Pietro Buzzega +3
The fine-tuning of deep pre-trained models has revealed compositional properties, with multiple specialized modules that can be arbitrarily composed into a single, multi-task model…
Monocular Per-Object Distance Estimation with Masked Object Modeling
Aniello Panariello, Gianluca Mancusi, Fedy Haj Ali +3
Per-object distance estimation is critical in surveillance and autonomous driving, where safety is crucial. While existing methods rely on geometric or deep supervised features, on…
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
Lorenzo Baraldi, Roberto Amoroso, Marcella Cornia +2
The use of self-supervised pre-training has emerged as a promising approach to enhance the performance of many different visual tasks. In this context, recent approaches have emplo…