282 citations · 673 across the 75 of their papers we have counts for
17 papers · 2 filters
Is Multiple Object Tracking a Matter of Specialization?
Gianluca Mancusi, Mattia Bernardi, Aniello Panariello +3
End-to-end transformer-based trackers have achieved remarkable performance on most human-related datasets. However, training these trackers in heterogeneous scenarios poses signifi…
KRONC: Keypoint-based Robust Camera Optimization for 3D Car Reconstruction
Davide Di Nucci, Alessandro Simoni, Matteo Tomei +3
The three-dimensional representation of objects or scenes starting from a set of images has been a widely discussed topic for years and has gained additional attention after the di…
BRIDGE: Bridging Gaps in Image Captioning Evaluation with Stronger Visual Cues
Sara Sarto, Marcella Cornia, Lorenzo Baraldi +1
Effectively aligning with human judgment when evaluating machine-generated image captions represents a complex yet intriguing challenge. Existing evaluation metrics like CIDEr or C…
Contrasting Deepfakes Diffusion via Contrastive Learning and Global-Local Similarities
Lorenzo Baraldi, Federico Cocchi, Marcella Cornia +2
Discerning between authentic content and that generated by advanced AI methods has become increasingly challenging. While previous research primarily addresses the detection of fak…
Mask and Compress: Efficient Skeleton-based Action Recognition in Continual Learning
Matteo Mosconi, Andriy Sorokin, Aniello Panariello +6
The use of skeletal data allows deep learning models to perform action recognition efficiently and effectively. Herein, we believe that exploring this problem within the context of…
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
Davide Caffagni, Federico Cocchi, Nicholas Moratelli +4
Multimodal LLMs are the natural evolution of LLMs, and enlarge their capabilities so as to work beyond the pure textual modality. As research is being carried out to design novel a…