activity
20162024
most citedWorking Memory Connections for LSTM

282 citations · 324 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

15 papers · 1 filter

cs.CV20222 cited

CaMEL: Mean Teacher Learning for Image Captioning

Manuele Barraco, Matteo Stefanini, Marcella Cornia +3

Describing images in natural language is a fundamental step towards the automatic modeling of connections between the visual and textual modalities. In this paper we present CaMEL,…

cs.CV20218 cited

Learning to Select: A Fully Attentive Approach for Novel Object Captioning

Marco Cagrandi, Marcella Cornia, Matteo Stefanini +2

Image captioning models have lately shown impressive results when applied to standard datasets. Switching to real-life scenarios, however, constitutes a challenge due to the larger…

cs.CV2021

Out of the Box: Embodied Navigation in the Real World

Roberto Bigazzi, Federico Landi, Marcella Cornia +3

The research field of Embodied AI has witnessed substantial progress in visual navigation and exploration thanks to powerful simulating platforms and the availability of 3D data of…

cs.CV20212 cited

Revisiting The Evaluation of Class Activation Mapping for Explainability: A Novel Metric and Experimental Analysis

Samuele Poppi, Marcella Cornia, Lorenzo Baraldi +1

As the request for deep learning solutions increases, the need for explainability is even more fundamental. In this setting, particular attention has been given to visualization te…

cs.CV2021

RMS-Net: Regression and Masking for Soccer Event Spotting

Matteo Tomei, Lorenzo Baraldi, Simone Calderara +2

The recently proposed action spotting task consists in finding the exact timestamp in which an event occurs. This task fits particularly well for soccer videos, where events corres…

cs.CV20201 cited

Inter-Homines: Distance-Based Risk Estimation for Human Safety

Matteo Fabbri, Fabio Lanzi, Riccardo Gasparini +3

In this document, we report our proposal for modeling the risk of possible contagiousity in a given area monitored by RGB cameras where people freely move and interact. Our system,…