282 citations · 326 across the 38 of their papers we have counts for
6 papers · 1 filter
Meshed-Memory Transformer for Image Captioning
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi +1
Transformer-based architectures represent the state of the art in sequence modeling tasks like machine translation and language understanding. Their applicability to multi-modal co…
Video action detection by learning graph-based spatio-temporal interactions
Matteo Tomei, Lorenzo Baraldi, Simone Calderara +2
Action Detection is a complex task that aims to detect and classify human actions in video clips. Typically, it has been addressed by processing fine-grained features extracted fro…
SMArT: Training Shallow Memory-aware Transformers for Robotic Explainability
Marcella Cornia, Lorenzo Baraldi, Rita Cucchiara
The ability to generate natural language explanations conditioned on the visual perception is a crucial step towards autonomous agents which can explain themselves and communicate…
Embodied Vision-and-Language Navigation with Dynamic Convolutional Filters
Federico Landi, Lorenzo Baraldi, Massimiliano Corsini +1
In Vision-and-Language Navigation (VLN), an embodied agent needs to reach a target destination with the only guidance of a natural language instruction. To explore the environment…
A Deep Learning based approach to VM behavior identification in cloud systems
Matteo Stefanini, Riccardo Lancellotti, Lorenzo Baraldi +1
Cloud computing data centers are growing in size and complexity to the point where monitoring and management of the infrastructure become a challenge due to scalability issues. A p…
M-VAD Names: a Dataset for Video Captioning with Naming
Stefano Pini, Marcella Cornia, Federico Bolelli +2
Current movie captioning architectures are not capable of mentioning characters with their proper name, replacing them with a generic "someone" tag. The lack of movie description d…