activity
20162026
most citedWorking Memory Connections for LSTM

282 citations · 328 across the 54 of their papers we have counts for

collaborators
Showing 2024Show all

10 papers · 1 filter

cs.CV2024

Perceive, Query & Reason: Enhancing Video QA with Question-Guided Temporal Queries

Roberto Amoroso, Gengyuan Zhang, Rajat Koner +3

Video Question Answering (Video QA) is a challenging video understanding task that requires models to comprehend entire videos, identify the most relevant information based on cont…

cs.CV20242 cited

Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis

Davide Bucciarelli, Nicholas Moratelli, Marcella Cornia +2

The task of image captioning demands an algorithm to generate natural language descriptions of visual inputs. Recent advancements have seen a convergence between image captioning r…

cs.CV2024

Causal Graphical Models for Vision-Language Compositional Understanding

Fiorenzo Parascandolo, Nicholas Moratelli, Enver Sangineto +2

Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language, usually modeling an image capt…

cs.CV2024

Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation

Luca Barsellotti, Lorenzo Bianchi, Nicola Messina +5

Open-Vocabulary Segmentation (OVS) aims at segmenting images from free-form textual concepts without predefined training classes. While existing vision-language models such as CLIP…

cs.CV2024

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

Federico Cocchi, Nicholas Moratelli, Marcella Cornia +2

Multimodal LLMs (MLLMs) are the natural extension of large language models to handle multimodal inputs, combining text and image data. They have recently garnered attention due to…

cs.CV2024

Personalized Instance-based Navigation Toward User-Specific Objects in Realistic Environments

Luca Barsellotti, Roberto Bigazzi, Marcella Cornia +2

In the last years, the research interest in visual navigation towards objects in indoor environments has grown significantly. This growth can be attributed to the recent availabili…