1 citations · 1 across the 2 of their papers we have counts for
6 papers
Understanding How MLLMs Describe Artworks Using Token Activation Maps
Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi +3
Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When an MLLM names a style, identi…
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
Ivan Rinaldi, Matteo Mendula, Nicola Fanelli +4
Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-cond…
Label Anything: Multi-Class Few-Shot Semantic Segmentation with Visual Prompts
Pasquale De Marinis, Nicola Fanelli, Raffaele Scaringi +4
Few-shot semantic segmentation aims to segment objects from previously unseen classes using only a limited number of labeled examples. In this paper, we introduce Label Anything, a…
ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval
Nicola Fanelli, Gennaro Vessio, Giovanna Castellano
Analyzing digitized artworks presents unique challenges, requiring not only visual interpretation but also a deep understanding of rich artistic, contextual, and historical knowled…
I Dream My Painting: Connecting MLLMs and Diffusion Models via Prompt Generation for Text-Guided Multi-Mask Inpainting
Nicola Fanelli, Gennaro Vessio, Giovanna Castellano
Inpainting focuses on filling missing or corrupted regions of an image to blend seamlessly with its surrounding content and style. While conditional diffusion models have proven ef…
Art2Mus: Bridging Visual Arts and Music through Cross-Modal Generation
Ivan Rinaldi, Nicola Fanelli, Giovanna Castellano +1
Artificial Intelligence and generative models have revolutionized music creation, with many models leveraging textual or visual prompts for guidance. However, existing image-to-mus…