110 citations · 128 across the 13 of their papers we have counts for
18 papers · 1 filter
Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation
Emanuele Colonna, Moises Diaz, Gennaro Vessio +2
Sign language production, which generates continuous 3D skeletal motion from spoken language input, must simultaneously satisfy two constraints: semantic fidelity, so that a deaf v…
Understanding How MLLMs Describe Artworks Using Token Activation Maps
Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi +3
Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When an MLLM names a style, identi…
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
Ivan Rinaldi, Matteo Mendula, Nicola Fanelli +4
Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-cond…
DistillFSS: Synthesizing Few-Shot Knowledge into a Lightweight Segmentation Model
Pasquale De Marinis, Pieter M. Blok, Uzay Kaymak +3
Cross-Domain Few-Shot Semantic Segmentation (CD-FSS) seeks to segment unknown classes in unseen domains using only a few annotated examples. This setting is inherently challenging:…
Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA
Pasquale De Marinis, Gennaro Vessio, Giovanna Castellano
Few-shot semantic segmentation (FSS) aims to segment novel classes in query images using only a small annotated support set. While prior research has mainly focused on improving de…
Matching-Based Few-Shot Semantic Segmentation Models Are Interpretable by Design
Pasquale De Marinis, Uzay Kaymak, Rogier Brussee +2
Few-Shot Semantic Segmentation (FSS) models achieve strong performance in segmenting novel classes with minimal labeled examples, yet their decision-making processes remain largely…