4 citations · 8 across the 13 of their papers we have counts for
14 papers
Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation
Emanuele Colonna, Moises Diaz, Gennaro Vessio +2
Sign language production, which generates continuous 3D skeletal motion from spoken language input, must simultaneously satisfy two constraints: semantic fidelity, so that a deaf v…
Understanding How MLLMs Describe Artworks Using Token Activation Maps
Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi +3
Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When an MLLM names a style, identi…
Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models
Lucrezia Laraspata, Giovanna Castellano, Gennaro Vessio
Hallucinations -- factually incorrect or unverifiable outputs -- remain one of the most challenging limitations of Large Language Models (LLMs), especially in knowledge-intensive t…
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
Ivan Rinaldi, Matteo Mendula, Nicola Fanelli +4
Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-cond…
DistillFSS: Synthesizing Few-Shot Knowledge into a Lightweight Segmentation Model
Pasquale De Marinis, Pieter M. Blok, Uzay Kaymak +3
Cross-Domain Few-Shot Semantic Segmentation (CD-FSS) seeks to segment unknown classes in unseen domains using only a few annotated examples. This setting is inherently challenging:…
Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA
Pasquale De Marinis, Gennaro Vessio, Giovanna Castellano
Few-shot semantic segmentation (FSS) aims to segment novel classes in query images using only a small annotated support set. While prior research has mainly focused on improving de…