works on

From the 1 of 12 linked papers with an AI index.

activity
20242026
collaborators

12 papers

cs.CV2026

Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation

Emanuele Colonna, Moises Diaz, Gennaro Vessio +2

The paper presents PIDiffSign, a diffusion-based model that generates 3D sign language motion from spoken language while enforcing anatomical constraints to ensure realistic and bi…

cs.CV2026

Understanding How MLLMs Describe Artworks Using Token Activation Maps

Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi +3

Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When an MLLM names a style, identi…

cs.CL2026

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models

Lucrezia Laraspata, Giovanna Castellano, Gennaro Vessio

Hallucinations -- factually incorrect or unverifiable outputs -- remain one of the most challenging limitations of Large Language Models (LLMs), especially in knowledge-intensive t…

cs.CV2026

Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA

Pasquale De Marinis, Gennaro Vessio, Giovanna Castellano

Few-shot semantic segmentation (FSS) aims to segment novel classes in query images using only a small annotated support set. While prior research has mainly focused on improving de…

physics.optics2026

Modeling and benchmarking quantum optical neurons for efficient neural computation

Andrea Andrisani, Gennaro Vessio, Fabrizio Sgobba +3

Quantum optical neurons (QONs) are emerging as promising computational units that leverage photonic interference to perform neural operations in an energy-efficient and physically…

cs.CV2026

Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment

Ivan Rinaldi, Matteo Mendula, Nicola Fanelli +4

Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-cond…