From the 1 of 12 linked papers with an AI index.
12 papers
Physics-Informed Diffusion for Biomechanically Plausible 3D Sign Language Generation
Emanuele Colonna, Moises Diaz, Gennaro Vessio +2
The paper presents PIDiffSign, a diffusion-based model that generates 3D sign language motion from spoken language while enforcing anatomical constraints to ensure realistic and bi…
Understanding How MLLMs Describe Artworks Using Token Activation Maps
Nicola Fanelli, Pasquale De Marinis, Raffaele Scaringi +3
Multimodal Large Language Models (MLLMs) describe artworks with remarkable fluency, yet the visual reasoning behind their outputs remains opaque. When an MLLM names a style, identi…
Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models
Lucrezia Laraspata, Giovanna Castellano, Gennaro Vessio
Hallucinations -- factually incorrect or unverifiable outputs -- remain one of the most challenging limitations of Large Language Models (LLMs), especially in knowledge-intensive t…
Take a Peek: Efficient Encoder Adaptation for Few-Shot Semantic Segmentation via LoRA
Pasquale De Marinis, Gennaro Vessio, Giovanna Castellano
Few-shot semantic segmentation (FSS) aims to segment novel classes in query images using only a small annotated support set. While prior research has mainly focused on improving de…
Modeling and benchmarking quantum optical neurons for efficient neural computation
Andrea Andrisani, Gennaro Vessio, Fabrizio Sgobba +3
Quantum optical neurons (QONs) are emerging as promising computational units that leverage photonic interference to perform neural operations in an energy-efficient and physically…
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
Ivan Rinaldi, Matteo Mendula, Nicola Fanelli +4
Music generation has advanced markedly through multimodal deep learning, enabling models to synthesize audio from text and, more recently, from images. However, existing image-cond…