8 papers
DIETA: A Decoder-only transformer-based model for Italian-English machine TrAnslation
Pranav Kasela, Marco Braga, Alessandro Ghiotto +3
In this paper, we present DIETA, a small, decoder-only Transformer model with 0.5 billion parameters, specifically designed and trained for Italian-English machine translation. We…
Localized Gaussians as Self-Attention Weights for Point Clouds Correspondence
Alessandro Riva, Alessandro Raganato, Simone Melzi
Current data-driven methodologies for point cloud matching demand extensive training time and computational resources, presenting significant challenges for model deployment and ap…
Blending Concepts with Text-to-Image Diffusion Models
Lorenzo Olearo, Giorgio Longari, Alessandro Raganato +2
Diffusion models have dramatically advanced text-to-image generation in recent years, translating abstract concepts into high-fidelity images with remarkable ease. In this work, we…
Reasoning Capabilities and Invariability of Large Language Models
Alessandro Raganato, Rafael Peñaloza, Marco Viviani +1
Large Language Models (LLMs) have shown remarkable capabilities in manipulating natural language across multiple applications, but their ability to handle simple reasoning tasks is…
Investigating Task Arithmetic for Zero-Shot Information Retrieval
Marco Braga, Pranav Kasela, Alessandro Raganato +1
Large Language Models (LLMs) have shown impressive zero-shot performance across a variety of Natural Language Processing tasks, including document re-ranking. However, their effect…
SemEval-2025 Task 3: Mu-SHROOM, the Multilingual Shared Task on Hallucinations and Related Observable Overgeneration Mistakes
Raúl Vázquez, Timothee Mickus, Elaine Zosa +15
We present the Mu-SHROOM shared task which is focused on detecting hallucinations and other overgeneration mistakes in the output of instruction-tuned large language models (LLMs).…