activity
20242026
collaborators

6 papers

cs.CL2026

eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi +1

We present eCREAM-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals. The corpus, in its current version, is c…

cs.CL2026

DECSELFMASK: Leveraging Unlabeled Text via Self-Relevance-Guided Masking for Decoder-Only Classification

Pietro Ferrazzi, Matteo Merler, Giovanni Bonetta +2

Classification tasks require annotated data, which can often be expensive, time-consuming, or even unfeasible to collect. This is the case of the medical domain, where large datase…

cs.CL2026

Small LLMs for Medical NLP: a Systematic Analysis of Few-Shot, Constraint Decoding, Fine-Tuning and Continual Pre-Training in Italian

Pietro Ferrazzi, Mattia Franzin, Alberto Lavelli +1

Large Language Models (LLMs) consistently excel in diverse medical Natural Language Processing (NLP) tasks, yet their substantial computational requirements often limit deployment…

cs.CL2025

Converting Annotated Clinical Cases into Structured Case Report Forms

Pietro Ferrazzi, Alberto Lavelli, Bernardo Magnini

Case Report Forms (CRFs) are largely used in medical research as they ensure accuracy, reliability, and validity of results in clinical studies. However, publicly available, wellan…

cs.CL2025

Low-resource Information Extraction with the European Clinical Case Corpus

Soumitra Ghosh, Begona Altuna, Saeed Farzi +5

We present E3C-3.0, a multilingual dataset in the medical domain, comprising clinical cases annotated with diseases and test-result relations. The dataset includes both native text…

cs.CL2024

Medical mT5: An Open-Source Multilingual Text-to-Text LLM for The Medical Domain

Iker García-Ferrero, Rodrigo Agerri, Aitziber Atutxa Salazar +10

Research on language technology for the development of medical applications is currently a hot topic in Natural Language Understanding and Generation. Thus, a number of large langu…