6 papers · 1 filter
eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian
Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi +1
We present eCREAM-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals. The corpus, in its current version, is c…
DECSELFMASK: Leveraging Unlabeled Text via Self-Relevance-Guided Masking for Decoder-Only Classification
Pietro Ferrazzi, Matteo Merler, Giovanni Bonetta +2
Classification tasks require annotated data, which can often be expensive, time-consuming, or even unfeasible to collect. This is the case of the medical domain, where large datase…
Toward Automatic Filling of Case Report Forms: A Case Study on Data from an Italian Emergency Department
Gabriela Anna Kaczmarek, Pietro Ferrazzi, Lorenzo Porta +2
Case Report Forms (CRFs) collect data about patients and are at the core of well-established practices to conduct research in clinical settings. With the recent progress of languag…
Small LLMs for Medical NLP: a Systematic Analysis of Few-Shot, Constraint Decoding, Fine-Tuning and Continual Pre-Training in Italian
Pietro Ferrazzi, Mattia Franzin, Alberto Lavelli +1
Large Language Models (LLMs) consistently excel in diverse medical Natural Language Processing (NLP) tasks, yet their substantial computational requirements often limit deployment…
Converting Annotated Clinical Cases into Structured Case Report Forms
Pietro Ferrazzi, Alberto Lavelli, Bernardo Magnini
Case Report Forms (CRFs) are largely used in medical research as they ensure accuracy, reliability, and validity of results in clinical studies. However, publicly available, wellan…
Low-resource Information Extraction with the European Clinical Case Corpus
Soumitra Ghosh, Begona Altuna, Saeed Farzi +5
We present E3C-3.0, a multilingual dataset in the medical domain, comprising clinical cases annotated with diseases and test-result relations. The dataset includes both native text…