most citedBiomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario

23 citations · 31 across the 4 of their papers we have counts for

collaborators

5 papers

cs.CL2022

Sequence-to-Sequence Resources for Catalan

Ona de Gibert, Ksenia Kharitonova, Blanca Calvo Figueras +2

In this work, we introduce sequence-to-sequence language resources for Catalan, a moderately under-resourced language, towards two tasks, namely: Summarization and Machine Translat…

cs.CL20214 cited

Spanish Legalese Language Model and Corpora

Asier Gutiérrez-Fandiño, Jordi Armengol-Estapé, Aitor Gonzalez-Agirre +1

There are many Language Models for the English language according to its worldwide relevance. However, for the Spanish language, even if it is a widely spoken language, there are v…

cs.CL202123 cited

Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario

Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño +4

This work presents biomedical and clinical language models for Spanish by experimenting with different pretraining choices, such as masking at word and subword level, varying the v…

cs.CL20214 cited

Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models

Casimiro Pio Carrino, Jordi Armengol-Estapé, Ona de Gibert Bonet +4

We introduce CoWeSe (the Corpus Web Salud Español), the largest Spanish biomedical corpus to date, consisting of 4.5GB (about 750M tokens) of clean plain text. CoWeSe is the result…

cs.CL2021

On the Multilingual Capabilities of Very Large-Scale English Language Models

Jordi Armengol-Estapé, Ona de Gibert Bonet, Maite Melero

Generative Pre-trained Transformers (GPTs) have recently been scaled to unprecedented sizes in the history of machine learning. These models, solely trained on the language modelin…