most citedAutomatic Spanish Translation of the SQuAD Dataset for Multilingual Question Answering

42 citations · 77 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CL202123 cited

Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario

Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño +4

This work presents biomedical and clinical language models for Spanish by experimenting with different pretraining choices, such as masking at word and subword level, varying the v…

cs.CL20214 cited

Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models

Casimiro Pio Carrino, Jordi Armengol-Estapé, Ona de Gibert Bonet +4

We introduce CoWeSe (the Corpus Web Salud Español), the largest Spanish biomedical corpus to date, consisting of 4.5GB (about 750M tokens) of clean plain text. CoWeSe is the result…

cs.CL20213 cited

Are Multilingual Models the Best Choice for Moderately Under-resourced Languages? A Comprehensive Assessment for Catalan

Jordi Armengol-Estapé, Casimiro Pio Carrino, Carlos Rodriguez-Penagos +5

Multilingual language models have been a crucial breakthrough as they considerably reduce the need of data for under-resourced languages. Nevertheless, the superiority of language-…

cs.CL20215 cited

Spanish Biomedical and Clinical Language Embeddings

Asier Gutiérrez-Fandiño, Jordi Armengol-Estapé, Casimiro Pio Carrino +3

We computed both Word and Sub-word Embeddings using FastText. For Sub-word embeddings we selected Byte Pair Encoding (BPE) algorithm to represent the sub-words. We evaluated the Bi…

cs.CL201942 cited

Automatic Spanish Translation of the SQuAD Dataset for Multilingual Question Answering

Casimiro Pio Carrino, Marta R. Costa-jussà, José A. R. Fonollosa

Recently, multilingual question answering became a crucial research topic, and it is receiving increased interest in the NLP community. However, the unavailability of large-scale d…