activity
20202022
most citedMarIA: Spanish Language Models

32 citations · 69 across the 11 of their papers we have counts for

collaborators

11 papers

cs.CL2022★ 1 cited

esCorpius: A Massive Spanish Crawling Corpus

Asier Gutiérrez-Fandiño, David Pérez-Fernández, Jordi Armengol-Estapé +2

In the recent years, transformer-based models have lead to significant advances in language modelling for natural language processing. However, they require a vast amount of data t…

cs.CV2021

The Large Labelled Logo Dataset (L3D): A Multipurpose and Hand-Labelled Continuously Growing Dataset

Asier Gutiérrez-Fandiño, David Pérez-Fernández, Jordi Armengol-Estapé

In this work, we present the Large Labelled Logo Dataset (L3D), a multipurpose, hand-labelled, continuously growing dataset. It is composed of around 770k of color 256x256 RGB imag…

cs.CL2021

FinEAS: Financial Embedding Analysis of Sentiment

Asier Gutiérrez-Fandiño, Miquel Noguer i Alonso, Petter Kolm +1

We introduce a new language representation model in finance called Financial Embedding Analysis of Sentiment (FinEAS). In financial markets, news and investor sentiment are signifi…

cs.CL2021★ 4 cited

Spanish Legalese Language Model and Corpora

Asier Gutiérrez-Fandiño, Jordi Armengol-Estapé, Aitor Gonzalez-Agirre +1

There are many Language Models for the English language according to its worldwide relevance. However, for the Spanish language, even if it is a widely spoken language, there are v…

cs.CL2021★ 23 cited

Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario

Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño +4

This work presents biomedical and clinical language models for Spanish by experimenting with different pretraining choices, such as masking at word and subword level, varying the v…

cs.CL2021★ 4 cited

Spanish Biomedical Crawled Corpus: A Large, Diverse Dataset for Spanish Biomedical Language Models

Casimiro Pio Carrino, Jordi Armengol-Estapé, Ona de Gibert Bonet +4

We introduce CoWeSe (the Corpus Web Salud Español), the largest Spanish biomedical corpus to date, consisting of 4.5GB (about 750M tokens) of clean plain text. CoWeSe is the result…