23 citations · 29 across the 2 of their papers we have counts for
3 papers
Biomedical and Clinical Language Models for Spanish: On the Benefits of Domain-Specific Pretraining in a Mid-Resource Scenario
Casimiro Pio Carrino, Jordi Armengol-Estapé, Asier Gutiérrez-Fandiño +4
This work presents biomedical and clinical language models for Spanish by experimenting with different pretraining choices, such as masking at word and subword level, varying the v…
XED: A Multilingual Dataset for Sentiment Analysis and Emotion Detection
Emily Öhman, Marc Pàmies, Kaisla Kajava +1
We introduce XED, a multilingual fine-grained emotion dataset. The dataset consists of human-annotated Finnish (25k) and English sentences (30k), as well as projected annotations f…
LT@Helsinki at SemEval-2020 Task 12: Multilingual or language-specific BERT?
Marc Pàmies, Emily Öhman, Kaisla Kajava +1
This paper presents the different models submitted by the LT@Helsinki team for the SemEval 2020 Shared Task 12. Our team participated in sub-tasks A and C; titled offensive languag…