20 citations · 20 across the 2 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2024
PeLLE: Encoder-based language models for Brazilian Portuguese based on open data
Guilherme Lamartine de Mello, Marcelo Finger, and Felipe Serras +4
In this paper we present PeLLE, a family of large language models based on the RoBERTa architecture, for Brazilian Portuguese, trained on curated, open data from the Carolina corpu…
cs.CL2022★ 20 cited
Pirá: A Bilingual Portuguese-English Dataset for Question-Answering about the Ocean
André F. A. Paschoal, Paulo Pirozelli, Valdinei Freire +8
Current research in natural language processing is highly dependent on carefully produced corpora. Most existing resources focus on English; some resources focus on languages such…