2 papers
cs.CL2026
NorBERTo: A ModernBERT Model Trained for Portuguese with 331 Billion Tokens Corpus
Enzo S. N. Silva, Pablo B. Costa, Raphael C. Vlasman +8
High-quality corpora are essential for advancing Natural Language Processing (NLP) in Portuguese. Building on previous encoder-only models such as BERTimbau and Albertina PT-BR, we…
cs.CE2026
Developing an ESG-Oriented Large Language Model through ESG Practices
Gabriel Assis, Ayrton Surica, Pedro Kroll +5
Environmental, Social, and Governance (ESG) considerations play a central role in contemporary financial decision-making. In parallel, Large Language Model (LLM) applications in th…