Tucano: Advancing Neural Text Generation for Portuguese
arXiv:2411.07854 · doi:10.1016/j.patter.2025.101325
Abstract
Significant advances have been made in natural language processing in recent years. However, our current deep learning approach to language modeling requires substantial resources in terms of data and computation. One of the side effects of this data-hungry paradigm is the current schism between languages, separating those considered high-resource, where most of the development happens and resources are available, and the low-resource ones, which struggle to attain the same level of performance and autonomy. This study aims to introduce a new set of resources to stimulate the future development of neural text generation in Portuguese. In this work, we document the development of GigaVerbo, a concatenation of deduplicated Portuguese text corpora amounting to 200 billion tokens. Via this corpus, we trained a series of decoder-transformers named Tucano. Our models perform equal or superior to other Portuguese and multilingual language models of similar size in several Portuguese benchmarks. The evaluation of our models also reveals that model performance on many currently available benchmarks used by the Portuguese NLP community has little to no correlation with the scaling of token ingestion during training, highlighting the limitations of such evaluations when it comes to the assessment of Portuguese generative language models. All derivatives of our study are openly released on GitHub and Hugging Face. See https://nkluge-correa.github.io/Tucano/
References in corpus (19)
- Sequence to Sequence Learning with Neural Networks
- Training language models to follow instructions with human feedback
- LLaMA: Open and Efficient Foundation Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Estimating the Carbon Footprint of BLOOM, a 176B Parameter Language Model
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- Root Mean Square Layer Normalization
- A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
- The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset
- Textbooks Are All You Need II: phi-1.5 technical report
- Energy Usage Reports: Environmental awareness as part of algorithmic accountability
- Cerebras-GPT: Open Compute-Optimal Language Models Trained on the Cerebras Wafer-Scale Cluster
- PolyLM: An Open Source Polyglot Large Language Model
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Orca-Math: Unlocking the potential of SLMs in Grade School Math
- A New Massive Multilingual Dataset for High-Performance Language Technologies
- Cabrita: closing the gap for foreign languages
- Sabiá-2: A New Generation of Portuguese Large Language Models