33 citations · 62 across the 12 of their papers we have counts for
14 papers
Fine PT-PT Web: A High-Quality 41 Billion Tokens Data Collection of the European Portuguese Web
Gonçalo Vinagre, Rui Pedro Guerra, Pedro Gomes +7
Curating Web corpora for regional language variants like European Portuguese (PT-PT) is heavily bottlenecked by dialectal overlap (mainly with PT-BR) and data processing scale. Thi…
SEQUOR: A Multi-Turn Benchmark for Realistic Constraint Following
Beatriz Canaverde, Duarte M. Alves, José Pombal +2
In a conversation, a helpful assistant must reliably follow user directives, even as they refine, modify, or contradict earlier requests. Yet most instruction-following benchmarks…
A Recipe for Long-Context Reasoning in Large Language Models via On-Policy Optimization and Distillation
Miguel Moura Ramos, Duarte M. Alves, André F. T. Martins
Existing approaches to post-train models for long-context tasks face complementary limitations: (i) supervised fine-tuning (SFT) provides stable supervision but suffers from exposu…
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
Afonso Simplício, Gonçalo Vinagre, Miguel Moura Ramos +19
Despite rapid progress in open large language models (LLMs), European Portuguese (pt-PT) remains underrepresented in both training data and native evaluation, with machine-translat…
EuroLLM-22B: Technical Report
Miguel Moura Ramos, Duarte M. Alves, Hippolyte Gisserot-Boukhlef +15
This report presents EuroLLM-22B, a large language model trained from scratch to support the needs of European citizens by covering all 24 official European Union languages and 11…
Should We Still Pretrain Encoders with Masked Language Modeling?
Hippolyte Gisserot-Boukhlef, Nicolas Boizard, Manuel Faysse +5
Learning high-quality text representations is fundamental to a wide range of NLP tasks. While encoder pretraining has traditionally relied on Masked Language Modeling (MLM), recent…