activity
20242026
collaborators

5 papers

cs.CL2026

Teaching LLMs Brazilian Healthcare: Injecting Knowledge from Official Clinical Guidelines

Hugo Abonizio, Filipe Rocha Lopes, Roberto Lotufo +1

Brazil's Unified Health System (SUS) relies on official clinical guidelines that define diagnostic criteria, treatments, dosages, and monitoring procedures for over 200 million cit…

cs.CL2025

PublicHearingBR: A Brazilian Portuguese Dataset of Public Hearing Transcripts for Summarization of Long Documents

Leandro Carísio Fernandes, Guilherme Zeferino Rodrigues Dobins, Roberto Lotufo +1

This paper introduces PublicHearingBR, a Brazilian Portuguese dataset designed for summarizing long documents. The dataset consists of transcripts of public hearings held by the Br…

cs.CL2025

Comparing Knowledge Injection Methods for LLMs in a Low-Resource Regime

Hugo Abonizio, Thales Almeida, Roberto Lotufo +1

Large language models (LLMs) often require vast amounts of text to effectively acquire new knowledge. While continuing pre-training on large corpora or employing retrieval-augmente…

cs.CL2024

ptt5-v2: A Closer Look at Continued Pretraining of T5 Models for the Portuguese Language

Marcos Piau, Roberto Lotufo, Rodrigo Nogueira

Despite advancements in Natural Language Processing (NLP) and the growing availability of pretrained models, the English language remains the primary focus of model development. Co…

cs.CL2024

MLissard: Multilingual Long and Simple Sequential Reasoning Benchmarks

Mirelle Bueno, Roberto Lotufo, Rodrigo Nogueira

Language models are now capable of solving tasks that require dealing with long sequences consisting of hundreds of thousands of tokens. However, they often fail on tasks that requ…