1 citations · 1 across the 3 of their papers we have counts for
3 papers
Quati: A Brazilian Portuguese Information Retrieval Dataset from Native Speakers
Mirelle Bueno, Eduardo Seiti de Oliveira, Rodrigo Nogueira +2
Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset…
Lissard: Long and Simple Sequential Reasoning Datasets
Mirelle Bueno, Roberto Lotufo, Rodrigo Nogueira
Language models are now capable of solving tasks that require dealing with long sequences consisting of hundreds of thousands of tokens. However, they often fail on tasks that requ…
Induced Natural Language Rationales and Interleaved Markup Tokens Enable Extrapolation in Large Language Models
Mirelle Bueno, Carlos Gemmell, Jeffrey Dalton +2
The ability to extrapolate, i.e., to make predictions on sequences that are longer than those presented as training examples, is a challenging problem for current deep learning mod…