23 citations · 45 across the 15 of their papers we have counts for
10 papers · 1 filter
Visconde: Multi-document QA with GPT-3 and Neural Reranking
Jayr Pereira, Robson Fidalgo, Roberto Lotufo +1
This paper proposes a question-answering system that can answer questions whose supporting evidence is spread over multiple (potentially long) documents. The system, called Viscond…
MonoByte: A Pool of Monolingual Byte-level Language Models
Hugo Abonizio, Leandro Rodrigues de Souza, Roberto Lotufo +1
The zero-shot cross-lingual ability of models pretrained on multilingual and even monolingual corpora has spurred many hypotheses to explain this intriguing empirical result. Howev…
Induced Natural Language Rationales and Interleaved Markup Tokens Enable Extrapolation in Large Language Models
Mirelle Bueno, Carlos Gemmell, Jeffrey Dalton +2
The ability to extrapolate, i.e., to make predictions on sequences that are longer than those presented as training examples, is a challenging problem for current deep learning mod…
Billions of Parameters Are Worth More Than In-domain Training Data: A case study in the Legal Case Entailment Task
Guilherme Moraes Rosa, Luiz Bonifacio, Vitor Jeronymo +3
Recent work has shown that language models scaled to billions of parameters, such as GPT-3, perform remarkably well in zero-shot and few-shot scenarios. In this work, we experiment…
On the ability of monolingual models to learn language-agnostic representations
Leandro Rodrigues de Souza, Rodrigo Nogueira, Roberto Lotufo
Pretrained multilingual models have become a de facto default approach for zero-shot cross-lingual transfer. Previous work has shown that these models are able to achieve cross-lin…
mMARCO: A Multilingual Version of the MS MARCO Passage Ranking Dataset
Luiz Bonifacio, Vitor Jeronymo, Hugo Queiroz Abonizio +4
The MS MARCO ranking dataset has been widely used for training deep learning models for IR tasks, achieving considerable effectiveness on diverse zero-shot scenarios. However, this…