200 citations · 512 across the 22 of their papers we have counts for
35 papers
Sabiá-3 Technical Report
Hugo Abonizio, Thales Sales Almeida, Thiago Laitz +4
This report presents Sabiá-3, our new flagship language model, and Sabiazinho-3, a more cost-effective sibling. The models were trained on a large brazilian-centric corpus. Evaluat…
NeuralSearchX: Serving a Multi-billion-parameter Reranker for Multilingual Metasearch at a Low Cost
Thales Sales Almeida, Thiago Laitz, João Seródio +3
The widespread availability of search API's (both free and commercial) brings the promise of increased coverage and quality of search results for metasearch engines, while decreasi…
mRobust04: A Multilingual Version of the TREC Robust 2004 Benchmark
Vitor Jeronymo, Mauricio Nascimento, Roberto Lotufo +1
Robust 2004 is an information retrieval benchmark whose large number of judgments per query make it a reliable evaluation dataset. In this paper, we present mRobust04, a multilingu…
MonoByte: A Pool of Monolingual Byte-level Language Models
Hugo Abonizio, Leandro Rodrigues de Souza, Roberto Lotufo +1
The zero-shot cross-lingual ability of models pretrained on multilingual and even monolingual corpora has spurred many hypotheses to explain this intriguing empirical result. Howev…
Billions of Parameters Are Worth More Than In-domain Training Data: A case study in the Legal Case Entailment Task
Guilherme Moraes Rosa, Luiz Bonifacio, Vitor Jeronymo +3
Recent work has shown that language models scaled to billions of parameters, such as GPT-3, perform remarkably well in zero-shot and few-shot scenarios. In this work, we experiment…
InPars: Data Augmentation for Information Retrieval using Large Language Models
Luiz Bonifacio, Hugo Abonizio, Marzieh Fadaee +1
The information retrieval community has recently witnessed a revolution due to large pretrained transformer models. Another key ingredient for this revolution was the MS MARCO data…