200 citations · 671 across the 38 of their papers we have counts for
8 papers · 1 filter
BLUEX: A benchmark based on Brazilian Leading Universities Entrance eXams
Thales Sales Almeida, Thiago Laitz, Giovana K. Bonás +1
One common trend in recent studies of language models (LMs) is the use of standardized tests for evaluation. However, despite being the fifth most spoken language worldwide, few su…
InPars Toolkit: A Unified and Reproducible Synthetic Data Generation Pipeline for Neural Information Retrieval
Hugo Abonizio, Luiz Bonifacio, Vitor Jeronymo +3
Recent work has explored Large Language Models (LLMs) to overcome the lack of training data for Information Retrieval (IR) tasks. The generalization abilities of these models have…
A Personalized Dense Retrieval Framework for Unified Information Access
Hansi Zeng, Surya Kallumadi, Zaid Alibadi +2
Developing a universal model that can efficiently and effectively respond to a wide range of information access requests -- from retrieval to recommendation to question answering -…
Sabiá: Portuguese Large Language Models
Ramon Pires, Hugo Abonizio, Thales Sales Almeida +1
As the capabilities of language models continue to advance, it is conceivable that "one-size-fits-all" model will remain as the main paradigm. For instance, given the vast number o…
Evaluating GPT-3.5 and GPT-4 Models on Brazilian University Admission Exams
Desnes Nunes, Ricardo Primi, Ramon Pires +2
The present study aims to explore the capabilities of Language Models (LMs) in tackling high-stakes multiple-choice tests, represented here by the Exame Nacional do Ensino Médio (E…
NeuralMind-UNICAMP at 2022 TREC NeuCLIR: Large Boring Rerankers for Cross-lingual Retrieval
Vitor Jeronymo, Roberto Lotufo, Rodrigo Nogueira
This paper reports on a study of cross-lingual information retrieval (CLIR) using the mT5-XXL reranker on the NeuCLIR track of TREC 2022. Perhaps the biggest contribution of this s…