1 citations · 3 across the 4 of their papers we have counts for
4 papers
CLERC: A Dataset for Legal Case Retrieval and Retrieval-Augmented Analysis Generation
Abe Bohan Hou, Orion Weller, Guanghui Qin +5
Legal professionals need to write analyses that rely on citations to relevant precedents, i.e., previous case decisions. Intelligent systems assisting legal professionals in writin…
FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions
Orion Weller, Benjamin Chang, Sean MacAvaney +5
Modern Language Models (LMs) are capable of following long and complex instructions that enable a large and diverse set of user requests. While Information Retrieval (IR) models us…
MegaWika: Millions of reports and their sources across 50 diverse languages
Samuel Barham, Orion Weller, Michelle Yuan +9
To foster the development of new models for collaborative AI-assisted report generation, we introduce MegaWika, consisting of 13 million Wikipedia articles in 50 diverse languages,…
Synthetic Cross-language Information Retrieval Training Data
James Mayfield, Eugene Yang, Dawn Lawrie +5
A key stumbling block for neural cross-language information retrieval (CLIR) systems has been the paucity of training data. The appearance of the MS MARCO monolingual training set…