most citedOpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs

6 citations · 8 across the 2 of their papers we have counts for

collaborators

5 papers

cs.CL2025

Frustratingly Simple Retrieval Improves Challenging, Reasoning-Intensive Benchmarks

Xinxi Lyu, Michael Duan, Rulin Shao +2

Retrieval-augmented Generation (RAG) has primarily been studied in limited settings, such as factoid question answering; more challenging, reasoning-intensive benchmarks have seen…

cs.AI20252 cited

ReasonIR: Training Retrievers for Reasoning Tasks

Rulin Shao, Rui Qiao, Varsha Kishore +8

We present ReasonIR-8B, the first retriever specifically trained for general reasoning tasks. Existing retrievers have shown limited gains on reasoning tasks, in part because exist…

cs.CV2025

ICONS: Influence Consensus for Vision-Language Data Selection

Xindi Wu, Mengzhou Xia, Rulin Shao +3

Training vision-language models via instruction tuning relies on large data mixtures spanning diverse tasks and domains, yet these mixtures frequently include redundant information…

cs.CL2024

Improving Factuality with Explicit Working Memory

Mingda Chen, Yang Li, Karthik Padthe +5

Large language models can generate factually inaccurate content, a problem known as hallucination. Recent works have built upon retrieved-augmented generation to improve factuality…

cs.CL20246 cited

OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs

Akari Asai, Jacqueline He, Rulin Shao +22

Scientific progress depends on researchers' ability to synthesize the growing body of literature. Can large language models (LMs) assist scientists in this task? We introduce OpenS…