collaborators

11 papers

cs.CL2026

Spokes: Optimizing for Diverse Pretraining Data Selection

Clarence Lee, Yejin Choi, Luke Zettlemoyer +2

Diversity plays a critical role in data selection, improving performance under fixed data budgets by reducing redundancy and repetition. However, optimizing for diversity is inhere…

cs.CL2026

DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Rulin Shao, Akari Asai, Shannon Zejiang Shen +18

Deep research agents perform multi-step research to produce long-form, well-attributed answers. However, most open deep research agents are trained on easily verifiable short-form…

cs.CR2026

A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage

Rui Xin, Niloofar Mireshghallah, Shuyue Stella Li +6

Sanitizing sensitive text data typically involves removing personally identifiable information (PII) or generating synthetic data under the assumption that these methods adequately…

cs.CL2025

2 OLMo 2 Furious

Team OLMo, Pete Walsh, Luca Soldaini +40

We present OLMo 2, the next generation of our fully open language models. OLMo 2 includes a family of dense autoregressive language models at 7B, 13B and 32B scales with fully rele…

cs.CL2025

Precise Information Control in Long-Form Text Generation

Jacqueline He, Howard Yen, Margaret Li +7

A central challenge in language models (LMs) is faithfulness hallucination: the generation of information unsubstantiated by input context. To study this problem, we propose Precis…

cs.CL2025

ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data

Tong Chen, Faeze Brahman, Jiacheng Liu +5

Language models (LMs) can memorize and reproduce segments from their pretraining data verbatim even in non-adversarial settings, raising concerns about copyright, plagiarism, priva…