collaborators

5 papers

cs.CL2026

When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering

Doeun Lee, Muge Zhang, Yi Yu +11

Across medical specialties, clinical practice is anchored in evidence-based guidelines that codify best studied diagnostic and treatment pathways. These pathways routinely fall sho…

cs.CL2026

HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

Emmy Liu, Varun Gangal, Michael Yu +4

Hallucination remains a central failure mode of large language models, but existing benchmarks operationalize it inconsistently across summarization, question answering, retrieval-…

cs.CL2026

To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval

Karan Singh, Michael Yu, Varun Gangal +4

Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-intensive situations. In this work, we system…

cs.CL2025

A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users

Nishant Balepur, Matthew Shu, Yoo Yeon Sung +5

To assist users in complex tasks, LLMs generate plans: step-by-step instructions towards a goal. While alignment methods aim to ensure LLM plans are helpful, they train (RLHF) or e…

cs.CL2025

TESS 2: A Large-Scale Generalist Diffusion Language Model

Jaesung Tae, Hamish Ivison, Sachin Kumar +1

We introduce TESS 2, a general instruction-following diffusion language model that outperforms contemporary instruction-tuned diffusion models, as well as matches and sometimes exc…