5 papers
When Cases Get Rare: A Retrieval Benchmark for Off-Guideline Clinical Question Answering
Doeun Lee, Muge Zhang, Yi Yu +11
Across medical specialties, clinical practice is anchored in evidence-based guidelines that codify best studied diagnostic and treatment pathways. These pathways routinely fall sho…
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models
Emmy Liu, Varun Gangal, Michael Yu +4
Hallucination remains a central failure mode of large language models, but existing benchmarks operationalize it inconsistently across summarization, question answering, retrieval-…
To Memorize or to Retrieve: Scaling the Interaction Between Pretraining and Retrieval
Karan Singh, Michael Yu, Varun Gangal +4
Retrieval-augmented generation (RAG) improves language model (LM) performance by providing relevant context at test time for knowledge-intensive situations. In this work, we system…
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
Nishant Balepur, Matthew Shu, Yoo Yeon Sung +5
To assist users in complex tasks, LLMs generate plans: step-by-step instructions towards a goal. While alignment methods aim to ensure LLM plans are helpful, they train (RLHF) or e…
TESS 2: A Large-Scale Generalist Diffusion Language Model
Jaesung Tae, Hamish Ivison, Sachin Kumar +1
We introduce TESS 2, a general instruction-following diffusion language model that outperforms contemporary instruction-tuned diffusion models, as well as matches and sometimes exc…