2 papers
cs.CL2025
ExpertLongBench: Benchmarking Language Models on Expert-Level Long-Form Generation Tasks with Structured Checklists
Jie Ruan, Inderjeet Nair, Shuyang Cao +14
This paper introduces ExpertLongBench, an expert-level benchmark containing 11 tasks from 9 domains that reflect realistic expert workflows and applications. Beyond question answer…
cs.CL2025
Evaluating the Retrieval Robustness of Large Language Models
Shuyang Cao, Karthik Radhakrishnan, David Rosenberg +4
Retrieval-augmented generation (RAG) generally enhances large language models' (LLMs) ability to solve knowledge-intensive tasks. But RAG may also lead to performance degradation d…