2 papers
cs.AI2026
AcCoRD: Evaluating User-Agent Collaboration Under Realistic User Preference Dynamics
Tejas Srinivasan, Shikib Mehri, Nandita Shankar Naik +3
User preferences in user-agent collaboration are rarely static and fully-specified upfront: preferences are formed, revealed, adjusted, and relaxed during interaction. Existing ben…
cs.CL2024
LMUnit: Fine-grained Evaluation with Natural Language Unit Tests
Jon Saad-Falcon, Rajan Vivek, William Berrios +6
As language models become integral to critical workflows, assessing their behavior remains a fundamental challenge -- human evaluation is costly and noisy, while automated metrics…