3 papers
cs.AI2026
RIFT: A RubrIc Failure Mode Taxonomy and Automated Diagnostics
Zhengyang Qi, Charles Dickens, Derek Pham +4
Rubric-based evaluation is widely used in LLM benchmarks and training pipelines for open-ended, less verifiable tasks. While prior work has demonstrated the effectiveness of rubric…
cs.AI2026
Benchmarking Agents in Insurance Underwriting Environments
Amanda Dsouza, Ramya Ramakrishnan, Charles Dickens +2
As AI agents integrate into enterprise applications, their evaluation demands benchmarks that reflect the complexity of real-world operations. Instead, existing benchmarks overemph…
cs.SE2025
Automating Benchmark Design
Amanda Dsouza, Harit Vishwakarma, Zhengyang Qi +6
The rapid progress and widespread deployment of LLMs and LLM-powered agents has outpaced our ability to evaluate them. Hand-crafted, static benchmarks are the primary tool for asse…