From the 1 of 21 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
ORCA-bench: How Ready Are Language Model Agents for Oncall?
Albert Gong, Kyuseong Choi, Abhineet Agarwal +5
The paper presents ORCA-bench, a benchmark that evaluates large language model agents on on-call root cause analysis tasks using real telemetry data from a live microservice system…
cs.CL2024
FairPair: A Robust Evaluation of Biases in Language Models through Paired Perturbations
Jane Dwivedi-Yu, Raaz Dwivedi, Timo Schick
The accurate evaluation of differential treatment in language models to specific groups is critical to ensuring a positive and safe user experience. An ideal evaluation should have…