1 paper · 1 filter
Leanne Tan, Rohan Jaggi, Shaun Khoo +1
Evaluations (Evals) are a deployment bottleneck for real-world AI applications: public benchmarks rarely match a team's users, context, or policies, and human review is often tedio…