2 citations · 3 across the 9 of their papers we have counts for
4 papers · 1 filter
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
Tomer Keren, Nitay Calderon, Asaf Yehudai +3
As agent capabilities advance, existing benchmarks, such as -Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and lab…
General Agent Evaluation
Elron Bandel, Asaf Yehudai, Lilach Eden +12
General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measured how agent architecture shapes…
ELT-Bench-Verified: Benchmark Quality Issues Underestimate AI Agent Capabilities
Christopher Zanoli, Andrea Giovannini, Tengjun Jin +2
Constructing Extract-Load-Transform (ELT) pipelines is a labor-intensive data engineering task and a high-impact target for AI automation. On ELT-Bench, the first benchmark for end…
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
Ora Nova Fandina, Leshem Choshen, Eitan Farchi +3
Consider a scenario where a harmfulness evaluation metric intended to filter unsafe responses from a Large Language Model. When applied to individual harmful prompt-response pairs,…