collaborators

14 papers

cs.LG2026

Stop Guessing When to Stop Testing: Efficient Model Evaluation with Just Enough Data

Ofir Arviv, Kristjan Greenewald, Yotam Perlitz +3

The inherent rigidity of fixed-size benchmarks makes them an inefficient tool for model evaluation. Diverse evaluation objectives, including model ranking, model selection and test…

cs.AI2026

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

Xiaoyuan Liu, Jianhong Tu, Yuqi Chen +26

Agent systems are advancing quickly across domains, but their evaluation remains fragmented. Most benchmarks rely on fixed, LLM-centric harnesses that require heavy integration, cr…

cs.AI2026

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Tomer Keren, Nitay Calderon, Asaf Yehudai +3

As agent capabilities advance, existing benchmarks, such as -Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and lab…

cs.CL2026

Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents

Asaf Yehudai, Lilach Eden, Michal Shmueli-Scheuer

Agentic systems are becoming more capable: agents define strategies, take actions, and interact with different environments. This autonomy poses serious challenges for overseeing a…

cs.CL2026

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

Leyao Wang, Yanan He, Peng Chen +5

Deep research agents increasingly automate complex information-seeking tasks, producing evidence-grounded reports via multi-step reasoning, tool use, and synthesis. Their growing r…

cs.AI2026

General Agent Evaluation

Elron Bandel, Asaf Yehudai, Lilach Eden +12

General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measured how agent architecture shapes…