most citedGeneral Agent Evaluation

2 citations · 3 across the 6 of their papers we have counts for

collaborators

13 papers

cs.AI20261 cited

ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Ido Levy, Ben Wiesel, Sami Marreed +4

Autonomous web agents solve complex browsing tasks, yet existing benchmarks measure only whether an agent finishes a task, ignoring whether it does so safely or in a way enterprise…

cs.AI2026

Governance by Construction for Generalist Agents

Segev Shlomov, Iftach Shoham, Alon Oved +7

Enterprise agents are increasingly expected to operate autonomously across tools and interfaces, yet production deployments require governance by construction. Systems must specify…

cs.AI20262 cited

General Agent Evaluation

Elron Bandel, Asaf Yehudai, Lilach Eden +12

General-purpose agents perform tasks in unfamiliar environments without domain-specific manual customization. Yet no study has systematically measured how agent architecture shapes…

cs.AI2026

Agent Mentor: Framing Agent Knowledge through Semantic Trajectory Analysis

Roi Ben-Gigi, Yuval David, Fabiana Fournier +4

AI agent development relies heavily on natural language prompting to define agents' tasks, knowledge, and goals. These prompts are interpreted by Large Language Models (LLMs), whic…

cs.AI2026

AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems

Hadar Mulian, Sergey Zeltyn, Ido Levy +3

We introduce a comprehensive validation framework for LLM-based agentic systems that provides systematic diagnosis and improvement of reliability failures. The framework includes f…

cs.CL2026

TabAgent: A Framework for Replacing Agentic Generative Components with Tabular-Textual Classifiers

Ido Levy, Eilam Shapira, Yinon Goldshtein +3

Agentic systems, AI architectures that autonomously execute multi-step workflows to achieve complex goals, are often built using repeated large language model (LLM) calls for close…