1 citations · 1 across the 4 of their papers we have counts for
5 papers · 1 filter
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
Ido Levy, Ben Wiesel, Sami Marreed +4
Autonomous web agents solve complex browsing tasks, yet existing benchmarks measure only whether an agent finishes a task, ignoring whether it does so safely or in a way enterprise…
Governance by Construction for Generalist Agents
Segev Shlomov, Iftach Shoham, Alon Oved +7
Enterprise agents are increasingly expected to operate autonomously across tools and interfaces, yet production deployments require governance by construction. Systems must specify…
AgentFixer: From Failure Detection to Fix Recommendations in LLM Agentic Systems
Hadar Mulian, Sergey Zeltyn, Ido Levy +3
We introduce a comprehensive validation framework for LLM-based agentic systems that provides systematic diagnosis and improvement of reliability failures. The framework includes f…
From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production
Segev Shlomov, Alon Oved, Sami Marreed +9
Agents are rapidly advancing in automating digital work, but enterprises face a harder challenge: moving beyond prototypes to deployed systems that deliver measurable business valu…
From Grounding to Planning: Benchmarking Bottlenecks in Web Agents
Segev Shlomov, Ben wiesel, Aviad Sela +3
General web-based agents are increasingly essential for interacting with complex web environments, yet their performance in real-world web applications remains poor, yielding extre…