1 citations · 1 across the 2 of their papers we have counts for
4 papers
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
Ido Levy, Ben Wiesel, Sami Marreed +4
Autonomous web agents solve complex browsing tasks, yet existing benchmarks measure only whether an agent finishes a task, ignoring whether it does so safely or in a way enterprise…
Governance by Construction for Generalist Agents
Segev Shlomov, Iftach Shoham, Alon Oved +7
Enterprise agents are increasingly expected to operate autonomously across tools and interfaces, yet production deployments require governance by construction. Systems must specify…
From Benchmarks to Business Impact: Deploying IBM Generalist Agent in Enterprise Production
Segev Shlomov, Alon Oved, Sami Marreed +9
Agents are rapidly advancing in automating digital work, but enterprises face a harder challenge: moving beyond prototypes to deployed systems that deliver measurable business valu…
Towards Enterprise-Ready Computer Using Generalist Agent
Sami Marreed, Alon Oved, Avi Yaeli +6
This paper presents our ongoing work toward developing an enterprise-ready Computer Using Generalist Agent (CUGA) system. Our research highlights the evolutionary nature of buildin…