1 citations · 1 across the 3 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026★ 1 cited
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
Ido Levy, Ben Wiesel, Sami Marreed +4
Autonomous web agents solve complex browsing tasks, yet existing benchmarks measure only whether an agent finishes a task, ignoring whether it does so safely or in a way enterprise…
cs.AI2026
Balanced Thinking: Improving Chain of Thought Training in Vision Language Models
Shaked Perek, Ben Wiesel, Avihu Dekel +2
Multimodal reasoning in vision-language models (VLMs) typically relies on a two-stage process: supervised fine-tuning (SFT) and reinforcement learning (RL). In standard SFT, all to…
cs.AI2024
From Grounding to Planning: Benchmarking Bottlenecks in Web Agents
Segev Shlomov, Ben wiesel, Aviad Sela +3
General web-based agents are increasingly essential for interacting with complex web environments, yet their performance in real-world web applications remains poor, yielding extre…