3 papers
cs.CL2026
A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
Xianren Zhang, Shreyas Prasad, Di Wang +4
Web agents have shown great promise in performing many tasks on ecommerce website. To assess their capabilities, several benchmarks have been introduced. However, current benchmark…
cs.AI2025
ToolCritic: Detecting and Correcting Tool-Use Errors in Dialogue Systems
Hassan Hamad, Yingru Xu, Liang Zhao +2
Tool-augmented large language models (LLMs) are increasingly employed in real-world applications, but tool usage errors still hinder their reliability. We introduce ToolCritic, a d…
cs.LG2025
Reflect before Act: Proactive Error Correction in Language Models
Qiuhai Zeng, Sarvesh Rajkumar, Di Wang +2
Large Language Models (LLMs) have demonstrated remarkable capabilities in interactive decision-making tasks, but existing methods often struggle with error accumulation and lack ro…