3 papers
cs.SE2026
UXBench: Measuring the Actionability of LLM-Generated UX Critiques
Wenjie Wang, Yue Huang, Zipeng Ling +11
Large language models (LLMs) are increasingly deployed as UX judges that inspect interfaces, diagnose usability problems, and propose repairs. Yet no controlled benchmark measures…
cs.CL2026
A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains
Xianren Zhang, Shreyas Prasad, Di Wang +4
Web agents have shown great promise in performing many tasks on ecommerce website. To assess their capabilities, several benchmarks have been introduced. However, current benchmark…
cs.LG2025
Reflect before Act: Proactive Error Correction in Language Models
Qiuhai Zeng, Sarvesh Rajkumar, Di Wang +2
Large Language Models (LLMs) have demonstrated remarkable capabilities in interactive decision-making tasks, but existing methods often struggle with error accumulation and lack ro…