2 papers
cs.CL2026
WebChallenger: A Reliable and Efficient Generalist Web Agent
Jayoo Hwang, Xiaowen Zhang, Vedant Padwal
Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference cost is prohibitive for the…
cs.SE2026
CodeGolf Bench: A Multi-Language Benchmark for Evaluating Concise Code Generation Capabilities of Large Language Models
Vedant Padwal
This paper introduces Code Bench, a benchmark capable of evaluating Large Language Models (LLMs) concise code generation abilities in 60 programming languages. Based on code golf,…