Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader
Boshui Chen, Huiping Liu, Shaolei Zhang
Large language models increasingly generate complete websites from natural-language descriptions, and reinforcement learning has become a central approach to closing their remainin…
cs.AI2026
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
Yuxin Zhang, Ju Fan, Meihao Fan +2
Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that capture the complexity of…
cs.AI2026
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
Lingyue Fu, Xin Ding, Linyue Pan +9
Current evaluation for Large Language Model (LLM) code agents predominantly focus on generating functional code in single-turn scenarios, which fails to evaluate the agent's capabi…