Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement?
Jiajun Shi, Siyuan Tao, Yuhao Wu +18
Large language models (LLMs) increasingly interact with external environments and accumulate substantial behavioral experience, yet existing agent benchmarks largely evaluate them…
cs.CL2026
WebWorld: The Browser as a World Model for Self-Improving Web Code
Jiajun Wu, Jian Yang, Yaxin Du +7
VLM-driven self-improvement of web code has a structural flaw: the model that proposes the repair is the model that judges it, and visual plausibility under that judge is a poor pr…