2 papers
cs.CL2026
SwingArena: Competitive Programming Arena for Long-context GitHub Issue Solving
Wendong Xu, Jing Xiong, Chenyang Zhao +16
We present SwingArena, a competitive evaluation framework for Large Language Models (LLMs) that closely mirrors real-world software development workflows. Unlike traditional static…
cs.LG2024
CoPS: Empowering LLM Agents with Provable Cross-Task Experience Sharing
Chen Yang, Chenyang Zhao, Quanquan Gu +1
Sequential reasoning in agent systems has been significantly advanced by large language models (LLMs), yet existing approaches face limitations. Reflection-driven reasoning relies…