3 papers
cs.LG2025
Multi-Turn Code Generation Through Single-Step Rewards
Arnav Kumar Jain, Gonzalo Gonzalez-Pumariega, Wayne Chen +3
We address the problem of code generation from multi-turn execution feedback. Existing methods either generate code without feedback or use complex, hierarchical reinforcement lear…
cs.HC2024
Challenges in Trustworthy Human Evaluation of Chatbots
Wenting Zhao, Alexander M. Rush, Tanya Goyal
Open community-driven platforms like Chatbot Arena that collect user preference data from site visitors have gained a reputation as one of the most trustworthy publicly available b…
cs.SE2024
Commit0: Library Generation from Scratch
Wenting Zhao, Nan Jiang, Celine Lee +4
With the goal of benchmarking generative systems beyond expert software development ability, we introduce Commit0, a benchmark that challenges AI agents to write libraries from scr…