1 paper
Kaiyuan Liu, Youcheng Pan, Yang Xiang +4
Recently, LLM agents have made rapid progress in improving their programming capabilities. However, existing benchmarks lack the ability to automatically evaluate from users' persp…