2 citations · 2 across the 7 of their papers we have counts for
4 papers · 1 filter
NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
Jingzhe Ding, Shengda Long, Changxin Pu +46
Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks fail to rigorously evaluate the long-horizon capabilities re…
DiscoX: Benchmarking Discourse-Level Translation task in Expert Domains
Xiying Zhao, Zhoufutu Wen, Zhixuan Chen +20
The evaluation of discourse-level translation in expert domains remains inadequate, despite its centrality to knowledge dissemination and cross-lingual scholarly communication. Whi…
Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
Shuangshuang Ying, Yunwen Li, Xingwei Qu +21
Current preference learning methods achieve high accuracy on standard benchmarks but exhibit significant performance degradation when objective quality signals are removed. We intr…
Using Large Language Model to Solve and Explain Physics Word Problems Approaching Human Level
Jingzhe Ding, Yan Cen, Xinyuan Wei
Our work demonstrates that large language model (LLM) pre-trained on texts can not only solve pure math word problems, but also physics word problems, whose solution requires calcu…