4 papers
Excision Score: Evaluating Edits with Surgical Precision
Nikolai Gruzinov, Ksenia Sycheva, Earl T. Barr +1
Many tasks revolve around editing a document, whether code or text. We formulate the revision similarity problem to unify a wide range of machine learning evaluation problems whose…
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
Evgeniy Glukhov, Michele Conti, Egor Bogomolov +2
Reliable handling of code diffs is central to agents that edit and refactor repositories at scale. We introduce Diff-XYZ, a compact benchmark for code-diff understanding with three…
Challenge on Optimization of Context Collection for Code Completion
Dmitry Ustalov, Egor Bogomolov, Alexander Bezzubov +4
The rapid advancement of workflows and methods for software engineering using AI emphasizes the need for a systematic evaluation and analysis of their ability to leverage informati…
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings
Petr Tsvetkov, Aleksandra Eliseeva, Danny Dig +4
When a Commit Message Generation (CMG) system is integrated into the IDEs and other products at JetBrains, we perform online evaluation based on user acceptance of the generated me…