Showing cs.SEShow all
3 papers · 1 filter
cs.SE2025
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
Evgeniy Glukhov, Michele Conti, Egor Bogomolov +2
Reliable handling of code diffs is central to agents that edit and refactor repositories at scale. We introduce Diff-XYZ, a compact benchmark for code-diff understanding with three…
cs.SE2025
Challenge on Optimization of Context Collection for Code Completion
Dmitry Ustalov, Egor Bogomolov, Alexander Bezzubov +4
The rapid advancement of workflows and methods for software engineering using AI emphasizes the need for a systematic evaluation and analysis of their ability to leverage informati…
cs.SE2025
Towards Realistic Evaluation of Commit Message Generation by Matching Online and Offline Settings
Petr Tsvetkov, Aleksandra Eliseeva, Danny Dig +4
When a Commit Message Generation (CMG) system is integrated into the IDEs and other products at JetBrains, we perform online evaluation based on user acceptance of the generated me…