4 papers
Diff-XYZ: A Benchmark for Evaluating Diff Understanding
Evgeniy Glukhov, Michele Conti, Egor Bogomolov +2
Reliable handling of code diffs is central to agents that edit and refactor repositories at scale. We introduce Diff-XYZ, a compact benchmark for code-diff understanding with three…
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation
Farid Bagirov, Mikhail Arkhipov, Ksenia Sycheva +2
The application of Reinforcement Learning with Verifiable Rewards (RLVR) to mathematical and coding domains has demonstrated significant improvements in the reasoning and problem-s…
On Pretraining for Project-Level Code Completion
Maksim Sapronov, Evgeniy Glukhov
Repository-level pretraining is commonly used to enable large language models for code to leverage codebase-wide context. This enhances their ability to generate accurate and conte…
Challenge on Optimization of Context Collection for Code Completion
Dmitry Ustalov, Egor Bogomolov, Alexander Bezzubov +4
The rapid advancement of workflows and methods for software engineering using AI emphasizes the need for a systematic evaluation and analysis of their ability to leverage informati…