26 citations · 26 across the 8 of their papers we have counts for
3 papers · 1 filter
CRQBench: A Benchmark of Code Reasoning Questions
Elizabeth Dinella, Satish Chandra, Petros Maniatis
Large Language Models have demonstrated exceptional proficiency on coding tasks, but it is challenging to precisely evaluate their code reasoning ability. Existing benchmarks are i…
KGym: A Platform and Dataset to Benchmark Large Language Models on Linux Kernel Crash Resolution
Alex Mathai, Chenxi Huang, Petros Maniatis +4
Large Language Models (LLMs) are consistently improving at increasingly realistic software engineering (SE) tasks. In real-world software stacks, significant SE effort is spent dev…
AI-Assisted Assessment of Coding Practices in Modern Code Review
Manushree Vijayvergiya, Małgorzata Salawa, Ivan Budiselić +10
Modern code review is a process in which an incremental code contribution made by a code author is reviewed by one or more peers before it is committed to the version control syste…