2 papers
cs.SE2026
OmniCode: A Benchmark for Evaluating Software Engineering Agents
Atharv Sonwane, Eng-Shen Tu, Wei-Chung Lu +11
LLM-powered coding agents are redefining how real-world software is developed. To drive the research towards better coding agents, we require challenging benchmarks that can rigoro…
cs.LG2025
CoCoNUT: Structural Code Understanding does not fall out of a tree
Claas Beger, Saikat Dutta
Large Language Models (LLMs) have shown impressive performance across a wide array of tasks involving both structured and unstructured textual data. Recent results on various bench…