2 papers
cs.SE2026
A Framework for Evaluating Agentic Skills at Scale
Maksim Shaposhnikov, Nicolas Fortuin, Simon Stipcich +3
Agent skills -- structured, reusable knowledge artifacts that augment LLM agent capabilities -- have been rapidly adopted in industry, yet their cross-domain impact and use across…
cs.SE2026
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
Maria I. Gorinova, Macey Baker, Amy Heineike +3
Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era: they collapse model, harness, and enviro…