2 papers
cs.SE2026
From Anatomy to Smells: An Empirical Study of SKILL.md in Agent Skills
David Boram Hong, Aaron Imani, Iftekhar Ahmed
Agent Skills provide on-demand domain knowledge to LLM agents without requiring model retraining. Each Agent Skill is defined by a mandatory SKILLmd file containing metadata and…
cs.SE2026
Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks
Yuangang Li, Justin Tian Jin Chen, Ethan Yu +2
Large language models (LLMs) increasingly rely on explicit reasoning to solve coding tasks, yet evaluating the quality of this reasoning remains challenging. Existing reasoning eva…