Showing cs.SEShow all
3 papers · 1 filter
cs.SE2026
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
Xiaonan Xu, Wenjing Wu
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically…
cs.SE2026
Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?
Xiaonan Xu, Wenjing Wu
When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is warranted. BSG-VA (buggy-state…
cs.SE2026
Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation
Xiaonan Xu, Wenjing Wu
Agent skills, reusable instruction artefacts supplied to a tool-using language model, are increasingly optimised by shortening, structural rewriting, stronger-model compilation, an…