5 papers
What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations
Xiaonan Xu, Wenjing Wu
Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically…
Test-time reasoning effort and unauthorized tool use in language-model agents: a prespecified equivalence study
Xiaonan Xu, Wenjing Wu
Language-model agents that execute multi-step workflows through tool calls operate under access-control policies that restrict which operations each role may perform. The APIs serv…
Validation Evidence in LLM Repair Agents: How Much of What Passes Actually Tests the Bug?
Xiaonan Xu, Wenjing Wu
When a repair agent runs a test and sees it pass, the result is treated as evidence about the reported defect. We measure how often that treatment is warranted. BSG-VA (buggy-state…
Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation
Xiaonan Xu, Wenjing Wu
Agent skills, reusable instruction artefacts supplied to a tool-using language model, are increasingly optimised by shortening, structural rewriting, stronger-model compilation, an…
Large Language Model-Driven Cross-Domain Orchestration Using Multi-Agent Workflow
Xiaonan Xu, Haoshuo Chen, Jesse E. Simsarian +5
We showcase an application that leverages multiple agents, powered by large language models and integrated tools, to collaboratively solve complex network operation tasks across va…