30 papers
Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents
Jinwei Hu, Yi Qi, Xinmiao Huang +3
Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's o…
Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification
Yanni Dong, Minghua Liu, Meiling Zhu +3
Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metr…
Skill Coverage: A Test Adequacy Metric for Agent Skills
Boyin Tan, Xiaowei Huang, Youcheng Sun
Agent skills encode reusable procedural knowledge for large language model (LLM) agents, and existing benchmarks show that such skills can improve task-level performance. However,…
SkillFuzz: Fuzzing Skill Composition for Implicit Intents Discovery in Open Skill Marketplaces
Jinwei Hu, Yi Dong, Youcheng Sun +1
Large Language Model (LLM)-based agents increasingly automate software engineering tasks through reusable skills, natural-language instruction documents that guide planning and exe…
A Data Efficiency Study of Synthetic Fog for Object Detection Using the Clear2Fog Pipeline
Mohamed Ahmed Mohamed, Xiaowei Huang
Object detection in adverse weather is critical for the safety of autonomous vehicles; however, the scarcity of labelled, real-world foggy data remains a significant bottleneck. In…
SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings
Yingjie Wang, Yi Dong, Edmund Lau +3
Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires prohibitive sample budgets. Sub…