collaborators
Showing cs.AIShow all

13 papers · 1 filter

cs.AI2026

Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents

Jinwei Hu, Yi Qi, Xinmiao Huang +3

Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's o…

cs.AI2026

Beyond Semantic Equivalence: Logical Graphs for LLM Uncertainty Quantification

Yanni Dong, Minghua Liu, Meiling Zhu +3

Large Language Models often produce confidently stated yet unreliable outputs, posing critical challenges for deployment in safety-sensitive applications. Existing uncertainty metr…

cs.AI2026

Skill Coverage: A Test Adequacy Metric for Agent Skills

Boyin Tan, Xiaowei Huang, Youcheng Sun

Agent skills encode reusable procedural knowledge for large language model (LLM) agents, and existing benchmarks show that such skills can improve task-level performance. However,…

cs.AI2026

SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings

Yingjie Wang, Yi Dong, Edmund Lau +3

Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires prohibitive sample budgets. Sub…

cs.AI2026

Dive into Ambiguity: A*-Inspired Multi-Agents Commonsense Obfuscation Attack on LLM Prompts

Boxuan Wang, Zhuoyun Li, Xiaowei Huang +1

Large language models (LLMs) excel in reasoning and knowledge-intensive tasks but remain vulnerable to prompt-level adversarial attacks that preserve intent while triggering common…

cs.AI2026

Position: A Three-Layer Probabilistic Assume-Guarantee Architecture Is Structurally Required for Safe LLM Agent Deployment

S. Bensalem, Y. Dong, M. Franzle +6

This position paper argues that enforcing LLM agent safety within a single abstraction layer is not merely suboptimal but categorically insufficient for deployed LLM agents -- a st…