6 papers · 1 filter
Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents
Jinwei Hu, Yi Qi, Xinmiao Huang +3
Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's o…
SCARCE: Scalable Cascade Analysis for Rare-event Characterisation via Embeddings
Yingjie Wang, Yi Dong, Edmund Lau +3
Rare events govern the safety profile of modern AI systems, yet their probabilities are extremely difficult to estimate: direct Monte Carlo requires prohibitive sample budgets. Sub…
Chain-of-Thought as a Lens: Evaluating Structured Reasoning Alignment between Human Preferences and Large Language Models
Boxuan Wang, Zhuoyun Li, Xinmiao Huang +2
This paper primarily demonstrates a method to quantitatively assess the alignment between multi-step, structured reasoning in large language models and human preferences. We introd…
Rethinking Multi-Agent Intelligence Through the Lens of Small-World Networks
Boxuan Wang, Zhuoyun Li, Xiaowei Huang +1
Large language models (LLMs) have enabled multi-agent systems (MAS) in which multiple agents argue, critique, and coordinate to solve complex tasks, making communication topology a…
Enhancing Robustness of LLM-Driven Multi-Agent Systems through Randomized Smoothing
Jinwei Hu, Yi Dong, Zhengtao Ding +1
This paper presents a defense framework for enhancing the safety of large language model (LLM) empowered multi-agent systems (MAS) in safety-critical domains such as aerospace. We…
Trust-Oriented Adaptive Guardrails for Large Language Models
Jinwei Hu, Yi Dong, Xiaowei Huang
Guardrail, an emerging mechanism designed to ensure that large language models (LLMs) align with human values by moderating harmful or toxic responses, requires a sociotechnical ap…