5 papers · 1 filter
Skill Use or Skill Theater? Evaluating the Reasoning Backroom in Skill-Augmented Language Agents
Jinwei Hu, Yi Qi, Xinmiao Huang +3
Reusable skills are becoming a standard interface for extending language agents with task procedures. Yet evaluators usually infer skill use from visible reasoning or the agent's o…
Responsible Agentic AI Requires Explicit Provenance
Jinwei Hu, Xinmiao Huang, Qisong He +3
Agentic AI is rapidly proliferating across diverse real-world domains such as software engineering, yet public trust has not kept pace. The central reason is that responsibility, d…
PrefixGuard: From LLM-Agent Traces to Online Failure-Warning Monitors
Xinmiao Huang, Jinwei Hu, Rajarshi Roy +3
Large language model (LLM) agents now execute long, tool-using tasks where final outcome checks can arrive too late for intervention. Online warning requires lightweight prefix mon…
Enhancing Robustness of LLM-Driven Multi-Agent Systems through Randomized Smoothing
Jinwei Hu, Yi Dong, Zhengtao Ding +1
This paper presents a defense framework for enhancing the safety of large language model (LLM) empowered multi-agent systems (MAS) in safety-critical domains such as aerospace. We…
Trust-Oriented Adaptive Guardrails for Large Language Models
Jinwei Hu, Yi Dong, Xiaowei Huang
Guardrail, an emerging mechanism designed to ensure that large language models (LLMs) align with human values by moderating harmful or toxic responses, requires a sociotechnical ap…