collaborators

6 papers

cs.LG2026

Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control

Haoze Liu, Run Liu, Haiying Xu +6

Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM per…

cs.AI2026

SkillEval: Decomposing Agent Skill Quality into Interpretable Signals

Jiahui Han, Qinuo Li, Ziheng Peng +6

Agent skills provide reusable procedural knowledge that helps agents solve specialized tasks. As their use expands, evaluating skill quality becomes increasingly important. Existin…

cs.AI2026

Do LLMs Know Their Vulnerable Scenarios?

Ziheng Peng, Huiqi Deng, Haoran Jing +5

Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teami…

cs.AI2026

AgentDoG: A Diagnostic Guardrail Framework for AI Agent Safety and Security

Dongrui Liu, Qihan Ren, Chen Qian +40

The rise of AI agents introduces complex safety and security challenges arising from autonomous tool use and environmental interactions. Current guardrail models lack agentic risk…

cs.LG2026

Towards Attributions of Input Variables in a Coalition

Xinhao Zheng, Huiqi Deng, Quanshi Zhang

This paper focuses on the fundamental challenge of partitioning input variables in attribution methods for Explainable AI, particularly in Shapley value-based approaches. Previous…

cs.LG2025

Attribution Explanations for Deep Neural Networks: A Theoretical Perspective

Huiqi Deng, Hongbin Pei, Quanshi Zhang +1

Attribution explanation is a typical approach for explaining deep neural networks (DNNs), inferring an importance or contribution score for each input variable to the final output.…