11 papers
Multi-User Large Language Model Agents
Shu Yang, Shenzhe Zhu, Hao Zhu +5
Large language models (LLMs) and LLM-based agents are increasingly deployed as assistants in planning and decision making, yet most existing systems are implicitly optimized for a…
STARS: Skill-Triggered Audit for Request-Conditioned Invocation Safety in Agent Systems
Guijia Zhang, Shu Yang, Xilin Gong +1
Autonomous language-model agents increasingly rely on installable skills and tools to complete user tasks. Static skill auditing can expose capability surface before deployment, bu…
Hierarchical Alignment: Enforcing Hierarchical Instruction-Following in LLMs through Logical Consistency
Shu Yang, Zihao Zhou, Di Wang +1
Large language models increasingly operate under multiple instructions from heterogeneous sources with different authority levels, including system policies, user requests, tool ou…
RepreGuard: Detecting LLM-Generated Text by Revealing Hidden Representation Patterns
Xin Chen, Junchao Wu, Shu Yang +7
Detecting content generated by large language models (LLMs) is crucial for preventing misuse and building trustworthy AI systems. Although existing detection methods perform well,…
Understanding and Mitigating Political Stance Cross-topic Generalization in Large Language Models
Jiayi Zhang, Shu Yang, Junchao Wu +2
Fine-tuning Large Language Models on a political topic will significantly manipulate their political stance on various issues and unintentionally affect their stance on unrelated t…
Is Long-to-Short a Free Lunch? Investigating Inconsistency and Reasoning Efficiency in LRMs
Shu Yang, Junchao Wu, Xuansheng Wu +3
Large Reasoning Models (LRMs) have achieved remarkable performance on complex tasks by engaging in extended reasoning before producing final answers, yet this strength introduces t…