Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Agent Safety Alignment via Reinforcement Learning
Zeyang Sha, Hanling Tian, Zhuoer Xu +3
The emergence of autonomous Large Language Model (LLM) agents capable of tool usage has introduced new safety risks that go beyond traditional conversational misuse. These agents,…
cs.AI2024
TroubleLLM: Align to Red Team Expert
Zhuoer Xu, Jianping Zhang, Shiwen Cui +2
Large Language Models (LLMs) become the start-of-the-art solutions for a variety of natural language tasks and are integrated into real-world applications. However, LLMs can be pot…