5 papers
Skills-Coach: A Self-Evolving Skill Optimizer via Training-Free GRPO
Yu Tian, Jiawei Chen, Lifan Zheng +5
We introduce Skills-Coach, a novel automated framework designed to significantly enhance the self-evolution of skills within Large Language Model (LLM)-based agents. Addressing the…
Exploring the Secondary Risks of Large Language Models
Jiawei Chen, Zhengwei Fang, Yu Tian +4
Ensuring the safety and alignment of Large Language Models is a significant challenge with their growing integration into critical applications and societal functions. While prior…
Red Teaming Large Reasoning Models
Jiawei Chen, Yang Yang, Chao Yu +6
Large Reasoning Models (LRMs) have emerged as a powerful advancement in multi-step reasoning tasks, offering enhanced transparency and logical consistency through explicit chains o…
LCO: LLM-based Constraint Optimization for Safer Agentic LLMs in Real-world Tasks
Jiayong Wan, Jiawei Chen, Zhaoxia Yin +2
Large Language Models (LLMs) are increasingly acting as autonomous agents, but their continuous interaction with the environment can lead to in-context reward hacking (ICRH), a phe…
KG-DF: A Black-box Defense Framework against Jailbreak Attacks Based on Knowledge Graphs
Shuyuan Liu, Jiawei Chen, Xiao Yang +2
With the widespread application of large language models (LLMs) in various fields, the security challenges they face have become increasingly prominent, especially the issue of jai…