7 papers · 1 filter
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
Weikang Zhou, Xiao Wang, Limao Xiong +18
Jailbreak attacks are crucial for identifying and mitigating the security vulnerabilities of Large Language Models (LLMs). They are designed to bypass safeguards and elicit prohibi…
Navigating the OverKill in Large Language Models
Chenyu Shi, Xiao Wang, Qiming Ge +7
Large language models are meticulously aligned to be both helpful and harmless. However, recent research points to a potential overkill which means models may refuse to answer beni…
RealBehavior: A Framework for Faithfully Characterizing Foundation Models' Human-like Behavior Mechanisms
Enyu Zhou, Rui Zheng, Zhiheng Xi +7
Reports of human-like behaviors in foundation models are growing, with psychological theories providing enduring tools to investigate these behaviors. However, current research ten…
TRACE: A Comprehensive Benchmark for Continual Learning in Large Language Models
Xiao Wang, Yuansen Zhang, Tianze Chen +9
Aligned large language models (LLMs) demonstrate exceptional capabilities in task-solving, following instructions, and ensuring safety. However, the continual learning aspect of th…
Secrets of RLHF in Large Language Models Part I: PPO
Rui Zheng, Shihan Dou, Songyang Gao +24
Large language models (LLMs) have formulated a blueprint for the advancement of artificial general intelligence. Its primary objective is to function as a human-centric (helpful, h…
Farewell to Aimless Large-scale Pretraining: Influential Subset Selection for Language Model
Xiao Wang, Weikang Zhou, Qi Zhang +7
Pretrained language models have achieved remarkable success in various natural language processing tasks. However, pretraining has recently shifted toward larger models and larger…