5 papers
JailbreakSkill: Scaling Automated Red-Teaming with Reusable and Ever-Evolving Skills
Xiaoyu Wen, Jiajia Li, Zhida He +11
Automated red-teaming has produced a growing collection of attack strategies, yet they typically remain scattered across prompts and workflows, making them difficult to systematica…
BEACON: Cross-Domain Co-Training of Generative Robot Policies via Best-Effort Adaptation
Antong Zhang, Han Qi, Heng Yang
We introduce BEACON--Best-Effort Adaptation for Cross-Domain Co-Training--a theory-driven framework for training generative robot policies with abundant source demonstrations and l…
Not All Turns Matter: Credit Assignment for Multi-Turn Jailbreaking
Zhida He, Xiaoyu Wen, Han Qi +7
Deploying LLMs in multi-turn dialogues facilitates jailbreak attacks that distribute harmful intent across seemingly benign turns. Recent training-based multi-turn jailbreak method…
Sample-Efficient Reinforcement Learning from Human Feedback via Information-Directed Sampling
Han Qi, Haochen Yang, Qiaosheng Zhang +1
We study the problem of reinforcement learning from human feedback (RLHF), a critical problem in training large language models, from a theoretical perspective. Our main contributi…
Do We Truly Need So Many Samples? Multi-LLM Repeated Sampling Efficiently Scales Test-Time Compute
Jianhao Chen, Zishuo Xun, Bocheng Zhou +8
This paper presents a simple, effective, and cost-efficient strategy to improve LLM performance by scaling test-time compute. Our strategy builds upon the repeated-sampling-then-vo…