4 papers · 1 filter
Not All Synthetic Data Is Yours to Learn From
Sina Alemohammad, Li Chen, Richard G. Baraniuk +1
Can a language model improve from plain text sampled from itself, with no prompts, no teacher, no verifier, and no reward model? Yes, but only when the synthetic corpus is compatib…
Configurable Reward Model for Balanced Safety Alignment
Zhengping Jiang, Mehran Khodabandeh, Akash Bharadwaj +5
Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned LLMs and standalone safety…
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh +6
This paper introduces Jailbreak-Zero, a novel red teaming methodology that shifts the paradigm of Large Language Model (LLM) safety evaluation from a constrained example-based appr…
Model Whisper: Steering Vectors Unlock Large Language Models' Potential in Test-time
Xinyue Kang, Diwei Shi, Li Chen
It is a critical challenge to efficiently unlock the powerful reasoning potential of Large Language Models (LLMs) for specific tasks or new distributions. Existing test-time adapta…