4 citations · 5 across the 6 of their papers we have counts for
3 papers · 1 filter
Configurable Reward Model for Balanced Safety Alignment
Zhengping Jiang, Mehran Khodabandeh, Akash Bharadwaj +5
Aligning large language models (LLMs) to heterogeneous and rapidly evolving safety requirements remains a critical challenge. Existing instruction-tuned LLMs and standalone safety…
Jailbreak-Zero: A Path to Pareto Optimal Red Teaming for Large Language Models
Kai Hu, Abhinav Aggarwal, Mehran Khodabandeh +6
This paper introduces Jailbreak-Zero, a novel red teaming methodology that shifts the paradigm of Large Language Model (LLM) safety evaluation from a constrained example-based appr…
To Test Machine Comprehension, Start by Defining Comprehension
Jesse Dunietz, Gregory Burnham, Akash Bharadwaj +3
Many tasks aim to measure machine reading comprehension (MRC), often focusing on question types presumed to be difficult. Rarely, however, do task designers start by considering wh…