2.7k citations · 3.1k across the 27 of their papers we have counts for
3 papers · 1 filter
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
Maya Pavlova, Erik Brinkman, Krithika Iyer +7
Red teaming assesses how large language models (LLMs) can produce content that violates norms, policies, and rules set during their safety training. However, most existing automate…
Fairness-Aware Meta-Learning via Nash Bargaining
Yi Zeng, Xuelin Yang, Li Chen +4
To address issues of group-level fairness in machine learning, it is natural to adjust model parameters based on specific fairness objectives over a sensitive-attributed validation…
On Responsible Machine Learning Datasets with Fairness, Privacy, and Regulatory Norms
Surbhi Mittal, Kartik Thakral, Richa Singh +4
Artificial Intelligence (AI) has made its way into various scientific fields, providing astonishing improvements over existing algorithms for a wide variety of tasks. In recent yea…