2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 2 cited
Automated Red Teaming with GOAT: the Generative Offensive Agent Tester
Maya Pavlova, Erik Brinkman, Krithika Iyer +7
Red teaming assesses how large language models (LLMs) can produce content that violates norms, policies, and rules set during their safety training. However, most existing automate…
cs.LG2024
Backtracking Improves Generation Safety
Yiming Zhang, Jianfeng Chi, Hailey Nguyen +4
Text generation has a fundamental limitation almost by definition: there is no taking back tokens that have been generated, even when they are clearly problematic. In the context o…