10 citations · 10 across the 1 of their papers we have counts for
3 papers
Multi-Agent Risks from Advanced AI
Lewis Hammond, Alan Chan, Jesse Clifton +41
The rapid development of advanced AI agents and the imminent deployment of many instances of these agents will give rise to multi-agent systems of unprecedented complexity. These s…
Can Go AIs be adversarially robust?
Tom Tseng, Euan McLean, Kellin Pelrine +2
Prior work found that superhuman Go AIs can be defeated by simple adversarial strategies, especially "cyclic" attacks. In this paper, we study whether adding natural countermeasure…
Exploiting Novel GPT-4 APIs
Kellin Pelrine, Mohammad Taufeeque, Michał Zając +2
Language model attacks typically assume one of two extreme threat models: full white-box access to model weights, or black-box access limited to a text generation API. However, rea…