9 citations · 9 across the 1 of their papers we have counts for
1 paper · 1 filter
Kellin Pelrine, Mohammad Taufeeque, Michał Zając +2
Language model attacks typically assume one of two extreme threat models: full white-box access to model weights, or black-box access limited to a text generation API. However, rea…