9 citations · 9 across the 1 of their papers we have counts for
2 papers
cs.CR2023
Exploiting Novel GPT-4 APIs
Kellin Pelrine, Mohammad Taufeeque, Michał Zając +2
Language model attacks typically assume one of two extreme threat models: full white-box access to model weights, or black-box access limited to a text generation API. However, rea…
cs.LG2022★ 9 cited
imitation: Clean Imitation Learning Implementations
Adam Gleave, Mohammad Taufeeque, Juan Rocamonde +7
imitation provides open-source implementations of imitation and reward learning algorithms in PyTorch. We include three inverse reinforcement learning (IRL) algorithms, three imita…