1 paper · 1 filter
Sarthak Munshi, Manish Bhatt, Vineeth Sai Narajala +4
While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehensive understanding of AI safety r…