1 citations · 1 across the 5 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
Enyi Jiang, Anders Gjølbye, Yibo Jacky Zhang +1
Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones. But refusing on the pro…
cs.LG2024★ 1 cited
Towards Generalized Certified Robustness with Multi-Norm Training
Enyi Jiang, David S. Cheung, Gagandeep Singh
Existing certified training methods can only train models to be robust against a certain perturbation type (e.g. or ). However, an certifiably robust mod…
cs.LG2024
RAMP: Boosting Adversarial Robustness Against Multiple Perturbations for Universal Robustness
Enyi Jiang, Gagandeep Singh
Most existing works focus on improving robustness against adversarial attacks bounded by a single norm using adversarial training (AT). However, these AT models' multiple-nor…