1 citations · 1 across the 2 of their papers we have counts for
4 papers
CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment
Chengxiao Wang, Enyi Jiang, Xiaojing Liao +1
Improving the safety of large language models (LLMs) often comes at the expense of utility, as globally applied safety tuning may affect model responses to both harmful and benign…
When Behavioral Safety Evaluation Fails: A Representation-Level Perspective
Enyi Jiang, Anders Gjølbye, Anders Gjølbye +2
Safety evaluation of large language models (LLMs) is largely behavioral: a model is certified safe when it refuses harmful requests and answers benign ones. But refusing on the pro…
Towards Generalized Certified Robustness with Multi-Norm Training
Enyi Jiang, David S. Cheung, Gagandeep Singh
Existing certified training methods can only train models to be robust against a certain perturbation type (e.g. or ). However, an certifiably robust mod…
Robust Answers, Fragile Logic: Probing the Decoupling Hypothesis in LLM Reasoning
Enyi Jiang, Changming Xu, Nischay Singh +2
While Chain-of-Thought (CoT) prompting has become a cornerstone for complex reasoning in Large Language Models (LLMs), the faithfulness of the generated reasoning remains an open q…