8 citations · 8 across the 2 of their papers we have counts for
2 papers
cs.CR2024★ 8 cited
Safeguarding Large Language Models: A Survey
Yi Dong, Ronghui Mu, Yanghao Zhang +9
In the burgeoning field of Large Language Models (LLMs), developing a robust safety mechanism, colloquially known as "safeguards" or "guardrails", has become imperative to ensure t…
cs.CV2024
Towards Fairness-Aware Adversarial Learning
Yanghao Zhang, Tianle Zhang, Ronghui Mu +2
Although adversarial training (AT) has proven effective in enhancing the model's robustness, the recently revealed issue of fairness in robustness has not been well addressed, i.e.…