Showing cs.CRShow all
2 papers · 1 filter
cs.CR2026
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
Jianwei Li, Jung-Eun Kim
Backdoor attacks pose severe security threats to large language models (LLMs), where a model behaves normally under benign inputs but produces malicious outputs when a hidden trigg…
cs.CR2025
Trustworthy AI: Safety, Bias, and Privacy -- A Survey
Xingli Fang, Jianwei Li, Varun Mulchandani +1
The capabilities of artificial intelligence systems have been advancing to a great extent, but these systems still struggle with failure modes, vulnerabilities, and biases. In this…