2 citations · 6 across the 12 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.CR2024★ 2 cited
AI Risk Management Should Incorporate Both Safety and Security
Xiangyu Qi, Yangsibo Huang, Yi Zeng +22
The exposure of security vulnerabilities in safety-aligned language models, e.g., susceptibility to adversarial attacks, has shed light on the intricate interplay between AI safety…
cs.HC2024
OpenHEXAI: An Open-Source Framework for Human-Centered Evaluation of Explainable Machine Learning
Jiaqi Ma, Vivian Lai, Yiming Zhang +5
Recently, there has been a surge of explainable AI (XAI) methods driven by the need for understanding machine learning model behaviors in high-stakes scenarios. However, properly e…