2 papers
cs.CR2025
When "Correct" Is Not Safe: Can We Trust Functionally Correct Patches Generated by Code Agents?
Yibo Peng, James Song, Lei Li +6
Code agents are increasingly trusted to autonomously fix bugs on platforms such as GitHub, yet their security evaluation focuses almost exclusively on functional correctness. In th…
cs.CV2025
Saliency-Bench: A Comprehensive Benchmark for Evaluating Visual Explanations
Yifei Zhang, James Song, Siyi Gu +4
Explainable AI (XAI) has gained significant attention for providing insights into the decision-making processes of deep learning models, particularly for image classification tasks…