Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Backdoors in DRL: Four Environments Focusing on In-distribution Triggers
Chace Ashcraft, Ted Staley, Josh Carney +4
Backdoor attacks, or trojans, pose a security risk by concealing undesirable behavior in deep neural network models. Open-source neural networks are downloaded from the internet da…
cs.LG2025
Detecting Dataset Bias in Medical AI: A Generalized and Modality-Agnostic Auditing Framework
Nathan Drenkow, Mitchell Pavlak, Keith Harrigian +5
Artificial Intelligence (AI) is now firmly at the center of evidence-based medicine. Despite many success stories that edge the path of AI's rise in healthcare, there are comparabl…
cs.LG2025
Investigating the Treacherous Turn in Deep Reinforcement Learning
Chace Ashcraft, Kiran Karra, Josh Carney +1
The Treacherous Turn refers to the scenario where an artificial intelligence (AI) agent subtly, and perhaps covertly, learns to perform a behavior that benefits itself but is deeme…