10 papers
Detecting Neural Network Failures through Spectral Analysis of Internal Activations
Arunan J
Neural network misclassifications exhibit characteristic spectral instability in internal activations that is invisible at the output layer. This phenomenon is identified and forma…
Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement Learning
Shijie Liu, Andrew C. Cullen, Paul Montague +2
The current state-of-the-art backdoor attacks against Reinforcement Learning (RL) rely upon unrealistically permissive access models, that assume the attacker can read (or even wri…
TRACER: Persistent Regularization for Robust Multimodal Finetuning
Hesam Asadollahzadeh, Feng Liu, Christopher Leckie +1
Mainstream strategies for finetuning pretrained multimodal models often degrade out-of-distribution (OOD) robustness, a phenomenon known as catastrophic forgetting. In this paper,…
Mechanistic Anomaly Detection via Functional Attribution
Hugo Lyons Keenan, Christopher Leckie, Sarah Erfani
We can often verify the correctness of neural network outputs using ground truth labels, but we cannot reliably determine whether the output was produced by normal or anomalous int…
Fortifying Time Series: DTW-Certified Robust Anomaly Detection
Shijie Liu, Tansu Alpcan, Christopher Leckie +1
Time-series anomaly detection is critical for ensuring safety in high-stakes applications, where robustness is a fundamental requirement rather than a mere performance metric. Addr…
On the Bayes Inconsistency of Disagreement Discrepancy Surrogates
Neil G. Marchant, Andrew C. Cullen, Feng Liu +1
Deep neural networks often fail when deployed in real-world contexts due to distribution shift, a critical barrier to building safe and reliable systems. An emerging approach to ad…