2 papers
cs.LG2024
Limited but consistent gains in adversarial robustness by co-training object recognition models with human EEG
Manshan Guo, Bhavin Choksi, Sari Sadiya +4
In contrast to human vision, artificial neural networks (ANNs) remain relatively susceptible to adversarial attacks. To address this vulnerability, efforts have been made to transf…
cs.AI2024
Position: An Inner Interpretability Framework for AI Inspired by Lessons from Cognitive Neuroscience
Martina G. Vilas, Federico Adolfi, David Poeppel +1
Inner Interpretability is a promising emerging field tasked with uncovering the inner mechanisms of AI systems, though how to develop these mechanistic theories is still much debat…