4 papers
H-Node Attack and Defense in Large Language Models
Eric Yocam, Varghese Vaidyan, Yong Wang
We present H-Node Adversarial Noise Cancellation (H-Node ANC), a mechanistic framework that identifies, exploits, and defends hallucination representations in transformer-based lar…
Adaptive Activation Cancellation for Hallucination Mitigation in Large Language Models
Eric Yocam, Varghese Vaidyan, Gurcan Comert +3
Large Language Models frequently generate fluent but factually incorrect text. We propose Adaptive Activation Cancellation (AAC), a real-time inference-time framework that treats h…
Quantum Adversarial Machine Learning and Defense Strategies: Challenges and Opportunities
Eric Yocam, Anthony Rizi, Mahesh Kamepalli +3
As quantum computing continues to advance, the development of quantum-secure neural networks is crucial to prevent adversarial attacks. This paper proposes three quantum-secure des…
Causal Interpretability for Adversarial Robustness: A Hybrid Generative Classification Approach
Chunheng Zhao, Pierluigi Pisu, Gurcan Comert +3
Deep learning-based discriminative classifiers, despite their remarkable success, remain vulnerable to adversarial examples that can mislead model predictions. While adversarial tr…