4 papers
Matching Pairs: Attributing Fine-Tuned Models to their Pre-Trained Large Language Models
Myles Foley, Ambrish Rawat, Taesung Lee +3
The wide applicability and adaptability of generative large language models (LLMs) has enabled their rapid adoption. While the pre-trained models can perform many tasks, such model…
Adaptive Verifiable Training Using Pairwise Class Similarity
Shiqi Wang, Kevin Eykholt, Taesung Lee +2
Verifiable training has shown success in creating neural networks that are provably robust to a given amount of noise. However, despite only enforcing a single robustness criterion…
Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
Bryant Chen, Wilka Carvalho, Nathalie Baracaldo +5
While machine learning (ML) models are being increasingly trusted to make decisions in different and varying areas, the safety of systems using such models has become an increasing…
Defending Against Machine Learning Model Stealing Attacks Using Deceptive Perturbations
Taesung Lee, Benjamin Edwards, Ian Molloy +1
Machine learning models are vulnerable to simple model stealing attacks if the adversary can obtain output labels for chosen inputs. To protect against these attacks, it has been p…