2 papers
cs.LG2024
If You Don't Understand It, Don't Use It: Eliminating Trojans with Filters Between Layers
Adriano Hernandez
Large language models (LLMs) sometimes exhibit dangerous unintended behaviors. Finding and fixing these is challenging because the attack surface is massive -- it is not tractable…
cs.LG2023
Model Stitching: Looking For Functional Similarity Between Representations
Adriano Hernandez, Rumen Dangovski, Peter Y. Lu +1
Model stitching (Lenc & Vedaldi 2015) is a compelling methodology to compare different neural network representations, because it allows us to measure to what degree they may be in…