2 citations · 8 across the 14 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023
Understanding the Inner Workings of Language Models Through Representation Dissimilarity
Davis Brown, Charles Godfrey, Nicholas Konz +2
As language models are applied to an increasing number of real-world applications, understanding their inner workings has become an important issue in model trust, interpretability…
cs.LG2023
Attributing Learned Concepts in Neural Networks to Training Data
Nicholas Konz, Charles Godfrey, Madelyn Shapiro +3
By now there is substantial evidence that deep learning models learn certain human-interpretable features as part of their internal representations of data. As having the right (or…