3 citations · 6 across the 8 of their papers we have counts for
6 papers · 1 filter
Who's the Evil Twin? Differential Auditing for Undesired Behavior
Ishwar Balappanawar, Venkata Hasith Vattikuti, Greta Kintzley +2
Detecting hidden behaviors in neural networks poses a significant challenge due to minimal prior knowledge and potential adversarial obfuscation. We explore this problem by framing…
Studying Cross-cluster Modularity in Neural Networks
Satvik Golechha, Maheep Chaudhary, Joan Velja +2
An approach to improve neural network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We define a measure…
Training Neural Networks for Modularity aids Interpretability
Satvik Golechha, Dylan Cope, Nandi Schoots
An approach to improve network interpretability is via clusterability, i.e., splitting a model into disjoint clusters that can be studied independently. We find pretrained models t…
Progress Measures for Grokking on Real-world Tasks
Satvik Golechha
Grokking, a phenomenon where machine learning models generalize long after overfitting, has been primarily observed and studied in algorithmic tasks. This paper explores grokking i…
Challenges in Mechanistically Interpreting Model Representations
Satvik Golechha, James Dao
Mechanistic interpretability (MI) aims to understand AI models by reverse-engineering the exact algorithms neural networks learn. Most works in MI so far have studied behaviors and…
Predicting Treatment Adherence of Tuberculosis Patients at Scale
Mihir Kulkarni, Satvik Golechha, Rishi Raj +9
Tuberculosis (TB), an infectious bacterial disease, is a significant cause of death, especially in low-income countries, with an estimated ten million new cases reported globally i…