15 citations · 15 across the 4 of their papers we have counts for
3 papers · 1 filter
GRASP: Deterministic argument ranking in interaction graphs
Diganta Misra, Antonio Orvieto, Rediet Abebe +1
Large language models are increasingly deployed as automated judges to evaluate the strength of arguments. As this role expands, their legitimacy depends on consistency, transparen…
Explaining Grokking in Transformers through the Lens of Inductive Bias
Jaisidh Singh, Diganta Misra, Antonio Orvieto
We investigate grokking in transformers through the lens of inductive bias: dispositions arising from architecture or optimization that let the network prefer one solution over ano…
Using Shapley interactions to understand how models use structure
Divyansh Singhvi, Diganta Misra, Andrej Erkelens +3
Language is an intricately structured system, and a key goal of NLP interpretability is to provide methodological insights for understanding how language models represent this stru…