2 citations · 2 across the 3 of their papers we have counts for
3 papers
On the Prunability of Attention Heads in Multilingual BERT
Aakriti Budhraja, Madhura Pande, Pratyush Kumar +1
Large multilingual models, such as mBERT, have shown promise in crosslingual transfer. In this work, we employ pruning to quantify the robustness and interpret layer-wise importanc…
The heads hypothesis: A unifying statistical approach towards understanding multi-headed attention in BERT
Madhura Pande, Aakriti Budhraja, Preksha Nema +2
Multi-headed attention heads are a mainstay in transformer-based models. Different methods have been proposed to classify the role of each attention head based on the relations bet…
On the Importance of Local Information in Transformer Based Models
Madhura Pande, Aakriti Budhraja, Preksha Nema +2
The self-attention module is a key component of Transformer-based models, wherein each token pays attention to every other token. Recent studies have shown that these heads exhibit…