4 citations · 7 across the 3 of their papers we have counts for
5 papers
BERT on a Data Diet: Finding Important Examples by Gradient-Based Pruning
Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi +3
Current pre-trained language models rely on large datasets for achieving state-of-the-art performance. However, past research has shown that not all examples in a dataset are equal…
GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh +1
There has been a growing interest in interpreting the underlying dynamics of Transformers. While self-attention patterns were initially deemed as the primary option, recent studies…
AdapLeR: Speeding up Inference by Adaptive Length Reduction
Ali Modarressi, Hosein Mohebbi, Mohammad Taher Pilehvar
Pre-trained language models have shown stellar performance in various downstream tasks. But, this usually comes at the cost of high latency and computation, hindering their usage i…
Not All Models Localize Linguistic Knowledge in the Same Place: A Layer-wise Probing on BERToids' Representations
Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi +2
Most of the recent works on probing representations have focused on BERT, with the presumption that the findings might be similar to the other models. In this work, we extend the p…
Exploring the Role of BERT Token Representations to Explain Sentence Probing Results
Hosein Mohebbi, Ali Modarressi, Mohammad Taher Pilehvar
Several studies have been carried out on revealing linguistic features captured by BERT. This is usually achieved by training a diagnostic classifier on the representations obtaine…