10 citations · 22 across the 12 of their papers we have counts for
25 papers
BERT on a Data Diet: Finding Important Examples by Gradient-Based Pruning
Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi +3
Current pre-trained language models rely on large datasets for achieving state-of-the-art performance. However, past research has shown that not all examples in a dataset are equal…
Looking at the Overlooked: An Analysis on the Word-Overlap Bias in Natural Language Inference
Sara Rajaee, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar
It has been shown that NLI models are usually biased with respect to the word-overlap between premise and hypothesis; they take this feature as a primary cue for predicting the ent…
GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers
Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh +1
There has been a growing interest in interpreting the underlying dynamics of Transformers. While self-attention patterns were initially deemed as the primary option, recent studies…
On the Importance of Data Size in Probing Fine-tuned Models
Houman Mehrafarin, Sara Rajaee, Mohammad Taher Pilehvar
Several studies have investigated the reasons behind the effectiveness of fine-tuning, usually through the lens of probing. However, these studies often neglect the role of the siz…
AdapLeR: Speeding up Inference by Adaptive Length Reduction
Ali Modarressi, Hosein Mohebbi, Mohammad Taher Pilehvar
Pre-trained language models have shown stellar performance in various downstream tasks. But, this usually comes at the cost of high latency and computation, hindering their usage i…
Not All Models Localize Linguistic Knowledge in the Same Place: A Layer-wise Probing on BERToids' Representations
Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi +2
Most of the recent works on probing representations have focused on BERT, with the presumption that the findings might be similar to the other models. In this work, we extend the p…