activity
20172022
most citedTowards a Seamless Integration of Word Senses into Downstream NLP Applications

10 citations · 22 across the 12 of their papers we have counts for

collaborators

25 papers

cs.CL20224 cited

BERT on a Data Diet: Finding Important Examples by Gradient-Based Pruning

Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi +3

Current pre-trained language models rely on large datasets for achieving state-of-the-art performance. However, past research has shown that not all examples in a dataset are equal…

cs.CL2022

Looking at the Overlooked: An Analysis on the Word-Overlap Bias in Natural Language Inference

Sara Rajaee, Yadollah Yaghoobzadeh, Mohammad Taher Pilehvar

It has been shown that NLI models are usually biased with respect to the word-overlap between premise and hypothesis; they take this feature as a primary cue for predicting the ent…

cs.CL20223 cited

GlobEnc: Quantifying Global Token Attribution by Incorporating the Whole Encoder Layer in Transformers

Ali Modarressi, Mohsen Fayyaz, Yadollah Yaghoobzadeh +1

There has been a growing interest in interpreting the underlying dynamics of Transformers. While self-attention patterns were initially deemed as the primary option, recent studies…

cs.CL2022

On the Importance of Data Size in Probing Fine-tuned Models

Houman Mehrafarin, Sara Rajaee, Mohammad Taher Pilehvar

Several studies have investigated the reasons behind the effectiveness of fine-tuning, usually through the lens of probing. However, these studies often neglect the role of the siz…

cs.CL2022

AdapLeR: Speeding up Inference by Adaptive Length Reduction

Ali Modarressi, Hosein Mohebbi, Mohammad Taher Pilehvar

Pre-trained language models have shown stellar performance in various downstream tasks. But, this usually comes at the cost of high latency and computation, hindering their usage i…

cs.CL2021

Not All Models Localize Linguistic Knowledge in the Same Place: A Layer-wise Probing on BERToids' Representations

Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi +2

Most of the recent works on probing representations have focused on BERT, with the presumption that the findings might be similar to the other models. In this work, we extend the p…