activity
20182022
most citedOn the Importance of Local Information in Transformer Based Models

2 citations · 3 across the 6 of their papers we have counts for

collaborators

12 papers

cs.CL2022

T-STAR: Truthful Style Transfer using AMR Graph as Intermediate Representation

Anubhav Jangra, Preksha Nema, Aravindan Raghuveer

Unavailability of parallel corpora for training text style transfer (TST) models is a very challenging yet common scenario. Also, TST models implicitly need to preserve the content…

cs.CL2021

A Framework for Rationale Extraction for Deep QA models

Sahana Ramnath, Preksha Nema, Deep Sahni +1

As neural-network-based QA models become deeper and more complex, there is a demand for robust frameworks which can access a model's rationale for its prediction. Current technique…

cs.CL2021

The heads hypothesis: A unifying statistical approach towards understanding multi-headed attention in BERT

Madhura Pande, Aakriti Budhraja, Preksha Nema +2

Multi-headed attention heads are a mainstay in transformer-based models. Different methods have been proposed to classify the role of each attention head based on the relations bet…

cs.CL2020

Towards Interpreting BERT for Reading Comprehension Based QA

Sahana Ramnath, Preksha Nema, Deep Sahni +1

BERT and its variants have achieved state-of-the-art performance in various NLP tasks. Since then, various works have been proposed to analyze the linguistic information being capt…

cs.CL20202 cited

On the Importance of Local Information in Transformer Based Models

Madhura Pande, Aakriti Budhraja, Preksha Nema +2

The self-attention module is a key component of Transformer-based models, wherein each token pays attention to every other token. Recent studies have shown that these heads exhibit…

cs.CL2020

Towards Transparent and Explainable Attention Models

Akash Kumar Mohankumar, Preksha Nema, Sharan Narasimhan +3

Recent studies on interpretability of attention distributions have led to notions of faithful and plausible explanations for a model's predictions. Attention distributions can be c…