1 citations · 2 across the 5 of their papers we have counts for
6 papers · 1 filter
SPARTAN: Sparse Hierarchical Memory for Parameter-Efficient Transformers
Ameet Deshpande, Md Arafat Sultan, Anthony Ferritto +3
Fine-tuning pre-trained language models (PLMs) achieves impressive performance on a range of downstream tasks, and their sizes have consequently been getting bigger. Since a differ…
ALIGN-MLM: Word Embedding Alignment is Crucial for Multilingual Pre-training
Henry Tang, Ameet Deshpande, Karthik Narasimhan
Multilingual pre-trained models exhibit zero-shot cross-lingual transfer, where a model fine-tuned on a source language achieves surprisingly good performance on a target language.…
Guiding Attention for Self-Supervised Learning with Transformers
Ameet Deshpande, Karthik Narasimhan
In this paper, we propose a simple and effective technique to allow for efficient self-supervised learning with bi-directional Transformers. Our approach is motivated by recent stu…
Sentiment Analysis for Reinforcement Learning
Ameet Deshpande, Eve Fleisig
While reinforcement learning (RL) has been successful in natural language processing (NLP) domains such as dialogue generation and text-based games, it typically faces the problem…
CLEVR Parser: A Graph Parser Library for Geometric Learning on Language Grounded Image Scenes
Raeid Saqur, Ameet Deshpande
The CLEVR dataset has been used extensively in language grounded visual reasoning in Machine Learning (ML) and Natural Language Processing (NLP) domains. We present a graph parser…
Weight Initialization in Neural Language Models
Ameet Deshpande, Vedant Somani
Semantic Similarity is an important application which finds its use in many downstream NLP applications. Though the task is mathematically defined, semantic similarity's essence is…