activity
20202023
most citedWhy is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality

1 citations · 2 across the 3 of their papers we have counts for

collaborators

5 papers

cs.CL2023

When to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants

Anuj Diwan, Eunsol Choi, David Harwath

We present the first unified study of the efficiency of self-attention-based Transformer variants spanning text, speech and vision. We identify input length thresholds (tipping poi…

cs.CL20221 cited

Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality

Anuj Diwan, Layne Berry, Eunsol Choi +2

Recent visuolinguistic pre-trained models show promising progress on various end tasks such as image retrieval and video captioning. Yet, they fail miserably on the recently propos…

cs.CV20221 cited

Zero-shot Video Moment Retrieval With Off-the-Shelf Models

Anuj Diwan, Puyuan Peng, Raymond J. Mooney

For the majority of the machine learning community, the expensive nature of collecting high-quality human-annotated data and the inability to efficiently finetune very large state-…

cs.CL2021

Multilingual and code-switching ASR challenges for low resource Indian languages

Anuj Diwan, Rakesh Vaideeswaran, Sanket Shah +19

Recently, there is increasing interest in multilingual automatic speech recognition (ASR) where a speech recognition system caters to multiple low resource languages by taking adva…

eess.AS2020

Reduce and Reconstruct: ASR for Low-Resource Phonetic Languages

Anuj Diwan, Preethi Jyothi

This work presents a seemingly simple but effective technique to improve low-resource ASR systems for phonetic languages. By identifying sets of acoustically similar graphemes in t…