1 citations · 2 across the 3 of their papers we have counts for
5 papers
When to Use Efficient Self Attention? Profiling Text, Speech and Image Transformer Variants
Anuj Diwan, Eunsol Choi, David Harwath
We present the first unified study of the efficiency of self-attention-based Transformer variants spanning text, speech and vision. We identify input length thresholds (tipping poi…
Why is Winoground Hard? Investigating Failures in Visuolinguistic Compositionality
Anuj Diwan, Layne Berry, Eunsol Choi +2
Recent visuolinguistic pre-trained models show promising progress on various end tasks such as image retrieval and video captioning. Yet, they fail miserably on the recently propos…
Zero-shot Video Moment Retrieval With Off-the-Shelf Models
Anuj Diwan, Puyuan Peng, Raymond J. Mooney
For the majority of the machine learning community, the expensive nature of collecting high-quality human-annotated data and the inability to efficiently finetune very large state-…
Multilingual and code-switching ASR challenges for low resource Indian languages
Anuj Diwan, Rakesh Vaideeswaran, Sanket Shah +19
Recently, there is increasing interest in multilingual automatic speech recognition (ASR) where a speech recognition system caters to multiple low resource languages by taking adva…
Reduce and Reconstruct: ASR for Low-Resource Phonetic Languages
Anuj Diwan, Preethi Jyothi
This work presents a seemingly simple but effective technique to improve low-resource ASR systems for phonetic languages. By identifying sets of acoustically similar graphemes in t…