23 citations · 31 across the 10 of their papers we have counts for
9 papers
Learning Neural Networks with Sparse Activations
Pranjal Awasthi, Nishanth Dikkala, Pritish Kamath +1
A core component present in many successful neural network architectures, is an MLP block of two fully connected layers with a non-linear activation in between. An intriguing pheno…
ReMI: A Dataset for Reasoning with Multiple Images
Mehran Kazemi, Nishanth Dikkala, Ankit Anand +8
With the continuous advancement of large language models (LLMs), it is essential to create new benchmarks to effectively evaluate their expanding capabilities and identify areas fo…
Improving Length-Generalization in Transformers via Task Hinting
Pranjal Awasthi, Anupam Gupta
It has been observed in recent years that transformers have problems with length generalization for certain types of reasoning and arithmetic tasks. In particular, the performance…
The Sample Complexity of Multi-Distribution Learning for VC Classes
Pranjal Awasthi, Nika Haghtalab, Eric Zhao
Multi-distribution learning is a natural generalization of PAC learning to settings with multiple data distributions. There remains a significant gap between the known upper and lo…
Best-Effort Adaptation
Pranjal Awasthi, Corinna Cortes, Mehryar Mohri
We study a problem of best-effort adaptation motivated by several applications and considerations, which consists of determining an accurate predictor for a target domain, for whic…
Congested Bandits: Optimal Routing via Short-term Resets
Pranjal Awasthi, Kush Bhatia, Sreenivas Gollapudi +1
For traffic routing platforms, the choice of which route to recommend to a user depends on the congestion on these routes -- indeed, an individual's utility depends on the number o…