papers

Publications (39)

cs.CL2022

Transformer Memory as a Differentiable Search Index

Yi Tay, Vinh Q. Tran, Mostafa Dehghani +10

In this paper, we demonstrate that information retrieval can be accomplished with a single Transformer, in which all information about the corpus is encoded in the parameters of th…

quant-ph2025

Experimental demonstration of the Bell-type inequalities for four qubit Dicke state using IBM Quantum Processing Units

Tomis Prajapati, Harsh Mehta, Shreya Banerjee +2

Violation of the Bell-type inequalities is necessary to confirm the existence of nonlocality in nonclassical (entangled) states. We have designed a customized operator which is mad…

cs.LG2024

Optimal Linear Decay Learning Rate Schedules and Further Refinements

Aaron Defazio, Ashok Cutkosky, Harsh Mehta +1

Learning rate schedules used in practice bear little resemblance to those recommended by theory. We close much of this theory/practice gap, and as a consequence are able to derive…

cs.LG2022

ALX: Large Scale Matrix Factorization on TPUs

Harsh Mehta, Steffen Rendle, Walid Krichene +1

We present ALX, an open-source library for distributed matrix factorization using Alternating Least Squares, written in JAX. Our design allows for efficient use of the TPU architec…

cs.CV2020

Retouchdown: Adding Touchdown to StreetLearn as a Shareable Resource for Language Grounding Tasks in Street View

Harsh Mehta, Yoav Artzi, Jason Baldridge +2

The Touchdown dataset (Chen et al., 2019) provides instructions by human annotators for navigation through New York City streets and for resolving spatial descriptions at a given l…

cs.LG2022

Large Scale Transfer Learning for Differentially Private Image Classification

Harsh Mehta, Abhradeep Thakurta, Alexey Kurakin +1

Differential Privacy (DP) provides a formal framework for training machine learning models with individual example level privacy. In the field of deep learning, Differentially Priv…