5 papers
Provably Learning Attention with Queries
Satwik Bhattamishra, Kulin Shah, Michael Hahn +1
We study the problem of learning Transformer-based sequence models with black-box access to their outputs. In this setting, a learner may adaptively query the oracle with any seque…
How Global Calibration Strengthens Multiaccuracy
SÃlvia Casacuberta, Parikshit Gopalan, Varun Kanade +1
Multiaccuracy and multicalibration are multigroup fairness notions for prediction that have found numerous applications in learning and computational complexity. They can be achiev…
DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection
Tala Aljaafari, Varun Kanade, Philip Torr +1
Deploying reinforcement learning (RL) in safety-critical settings is constrained by brittleness under distribution shift. We study out-of-distribution (OOD) detection for RL time s…
Hardness of Learning Regular Languages in the Next Symbol Prediction Setting
Satwik Bhattamishra, Phil Blunsom, Varun Kanade
We study the learnability of languages in the Next Symbol Prediction (NSP) setting, where a learner receives only positive examples from a language together with, for every prefix,…
Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers
Charles London, Varun Kanade
Pause tokens, simple filler symbols such as "...", consistently improve Transformer performance on both language and mathematical tasks, yet their theoretical effect remains unexpl…