collaborators

5 papers

cs.LG2026

Provably Learning Attention with Queries

Satwik Bhattamishra, Kulin Shah, Michael Hahn +1

We study the problem of learning Transformer-based sequence models with black-box access to their outputs. In this setting, a learner may adaptively query the oracle with any seque…

cs.LG2026

How Global Calibration Strengthens Multiaccuracy

Sílvia Casacuberta, Parikshit Gopalan, Varun Kanade +1

Multiaccuracy and multicalibration are multigroup fairness notions for prediction that have found numerous applications in learning and computational complexity. They can be achiev…

cs.LG2025

DEEDEE: Fast and Scalable Out-of-Distribution Dynamics Detection

Tala Aljaafari, Varun Kanade, Philip Torr +1

Deploying reinforcement learning (RL) in safety-critical settings is constrained by brittleness under distribution shift. We study out-of-distribution (OOD) detection for RL time s…

cs.LG2025

Hardness of Learning Regular Languages in the Next Symbol Prediction Setting

Satwik Bhattamishra, Phil Blunsom, Varun Kanade

We study the learnability of languages in the Next Symbol Prediction (NSP) setting, where a learner receives only positive examples from a language together with, for every prefix,…

cs.LG2025

Pause Tokens Strictly Increase the Expressivity of Constant-Depth Transformers

Charles London, Varun Kanade

Pause tokens, simple filler symbols such as "...", consistently improve Transformer performance on both language and mathematical tasks, yet their theoretical effect remains unexpl…