most citedWhat can a Single Attention Layer Learn? A Study Through the Random Features Lens

3 citations · 6 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG20233 cited

What can a Single Attention Layer Learn? A Study Through the Random Features Lens

Hengyu Fu, Tianyu Guo, Yu Bai +1

Attention layers -- which map a sequence of inputs to a sequence of outputs -- are core building blocks of the Transformer architecture which has achieved significant breakthroughs…

cs.LG2023

Sample-Efficient Learning of POMDPs with Multiple Observations In Hindsight

Jiacheng Guo, Minshuo Chen, Huan Wang +3

This paper studies the sample-efficiency of learning in Partially Observable Markov Decision Processes (POMDPs), a challenging problem in reinforcement learning that is known to be…

cs.LG2023

Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection

Yu Bai, Fan Chen, Huan Wang +2

Neural sequence models based on the transformer architecture have demonstrated remarkable \emph{in-context learning} (ICL) abilities, where they can perform new tasks when prompted…

cs.LG20221 cited

The Role of Coverage in Online Reinforcement Learning

Tengyang Xie, Dylan J. Foster, Yu Bai +2

Coverage conditions -- which assert that the data logging distribution adequately covers the state space -- play a fundamental role in determining the sample complexity of offline…

cs.LG20221 cited

Sample-Efficient Learning of Correlated Equilibria in Extensive-Form Games

Ziang Song, Song Mei, Yu Bai

Imperfect-Information Extensive-Form Games (IIEFGs) is a prevalent model for real-world games involving imperfect information and sequential plays. The Extensive-Form Correlated Eq…

cs.LG20221 cited

Efficient and Differentiable Conformal Prediction with General Function Classes

Yu Bai, Song Mei, Huan Wang +2

Quantifying the data uncertainty in learning tasks is often done by learning a prediction interval or prediction set of the label given the input. Two commonly desired properties f…