462 citations · 767 across the 20 of their papers we have counts for
11 papers · 1 filter
Higher-order Linear Attention
Yifan Zhang, Zhen Qin, Mengdi Wang +1
The quadratic cost of scaled dot-product attention is a central obstacle to scaling autoregressive language models to long contexts. Linear-time attention and State Space Models (S…
Variance-Dependent Regret Bounds for Linear Bandits and Reinforcement Learning: Adaptivity and Computational Efficiency
Heyang Zhao, Jiafan He, Dongruo Zhou +2
Recently, several studies (Zhou et al., 2021a; Zhang et al., 2021b; Kim et al., 2021; Zhou and Gu, 2022) have provided variance-dependent regret bounds for linear contextual bandit…
Structure-informed Language Models Are Protein Designers
Zaixiang Zheng, Yifan Deng, Dongyu Xue +3
This paper demonstrates that language models are strong structure-based protein designers. We present LM-Design, a generic approach to reprogramming sequence-based protein language…
Learning Two-Player Mixture Markov Games: Kernel Function Approximation and Correlated Equilibrium
Chris Junchi Li, Dongruo Zhou, Quanquan Gu +1
We consider learning Nash equilibria in two-player zero-sum Markov Games with nonlinear function approximation, where the action-value function is approximated by a function in a R…
Towards Understanding Mixture of Experts in Deep Learning
Zixiang Chen, Yihe Deng, Yue Wu +2
The Mixture-of-Experts (MoE) layer, a sparsely-activated model controlled by a router, has achieved great success in deep learning. However, the understanding of such architecture…
The Power and Limitation of Pretraining-Finetuning for Linear Regression under Covariate Shift
Jingfeng Wu, Difan Zou, Vladimir Braverman +2
We study linear regression under covariate shift, where the marginal distribution over the input covariates differs in the source and the target domains, while the conditional dist…