activity
20182023
most citedFine-Tuning Language Models with Just Forward Passes

36 citations · 72 across the 25 of their papers we have counts for

collaborators
Showing cs.LGShow all

27 papers · 1 filter

cs.LG2023★ 8 cited

Teaching Arithmetic to Small Transformers

Nayoung Lee, Kartik Sreenivasan, Jason D. Lee +2

Large language models like GPT-4 exhibit emergent capabilities across general-purpose tasks, such as basic arithmetic, when trained on extensive text data, even though these tasks…

cs.LG2023

Settling the Sample Complexity of Online Reinforcement Learning

Zihan Zhang, Yuxin Chen, Jason D. Lee +1

A central issue lying at the heart of online reinforcement learning (RL) is data efficiency. While a number of recent works achieved asymptotically minimal regret in online RL, the…

cs.LG2023

Sample Complexity for Quadratic Bandits: Hessian Dependent Bounds and Optimal Algorithms

Qian Yu, Yining Wang, Baihe Huang +2

In stochastic zeroth-order optimization, a problem of practical relevance is understanding how to fully exploit the local geometry of the underlying objective function. We consider…

cs.LG2023★ 2 cited

Smoothing the Landscape Boosts the Signal for SGD: Optimal Sample Complexity for Learning Single Index Models

Alex Damian, Eshaan Nichani, Rong Ge +1

We focus on the task of learning a single index model with respect to the isotropic Gaussian distribution in dimensions. Prior work has shown that the samp…

cs.LG2023★ 3 cited

Reward-agnostic Fine-tuning: Provable Statistical Benefits of Hybrid Reinforcement Learning

Gen Li, Wenhao Zhan, Jason D. Lee +2

This paper studies tabular reinforcement learning (RL) in the hybrid setting, which assumes access to both an offline dataset and online interactions with the unknown environment.…

cs.LG2023★ 2 cited

Local Optimization Achieves Global Optimality in Multi-Agent Reinforcement Learning

Yulai Zhao, Zhuoran Yang, Zhaoran Wang +1

Policy optimization methods with function approximation are widely used in multi-agent reinforcement learning. However, it remains elusive how to design such algorithms with statis…