activity
20172023
most citedFine-Grained Analysis of Optimization and Generalization for Overparameterized Two-Layer Neural Networks

256 citations · 627 across the 30 of their papers we have counts for

collaborators
Showing 2023Show all

9 papers · 1 filter

cs.LG2023

Settling the Sample Complexity of Online Reinforcement Learning

Zihan Zhang, Yuxin Chen, Jason D. Lee +1

A central issue lying at the heart of online reinforcement learning (RL) is data efficiency. While a number of recent works achieved asymptotically minimal regret in online RL, the…

cs.LG2023

Active Representation Learning for General Task Space with Applications in Robotics

Yifang Chen, Yingbing Huang, Simon S. Du +2

Representation learning based on multi-task pretraining has become a powerful approach in many domains. In particular, task-aware representation learning aims to learn an optimal r…

cs.LG2023

Improved Active Multi-Task Representation Learning via Lasso

Yiping Wang, Yifang Chen, Kevin Jamieson +1

To leverage the copious amount of data from source tasks and overcome the scarcity of the target task samples, representation learning based on multi-task pretraining has become a…

cs.LG2023

A Black-box Approach for Non-stationary Multi-agent Reinforcement Learning

Haozhe Jiang, Qiwen Cui, Zhihan Xiong +2

We investigate learning the equilibria in non-stationary multi-agent systems and address the challenges that differentiate multi-agent learning from single-agent learning. Specific…

cs.CL2023

Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer

Yuandong Tian, Yiping Wang, Beidi Chen +1

Transformer architecture has shown impressive performance in multiple research domains and has become the backbone of many neural network models. However, there is limited understa…

cs.LG2023

Over-Parameterization Exponentially Slows Down Gradient Descent for Learning a Single Neuron

Weihang Xu, Simon S. Du

We revisit the problem of learning a single neuron with ReLU activation under Gaussian input with square loss. We particularly focus on the over-parameterization setting where the…