activity
20232025
most citedImplicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal Data

4 citations · 8 across the 3 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2025

Smoothed Agnostic Learning of Halfspaces over the Hypercube

Yiwen Kou, Raghu Meka

Agnostic learning of Boolean halfspaces is a fundamental problem in computational learning theory, but it is known to be computationally hard even for weak learning. Recent work [C…

cs.LG2024

Matching the Statistical Query Lower Bound for -Sparse Parity Problems with Sign Stochastic Gradient Descent

Yiwen Kou, Zixiang Chen, Quanquan Gu +1

The -sparse parity problem is a classical problem in computational complexity and algorithmic theory, serving as a key benchmark for understanding computational classes. In this…

cs.LG2024

Guided Discrete Diffusion for Electronic Health Record Generation

Jun Han, Zixiang Chen, Yongqian Li +4

Electronic health records (EHRs) are a pivotal data source that enables numerous applications in computational medicine, e.g., disease progression prediction, clinical trial design…

cs.LG2023

Fast Sampling via Discrete Non-Markov Diffusion Models with Predetermined Transition Time

Zixiang Chen, Huizhuo Yuan, Yongqian Li +3

Discrete diffusion models have emerged as powerful tools for high-quality data generation. Despite their success in discrete spaces, such as text generation tasks, the acceleration…

cs.LG20234 cited

Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal Data

Yiwen Kou, Zixiang Chen, Quanquan Gu

The implicit bias towards solutions with favorable properties is believed to be a key reason why neural networks trained by gradient-based optimization can generalize well. While t…

cs.LG20234 cited

Why Does Sharpness-Aware Minimization Generalize Better Than SGD?

Zixiang Chen, Junkai Zhang, Yiwen Kou +3

The challenge of overfitting, in which the model memorizes the training data and fails to generalize to test data, has become increasingly significant in the training of large neur…