collaborators

5 papers

cs.LG2026

Sharper Analysis of Single-Loop Methods for Bilevel Optimization

Yubo Zhou, Jun Shu, Luo Luo +4

Bilevel optimization underpins many machine learning applications, including hyperparameter optimization, meta-learning, neural architecture search, and reinforcement learning. Whi…

cs.LG2026

Leveraging Extragradient for Effective Sharpness-Aware Minimization in Deep Learning

Yao Fu, Chunxia Zhang, Junmin Liu +3

Generalization remains a pivotal challenge in deep learning, where traditional optimizers like Stochastic Gradient Descent (SGD) often converge to sharp minima, leading to overfitt…

cs.LG2026

Towards Understanding The Calibration Benefits of Sharpness-Aware Minimization

Chengli Tan, Yubo Zhou, Haishan Ye +7

Deep neural networks have been increasingly used in safety-critical applications such as medical diagnosis and autonomous driving. However, many studies suggest that they are prone…

cs.LG2026

Understanding the Generalization of Bilevel Programming in Hyperparameter Optimization: A Tale of Bias-Variance Decomposition

Yubo Zhou, Jun Shu, Junmin Liu +1

Gradient-based hyperparameter optimization (HPO) have emerged recently, leveraging bilevel programming techniques to optimize hyperparameter by estimating hypergradient w.r.t. vali…

math.ST2025

Zero-Order Sharpness-Aware Minimization

Yao Fu, Yihang Jin, Chunxia Zhang +3

Prompt learning has become a key method for adapting large language models to specific tasks with limited data. However, traditional gradient-based optimization methods for tuning…