6 papers
Near-Optimal Lower Bounds for Randomized Algorithms in Exact Value Zeroth-Order Convex Optimization
Haihan Zhang, Chenheng Zhang, Zhiquan Qi +1
Whether exact scalar feedback intrinsically incurs the additional dimension paid by known zeroth-order methods remains open even for Lipschitz convex optimization. For a univer…
What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear Regression
Wendao Wu, Fangqing Zhang, Haihan Zhang +1
Teacher-Student Knowledge Transfer (KT) is ubiquitous in modern machine learning, ranging from classical model compression via Knowledge Distillation (KD) to the emergent phenomeno…
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
Wenjie Ma, Andrei Cojocaru, Neel Kolhe +6
Recent advances in large language models (LLMs) for mathematical reasoning have largely focused on tasks with easily verifiable final answers while generating and verifying natural…
Learning Curves of Stochastic Gradient Descent in Kernel Regression
Haihan Zhang, Weicheng Lin, Yuanshi Liu +1
This paper considers a canonical problem in kernel regression: how good are the model performances when it is trained by the popular online first-order algorithms, compared to the…
Scaling Law for Stochastic Gradient Descent in Quadratically Parameterized Linear Regression
Shihong Ding, Haihan Zhang, Hanzhen Zhao +1
In machine learning, the scaling law describes how the model performance improves with the model and data size scaling up. From a learning theory perspective, this class of results…
Optimal Algorithms in Linear Regression under Covariate Shift: On the Importance of Precondition
Yuanshi Liu, Haihan Zhang, Qian Chen +1
A common pursuit in modern statistical learning is to attain satisfactory generalization out of the source data distribution (OOD). In theory, the challenge remains unsolved even u…