1 citations · 1 across the 7 of their papers we have counts for
6 papers · 1 filter
When Does -Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the Implicit Bias
Ye Su, Jian Li, Yong Liu
Benign overfitting is well-characterized in geometries, but its behavior under the implicit bias of greedy ensembles remains challenging. The analytical barrier s…
Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry
Ye Su, Huayi Tang, Zixuan Gong +1
While Mixture-of-Experts (MoE) architectures define the state-of-the-art, their theoretical success is often attributed to heuristic efficiency rather than geometric expressivity.…
Geometric Capacity of Transformers: A Tropical Geometry Perspective
Ye Su, Yong Liu
To quantify the geometric capacity of transformers, we develop a tropical-geometric framework for analyzing the spatial partitions induced by conditioned self-attention. In the zer…
Exact Finite-Sample Variance Decomposition of Subagging: A Spectral Filtering Perspective
Ye Su, Mingrui Ye, Yining Wang +2
Standard resampling ratios (e.g., ) are widely used as default baselines in ensemble learning for three decades. However, how these ratios interact with a base lea…
Effective Frontiers: A Unification of Neural Scaling Laws
Jiaxuan Zou, Zixuan Gong, Ye Su +2
Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity (), datasize (), and compute (). However, existing theoretical…
Variational Inference, Entropy, and Orthogonality: A Unified Theory of Mixture-of-Experts
Ye Su, Yong Liu
Mixture-of-Experts models enable large language models to scale efficiently, as they only activate a subset of experts for each input. Their core mechanisms, Top-k routing and auxi…