6 papers
Reducing Learner Redundancy in Boosting via Residual Orthogonalization
Ye Su, Jipeng Guo, Yong Liu +5
While sequential residual fitting is the bedrock of standard boosting frameworks, it inherently breeds learner redundancy by repeatedly revisiting correlated error components. To a…
Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry
Ye Su, Huayi Tang, Zixuan Gong +1
While Mixture-of-Experts (MoE) architectures define the state-of-the-art, their theoretical success is often attributed to heuristic efficiency rather than geometric expressivity.…
Geometric Capacity of Transformers: A Tropical Geometry Perspective
Ye Su, Yong Liu
To quantify the geometric capacity of transformers, we develop a tropical-geometric framework for analyzing the spatial partitions induced by conditioned self-attention. In the zer…
Exact Finite-Sample Variance Decomposition of Subagging: A Spectral Filtering Perspective
Ye Su, Mingrui Ye, Yining Wang +2
Standard resampling ratios (e.g., ) are widely used as default baselines in ensemble learning for three decades. However, how these ratios interact with a base lea…
Effective Frontiers: A Unification of Neural Scaling Laws
Jiaxuan Zou, Zixuan Gong, Ye Su +2
Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity (), datasize (), and compute (). However, existing theoretical…
Variational Inference, Entropy, and Orthogonality: A Unified Theory of Mixture-of-Experts
Ye Su, Yong Liu
Mixture-of-Experts models enable large language models to scale efficiently, as they only activate a subset of experts for each input. Their core mechanisms, Top-k routing and auxi…