collaborators

6 papers

cs.LG2026

Reducing Learner Redundancy in Boosting via Residual Orthogonalization

Ye Su, Jipeng Guo, Yong Liu +5

While sequential residual fitting is the bedrock of standard boosting frameworks, it inherently breeds learner redundancy by repeatedly revisiting correlated error components. To a…

cs.LG2026

Sparsity is Combinatorial Depth: Quantifying MoE Expressivity via Tropical Geometry

Ye Su, Huayi Tang, Zixuan Gong +1

While Mixture-of-Experts (MoE) architectures define the state-of-the-art, their theoretical success is often attributed to heuristic efficiency rather than geometric expressivity.…

cs.LG2026

Geometric Capacity of Transformers: A Tropical Geometry Perspective

Ye Su, Yong Liu

To quantify the geometric capacity of transformers, we develop a tropical-geometric framework for analyzing the spatial partitions induced by conditioned self-attention. In the zer…

cs.LG2026

Exact Finite-Sample Variance Decomposition of Subagging: A Spectral Filtering Perspective

Ye Su, Mingrui Ye, Yining Wang +2

Standard resampling ratios (e.g., ) are widely used as default baselines in ensemble learning for three decades. However, how these ratios interact with a base lea…

cs.LG2026

Effective Frontiers: A Unification of Neural Scaling Laws

Jiaxuan Zou, Zixuan Gong, Ye Su +2

Neural scaling laws govern the prediction power-law improvement of test loss with respect to model capacity (), datasize (), and compute (). However, existing theoretical…

cs.LG2026

Variational Inference, Entropy, and Orthogonality: A Unified Theory of Mixture-of-Experts

Ye Su, Yong Liu

Mixture-of-Experts models enable large language models to scale efficiently, as they only activate a subset of experts for each input. Their core mechanisms, Top-k routing and auxi…