collaborators

6 papers

cs.LG2026

When Does -Boosting Overfit Benignly? High-Dimensional Risk Asymptotics and the Implicit Bias

Ye Su, Jian Li, Yong Liu

Benign overfitting is well-characterized in geometries, but its behavior under the implicit bias of greedy ensembles remains challenging. The analytical barrier s…

cs.LG2026

On the Blessing of Pre-training in Weak-to-Strong Generalization

Wei Yao, Wang Zhaoyang, Gengze Xu +5

The paradigm of Weak-to-Strong Generalization (W2SG) suggests that a pre-trained strong model can surpass its weak supervisor, yet the decisive role of pre-training remains theoret…

cs.LG2025

Weak-to-Strong Generalization via Bregman Bias-Variance Decomposition

Gengze Xu, Wei Yao, Ziqiao Wang +1

Weak-to-strong generalization (W2SG) is the phenomenon in which a powerful student model, trained on labels produced by a weaker teacher, ultimately outperforms the teacher on the…

cs.LG2025

The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration

Wei Yao, Wenkai Yang, Gengze Xu +3

Weak-to-strong generalization, where weakly supervised strong models outperform their weaker teachers, offers a promising approach to aligning superhuman models with human values.…

cs.LG2025

On Weak-to-Strong Generalization and f-Divergence

Wei Yao, Gengze Xu, Huayi Tang +4

Weak-to-strong generalization (W2SG) has emerged as a promising paradigm for stimulating the capabilities of strong pre-trained models by leveraging supervision from weaker supervi…

cs.LG2025

Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL

Wei Yao, Wenkai Yang, Ziqiao Wang +2

As large language models advance toward superhuman performance, ensuring their alignment with human values and abilities grows increasingly complex. Weak-to-strong generalization o…