3 papers
cs.LG2025
On Weak-to-Strong Generalization and f-Divergence
Wei Yao, Gengze Xu, Huayi Tang +4
Weak-to-strong generalization (W2SG) has emerged as a promising paradigm for stimulating the capabilities of strong pre-trained models by leveraging supervision from weaker supervi…
cs.LG2025
Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
Wei Yao, Wenkai Yang, Ziqiao Wang +2
As large language models advance toward superhuman performance, ensuring their alignment with human values and abilities grows increasingly complex. Weak-to-strong generalization o…
cs.LG2025
The Capabilities and Limitations of Weak-to-Strong Generalization: Generalization and Calibration
Wei Yao, Wenkai Yang, Gengze Xu +3
Weak-to-strong generalization, where weakly supervised strong models outperform their weaker teachers, offers a promising approach to aligning superhuman models with human values.…