4 papers
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
Chenruo Liu, Kenan Tang, Yao Qin +1
This paper bridges distribution shift and AI safety through a comprehensive analysis of their conceptual and methodological synergies. While prior discussions often focus on narrow…
A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula
Chenruo Liu, Yijun Dong, Yiqiu Shen +1
Iterative self-improvement fine-tunes an autoregressive large language model (LLM) on reward-verified outputs generated by the LLM itself. In contrast to the empirical success of s…
Does Weak-to-strong Generalization Happen under Spurious Correlations?
Chenruo Liu, Yijun Dong, Qi Lei
We initiate a unified theoretical and algorithmic study of a key problem in weak-to-strong (W2S) generalization: when fine-tuning a strong pre-trained student with pseudolabels fro…
Superclass-Guided Representation Disentanglement for Spurious Correlation Mitigation
Chenruo Liu, Hongjun Liu, Zeyu Lai +3
To enhance group robustness to spurious correlations, prior work often relies on auxiliary group annotations and assumes identical sets of groups across training and test domains.…