7 papers
Hierarchical Domain Generalization
Chenxiao Yang, Zhiyuan Li, Shai Ben-David +1
We study hierarchical domain generalization as a problem of extrapolation from finite observed regions to an entire instance space, replacing i.i.d. sampling with arbitrary domain…
Overfitting and Generalizing with (PAC) Bayesian Prediction in Noisy Binary Classification
Xiaohan Zhu, Mesrob I. Ohannessian, Nathan Srebro
We consider a PAC-Bayes type learning rule for binary classification, balancing the training error of a randomized ''posterior'' predictor with its KL divergence to a pre-specified…
Weak-to-Strong Generalization Even in Random Feature Networks, Provably
Marko Medvedev, Kaifeng Lyu, Dingli Yu +3
Weak-to-Strong Generalization (Burns et al., 2024) is the phenomenon whereby a strong student, say GPT-4, learns a task from a weak teacher, say GPT-2, and ends up significantly ou…
Shift is Good: Mismatched Data Mixing Improves Test Performance
Marko Medvedev, Kaifeng Lyu, Zhiyuan Li +1
We consider training and testing on mixture distributions with different training and test proportions. We show that in many settings, and in some sense generically, distribution s…
PENCIL: Long Thoughts with Short Memory
Chenxiao Yang, Nathan Srebro, David McAllester +1
While state-of-the-art LLMs have demonstrated great promise of using long Chains-of-Thought (CoT) to boost reasoning, scaling it up to more challenging problems at test-time is fun…
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
Xiaohan Zhu, Nathan Srebro
We provide a complete characterization of the entire regularization curve of a modified two-part-code Minimum Description Length (MDL) learning rule for binary classification, base…