Unsupervised Post-Training of Foundation Models: A Survey
arXiv:2608.24982
Abstract
Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rather than an external oracle. We catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated target, or an internal evaluator. Beyond inventory, we show how the choice of internal signal and task structure determines whether post-training improves the model or recursively amplifies error. An orthogonal Input Visibility Update Persistence view maps deployment regimes and defines a unified framework for UPT selection and evaluation.
Accepted to Findings of EMNLP 2026. 20 pages, 3 figures, 8 tables