Approximate label symmetries improve data efficiency
arXiv:2605.28238
Abstract
Enforcing feature symmetries in machine learning (ML) models is a common strategy to mitigate data scarcity. Confirming expectations from statistical learning theory, we show that exact, as well as approximate, label symmetries can also improve data efficiency. We illustrate the idea for the s, p, d orbital densities of the electron in the hydrogen atom, for the three vibrational normal modes of the water molecule, and for its full 3D potential energy hypersurface. Resulting ML models of electron density and potential energies exhibit superior learning curves, demonstrating improved generalization efficiency. We observe that learning curves similarly improve even when label symmetries are not exact - up to the convergence floors set by the degree to which the symmetry is approximate. Further improvements are obtained for approximate label symmetries in the molecular potential energy surface, using a Hessian-based correction that suppresses the leading order term in the error.