Coding schemes in neural networks learning classification tasks
arXiv:2406.16689 · doi:10.1038/s41467-025-58276-6
Abstract
Neural networks posses the crucial ability to generate meaningful representations of task-dependent features. Indeed, with appropriate scaling, supervised learning in neural networks can result in strong, task-dependent feature learning. However, the nature of the emergent representations, which we call the `coding scheme', is still unclear. To understand the emergent coding scheme, we investigate fully-connected, wide neural networks learning classification tasks using the Bayesian framework where learning shapes the posterior distribution of the network weights. Consistent with previous findings, our analysis of the feature learning regime (also known as `non-lazy', `rich', or `mean-field' regime) shows that the networks acquire strong, data-dependent features. Surprisingly, the nature of the internal representations depends crucially on the neuronal nonlinearity. In linear networks, an analog coding scheme of the task emerges. Despite the strong representations, the mean predictor is identical to the lazy case. In nonlinear networks, spontaneous symmetry breaking leads to either redundant or sparse coding schemes. Our findings highlight how network properties such as scaling of weights and neuronal nonlinearity can profoundly influence the emergent representations.
References in corpus (8)
- Reconciling modern machine learning practice and the bias-variance trade-off
- A Mean Field View of the Landscape of Two-Layers Neural Networks
- Prevalence of Neural Collapse during the terminal phase of deep learning training
- The Principles of Deep Learning Theory
- Disentangling feature and lazy training in deep neural networks
- A statistical mechanics framework for Bayesian deep neural networks beyond the infinite-width limit
- Unified field theoretical approach to deep and recurrent neuronal networks
- Minimal model of permutation symmetry in unsupervised learning
Cited by in corpus (4)
- Summary statistics of learning link changing neural representations to behavior
- Statistical physics of deep learning: Optimal learning of a multi-layer perceptron near interpolation
- Statistical mechanics of extensive-width Bayesian neural networks near interpolation
- Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation