What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure
arXiv:2302.12239 · doi:10.1038/s41467-024-55158-1
Abstract
Deep neural networks drive the success of natural language processing. A fundamental property of language is its compositional structure, allowing humans to systematically produce forms for new meanings. For humans, languages with more compositional and transparent structures are typically easier to learn than those with opaque and irregular structures. However, this learnability advantage has not yet been shown for deep neural networks, limiting their use as models for human language learning. Here, we directly test how neural networks compare to humans in learning and generalizing different languages that vary in their degree of compositional structure. We evaluate the memorization and generalization capabilities of a large language model and recurrent neural networks, and show that both deep neural networks exhibit a learnability advantage for more structured linguistic input: neural networks exposed to more compositional languages show more systematic generalization, greater agreement between different agents, and greater similarity to human learners.
20 pages + supplementary material
References in corpus (12)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Training language models to follow instructions with human feedback
- On the Opportunities and Risks of Foundation Models
- Reconciling modern machine learning practice and the bias-variance trade-off
- Scaling Laws for Neural Language Models
- Emergent Abilities of Large Language Models
- Quantifying Memorization Across Neural Language Models
- Linguistic generalization and compositionality in modern artificial neural networks
- State-of-the-art generalisation research in NLP: A taxonomy and review
- Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
- To Code, or Not To Code? Exploring Impact of Code in Pre-training
- Emergent Communication: Generalization and Overfitting in Lewis Games