Cognitive Science in the era of Artificial Intelligence: A roadmap for reverse-engineering the infant language-learner
arXiv:1607.08723 · doi:10.1016/j.cognition.2017.11.008
Abstract
During their first years of life, infants learn the language(s) of their environment at an amazing speed despite large cross cultural variations in amount and complexity of the available language input. Understanding this simple fact still escapes current cognitive and linguistic theories. Recently, spectacular progress in the engineering science, notably, machine learning and wearable technology, offer the promise of revolutionizing the study of cognitive development. Machine learning offers powerful learning algorithms that can achieve human-like performance on many linguistic tasks. Wearable sensors can capture vast amounts of data, which enable the reconstruction of the sensory experience of infants in their natural environment. The project of 'reverse engineering' language development, i.e., of building an effective system that mimics infant's achievements appears therefore to be within reach. Here, we analyze the conditions under which such a project can contribute to our scientific understanding of early language development. We argue that instead of defining a sub-problem or simplifying the data, computational models should address the full complexity of the learning situation, and take as input the raw sensory signals available to infants. This implies that (1) accessible but privacy-preserving repositories of home data be setup and widely shared, and (2) models be evaluated at different linguistic levels through a benchmark of psycholinguist tests that can be passed by machines and humans alike, (3) linguistically and psychologically plausible learning architectures be scaled up to real data using probabilistic/optimization principles from machine learning. We discuss the feasibility of this approach and present preliminary results.
27 pages, 5 figures, 3 tables, supplementary materials
References in corpus (9)
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- WaveNet: A Generative Model for Raw Audio
- From Frequency to Meaning: Vector Space Models of Semantics
- Deep Neural Networks Rival the Representation of Primate IT Cortex for Core Visual Object Recognition
- Achieving Human Parity in Conversational Speech Recognition
- Exploring Nearest Neighbor Approaches for Image Captioning
- Surpassing Human-Level Face Verification Performance on LFW with GaussianFace
- Temporally coherent 4D reconstruction of complex dynamic scenes
- Are words easier to learn from infant- than adult-directed speech? A quantitative corpus-based investigation
Cited by in corpus (21)
- On the Opportunities and Risks of Foundation Models
- Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora
- Generative Adversarial Phonology: Modeling unsupervised phonetic and phonological learning with neural networks
- Self-supervised language learning from raw audio: Lessons from the Zero Resource Speech Challenge
- Speech-Image Semantic Alignment Does Not Depend on Any Prior Classification Tasks
- improving partition-block-based acoustic echo canceler in under-modeling scenarios
- Unsupervised Word Segmentation from Speech with Attention
- What they do when in doubt: a study of inductive biases in seq2seq learners
- Unsupervised Few-shot Learning via Self-supervised Training
- Development of collective behavior in newborn artificial agents
- BabySLM: language-acquisition-friendly benchmark of self-supervised spoken language models
- A Computational Model of Early Word Learning from the Infant's Point of View
- Learning Hierarchical Discrete Linguistic Units from Visually-Grounded Speech
- Improving Unsupervised Subword Modeling via Disentangled Speech Representation Learning and Transformation
- The Perceptimatic English Benchmark for Speech Perception Models
- A Hierarchical Subspace Model for Language-Attuned Acoustic Unit Discovery
- Learning Neural Models for Natural Language Processing in the Face of Distributional Shift
- How Familiar Does That Sound? Cross-Lingual Representational Similarity Analysis of Acoustic Word Embeddings
- Unsupervised Word Segmentation from Discrete Speech Units in Low-Resource Settings
- Voice Conversion Based Speaker Normalization for Acoustic Unit Discovery
- Competition in Cross-situational Word Learning: A Computational Study