Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding (Survey)
arXiv:2307.10246
Abstract
Can artificial intelligence unlock the secrets of the human brain? How do the inner mechanisms of deep learning models relate to our neural circuits? Is it possible to enhance AI by tapping into the power of brain recordings? These captivating questions lie at the heart of an emerging field at the intersection of neuroscience and artificial intelligence. Our survey dives into this exciting domain, focusing on human brain recording studies and cutting-edge cognitive neuroscience datasets that capture brain activity during natural language processing, visual perception, and auditory experiences. We explore two fundamental approaches: encoding models, which attempt to generate brain activity patterns from sensory inputs; and decoding models, which aim to reconstruct our thoughts and perceptions from neural signals. These techniques not only promise breakthroughs in neurological diagnostics and brain-computer interfaces but also offer a window into the very nature of cognition. In this survey, we first discuss popular representations of language, vision, and speech stimuli, and present a summary of neuroscience datasets. We then review how the recent advances in deep learning transformed this field, by investigating the popular deep learning based encoding and decoding architectures, noting their benefits and limitations across different sensory modalities. From text to images, speech to videos, we investigate how these models capture the brain's response to our complex, multimodal world. While our primary focus is on human studies, we also highlight the crucial role of animal models in advancing our understanding of neural mechanisms. Throughout, we mention the ethical implications of these powerful technologies, addressing concerns about privacy and cognitive liberty. We conclude with a summary and discussion of future trends in this rapidly evolving field.
61 pages, 22 figures
References in corpus (47)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Efficient Estimation of Word Representations in Vector Space
- Auto-Encoding Variational Bayes
- Denoising Diffusion Probabilistic Models
- Longformer: The Long-Document Transformer
- An efficient framework for learning sentence representations
- Make-A-Video: Text-to-Video Generation without Text-Video Data
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
- Decoding speech perception from non-invasive brain recordings
- MusicLM: Generating Music From Text
- Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain)
- From voxels to pixels and back: Self-supervision in natural-image reconstruction from fMRI
- Toward a realistic model of speech processing in the brain with self-supervised learning
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors
- Mind Reader: Reconstructing complex images from brain activities
- Disentangling Syntax and Semantics in the Brain with Deep Networks
- MuLan: A Joint Embedding of Music Audio and Natural Language
- Scaling laws for language encoding models in fMRI
- Brain decoding: toward real-time reconstruction of visual perception
- Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?
- Inducing brain-relevant bias in natural language processing models
- Cinematic Mindscapes: High-quality Video Reconstruction from Brain Activity
- Self-supervised models of audio effectively explain human cortical responses to speech
- Does injecting linguistic structure into language models lead to better alignment with brain recordings?
- Brain encoding models based on multimodal transformers can transfer across language and vision
- Brain2Word: Decoding Brain Activity for Language Generation
- Language models and brains align due to more than next-word prediction and word-level information
- Joint processing of linguistic properties in brains and language models
- Training language models to summarize narratives improves brain alignment
- Multimodal foundation models are better simulators of the human brain
- Contrast, Attend and Diffuse to Decode High-Resolution Images from Brain Activities
- A Penny for Your (visual) Thoughts: Self-Supervised Reconstruction of Natural Movies from Brain Activity
- Modeling Task Effects on Meaning Representation in the Brain via Zero-Shot MEG Prediction
- Visio-Linguistic Brain Encoding
- Decoding Natural Images from EEG for Object Recognition
- Beyond linear regression: mapping models in cognitive neuroscience should align with research goals
- BrainCLIP: Bridging Brain and Visual-Linguistic Representation Via CLIP for Generic Natural Visual Stimulus Decoding
- Improving visual image reconstruction from human brain activity using latent diffusion models via multiple decoded inputs
- Explaining black box text modules in natural language with language models
- A Dual-Stream Neural Network Explains the Functional Segregation of Dorsal and Ventral Visual Pathways in Human Brains
- Same Cause; Different Effects in the Brain
- Instruction-tuning Aligns LLMs to the Human Brain
- Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI
- Applicability of scaling laws to vision encoding models
- Brain-Like Language Processing via a Shallow Untrained Multihead Attention Network
- Functional Brain-to-Brain Transformation with No Shared Data
- MinD-3D: Reconstruct High-quality 3D objects in Human Brain