DREAM: Visual Decoding from Reversing Human Visual System
arXiv:2310.02265
Abstract
In this work we present DREAM, an fMRI-to-image method for reconstructing viewed images from brain activities, grounded on fundamental knowledge of the human visual system. We craft reverse pathways that emulate the hierarchical and parallel nature of how humans perceive the visual world. These tailored pathways are specialized to decipher semantics, color, and depth cues from fMRI data, mirroring the forward pathways from visual stimuli to fMRI recordings. To do so, two components mimic the inverse processes within the human visual system: the Reverse Visual Association Cortex (R-VAC) which reverses pathways of this brain region, extracting semantics from fMRI data; the Reverse Parallel PKM (R-PKM) component simultaneously predicting color and depth from fMRI signals. The experiments indicate that our method outperforms the current state-of-the-art models in terms of the consistency of appearance, structure, and semantics. Code will be made publicly available to facilitate further research in this field.
Project Page: https://weihaox.github.io/DREAM
References in corpus (13)
- LAION-5B: An open large-scale dataset for training next generation image-text models
- Adding Conditional Control to Text-to-Image Diffusion Models
- MixCo: Mix-up Contrastive Learning for Visual Representation
- Reconstructing the Mind's Eye: fMRI-to-Image with Contrastive Learning and Diffusion Priors
- Mind Reader: Reconstructing complex images from brain activities
- T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
- Decoding natural image stimuli from fMRI data with a surface-based convolutional network
- Brain Captioning: Decoding human brain activity into images and text
- Versatile Diffusion: Text, Images and Variations All in One Diffusion Model
- Improving visual image reconstruction from human brain activity using latent diffusion models via multiple decoded inputs
- BrainCLIP: Bridging Brain and Visual-Linguistic Representation Via CLIP for Generic Natural Visual Stimulus Decoding
- Natural scene reconstruction from fMRI signals using generative latent diffusion
- Controllable Mind Visual Diffusion Model