Towards artificial general intelligence via a multimodal foundation model
arXiv:2110.14378 · doi:10.1038/s41467-022-30761-2
Abstract
The fundamental goal of artificial intelligence (AI) is to mimic the core cognitive activities of human. Despite tremendous success in the AI research, most of existing methods have only single-cognitive ability. To overcome this limitation and take a solid step towards artificial general intelligence (AGI), we develop a foundation model pre-trained with huge multimodal data, which can be quickly adapted for various downstream cognitive tasks. To achieve this goal, we propose to pre-train our foundation model by self-supervised learning with weak semantic correlation data crawled from the Internet and show that promising results can be obtained on a wide range of downstream tasks. Particularly, with the developed model-interpretability tools, we demonstrate that strong imagination ability is now possessed by our foundation model. We believe that our work makes a transformative stride towards AGI, from our common practice of "weak or narrow AI" to that of "strong or generalized AI".
Published by Nature Communications, see https://www.nature.com/articles/s41467-022-30761-2
References in corpus (3)
Cited by in corpus (13)
- AutoTRIZ: Automating Engineering Innovation with TRIZ and Large Language Models
- CLIP4STR: A Simple Baseline for Scene Text Recognition with Pre-trained Vision-Language Model
- Optical Generative Models
- A Survey on State-of-the-art Deep Learning Applications and Challenges
- From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
- Graph Foundation Models: Concepts, Opportunities and Challenges
- An overview of domain-specific foundation model: key technologies, applications and challenges
- Language models in molecular discovery
- Potentials of the Metaverse for Robotized Applications in Industry 4.0 and Industry 5.0
- Efficient and Effective Adaptation of Multimodal Foundation Models in Sequential Recommendation
- CCMB: A Large-scale Chinese Cross-modal Benchmark
- Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning
- Tissue Concepts: supervised foundation models in computational pathology