Towards A Measure Of General Machine Intelligence
arXiv:2109.12075
Abstract
To build general-purpose artificial intelligence systems that can deal with unknown variables across unknown domains, we need benchmarks that measure how well these systems perform on tasks they have never seen before. A prerequisite for this is a measure of a task's generalization difficulty, or how dissimilar it is from the system's prior knowledge and experience. If the skill of an intelligence system in a particular domain is defined as it's ability to consistently generate a set of instructions (or programs) to solve tasks in that domain, current benchmarks do not quantitatively measure the efficiency of acquiring new skills, making it possible to brute-force skill acquisition by training with unlimited amounts of data and compute power. With this in mind, we first propose a common language of instruction, a programming language that allows the expression of programs in the form of directed acyclic graphs across a wide variety of real-world domains and computing platforms. Using programs generated in this language, we demonstrate a match-based method to both score performance and calculate the generalization difficulty of any given set of tasks. We use these to define a numeric benchmark called the generalization index, or the g-index, to measure and compare the skill-acquisition efficiency of any intelligence system on a set of real-world tasks. Finally, we evaluate the suitability of some well-known models as general intelligence systems by calculating their g-index scores.
31 pages, 15 Figures, 3 Tables; Sample Data and g-index Reference Code at https://github.com/mayahq/g-index-benchmark; g-index toy environment at https://github.com/mayahq/flatland; version 2 added a section about the toy environment; version 3 compressed images to reduce file size; version 4 updated description of flatland toy environment
References in corpus (11)
- Language Models are Few-Shot Learners
- Evaluating Large Language Models Trained on Code
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
- Neuro-Symbolic Program Synthesis
- Are we done with ImageNet?
- The Lovelace 2.0 Test of Artificial Creativity and Intelligence
- Unsupervised Translation of Programming Languages
- Program Synthesis with Large Language Models
- Comparison and Combination of State-of-the-art Techniques for Handwritten Character Recognition: Topping the MNIST Benchmark
- SPoC: Search-based Pseudocode to Code