Mathematical Capabilities of ChatGPT
arXiv:2301.13867
Abstract
We investigate the mathematical capabilities of two iterations of ChatGPT (released 9-January-2023 and 30-January-2023) and of GPT-4 by testing them on publicly available datasets, as well as hand-crafted ones, using a novel methodology. In contrast to formal mathematics, where large databases of formal proofs are available (e.g., the Lean Mathematical Library), current datasets of natural-language mathematics, used to benchmark language models, either cover only elementary mathematics or are very small. We address this by publicly releasing two new datasets: GHOSTS and miniGHOSTS. These are the first natural-language datasets curated by working researchers in mathematics that (1) aim to cover graduate-level mathematics, (2) provide a holistic overview of the mathematical capabilities of language models, and (3) distinguish multiple dimensions of mathematical reasoning. These datasets also test whether ChatGPT and GPT-4 can be helpful assistants to professional mathematicians by emulating use cases that arise in the daily professional activities of mathematicians. We benchmark the models on a range of fine-grained performance metrics. For advanced mathematics, this is the most detailed evaluation effort to date. We find that ChatGPT can be used most successfully as a mathematical assistant for querying facts, acting as a mathematical search engine and knowledge base interface. GPT-4 can additionally be used for undergraduate-level mathematics but fails on graduate-level difficulty. Contrary to many positive reports in the media about GPT-4 and ChatGPT's exam-solving abilities (a potential case of selection bias), their overall mathematical performance is well below the level of a graduate student. Hence, if your goal is to use ChatGPT to pass a graduate-level math exam, you would be better off copying from your average peer!
Added further evaluations on another ChatGPT version and on GPT-4. The GHOSTS and miniGHOSTS datasets are available at https://github.com/xyfrieder/science-GHOSTS
Cited by in corpus (18)
- Summary of ChatGPT-Related Research and Perspective Towards the Future of Large Language Models
- Let's have a chat! A Conversation with ChatGPT: Technology, Applications, and Limitations
- Enhancing STEM Learning with ChatGPT and Bing Chat as Objects to Think With: A Case Study
- How understanding large language models can inform the use of ChatGPT in physics education
- The Impact of AI in Physics Education: A Comprehensive Review from GCSE to University Levels
- ChatGPT in the classroom. Exploring its potential and limitations in a Functional Programming course
- Can we trust the evaluation on ChatGPT?
- Empirical assessment of ChatGPT's answering capabilities in natural science and engineering
- Towards AI-Assisted Synthesis of Verified Dafny Methods
- The Use of Generative Artificial Intelligence for Upper Secondary Mathematics Education Through the Lens of Technology Acceptance
- Human I/O: Towards a Unified Approach to Detecting Situational Impairments
- GPT-assisted learning of structure-property relationships by graph neural networks: Application to rare-earth doped phosphors
- Computational Argumentation-based Chatbots: a Survey
- Using AI Large Language Models for Grading in Education: A Hands-On Test for Physics
- Artificial Intelligence and Nuclear Weapons Proliferation: The Technological Arms Race for (In)visibility
- Unmasking the giant: A comprehensive evaluation of ChatGPT's proficiency in coding algorithms and data structures
- Do GPT Language Models Suffer From Split Personality Disorder? The Advent Of Substrate-Free Psychometrics
- AutoMathKG: The automated mathematical knowledge graph based on LLM and vector database