90 citations · 149 across the 10 of their papers we have counts for
Showing 2023Show all
2 papers · 1 filter
cs.AI2023
Testing GPT-4 with Wolfram Alpha and Code Interpreter plug-ins on math and science problems
Ernest Davis, Scott Aaronson
This report describes a test of the large language model GPT-4 with the Wolfram Alpha and the Code Interpreter plug-ins on 105 original problems in science and math, at the high sc…
cs.AI2023★ 6 cited
Benchmarks for Automated Commonsense Reasoning: A Survey
Ernest Davis
More than one hundred benchmarks have been developed to test the commonsense knowledge and commonsense reasoning abilities of artificial intelligence (AI) systems. However, these b…