Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
Large Language Models' Reasoning Stalls: An Investigation into the Capabilities of Frontier Models
Lachlan McGinness, Peter Baumgartner
Empirical methods to examine the capability of Large Language Models (LLMs) to use Automated Theorem Prover (ATP) reasoning strategies are studied. We evaluate the performance of S…
cs.AI2025
Large Language Models Imitate Logical Reasoning, but at what Cost?
Lachlan McGinness, Peter Baumgartner
We present a longitudinal study which evaluates the reasoning capability of frontier Large Language Models over an eighteen month period. We measured the accuracy of three leading…
cs.AI2025
The AlphaPhysics Term Rewriting System for Marking Algebraic Expressions in Physics Exams
Peter Baumgartner, Lachlan McGinness
We present our method for automatically marking Physics exams. The marking problem consists in assessing typed student answers for correctness with respect to a ground truth soluti…