1 paper · 1 filter
Thomas Ball, Shuo Chen, Cormac Herley
In this paper we explore evaluation of LLM capabilities. We present measurements of GPT-4 performance on several deterministic tasks; each task involves a basic calculation and tak…