5 papers
DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning
Lachlan McPheat, Navdeep Kaur, Robert Blackwell +3
We introduce DecompSR, decomposed spatial reasoning, a large benchmark dataset (over 5m datapoints) and generation framework designed to analyse compositional spatial reasoning abi…
QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi
Anthony G. Cohn, Robert E. Blackwell
We introduce an extensive qualitative spatial and temporal reasoning (QSTR) benchmark for evaluating large language models (LLMs). We pose questions concerning compositional reason…
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
Anthony G Cohn, Robert E Blackwell
We investigate the abilities of 28 Large language Models (LLMs) to reason about cardinal directions (CDs) using a benchmark generated from a set of templates, extensively testing a…
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
Robert E. Blackwell, Jon Barry, Anthony G. Cohn
Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark s…
Can Large Language Models Reason about the Region Connection Calculus?
Anthony G Cohn, Robert E Blackwell
Qualitative Spatial Reasoning is a well explored area of Knowledge Representation and Reasoning and has multiple applications ranging from Geographical Information Systems to Robot…