collaborators

6 papers

cs.AI2026

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning

Lachlan McPheat, Navdeep Kaur, Robert Blackwell +3

We introduce DecompSR, decomposed spatial reasoning, a large benchmark dataset (over 5m datapoints) and generation framework designed to analyse compositional spatial reasoning abi…

cs.AI2026

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi

Anthony G. Cohn, Robert E. Blackwell

We introduce an extensive qualitative spatial and temporal reasoning (QSTR) benchmark for evaluating large language models (LLMs). We pose questions concerning compositional reason…

cs.CL2025

Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited

Anthony G Cohn, Robert E Blackwell

We investigate the abilities of 28 Large language Models (LLMs) to reason about cardinal directions (CDs) using a benchmark generated from a set of templates, extensively testing a…

cs.CL2025

Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores

Robert E. Blackwell, Jon Barry, Anthony G. Cohn

Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark s…

cs.CL2025

An Empirical Study of Conformal Prediction in LLM with ASP Scaffolds for Robust Reasoning

Navdeep Kaur, Lachlan McPheat, Alessandra Russo +2

In this paper, we examine the use of Conformal Language Modelling (CLM) alongside Answer Set Programming (ASP) to enhance the performance of standard open-weight LLMs on complex mu…

cs.ET2025

City Models: Past, Present and Future Prospects

Helge Ritter, Otthein Herzog, Kurt Rothermel +2

We attempt to take a comprehensive look at the challenges of representing the spatio-temporal structures and dynamic processes defining a city's overall characteristics. For the ta…