Showing cs.CLShow all
3 papers · 1 filter
cs.CL2025
Evaluating the Ability of Large Language Models to Reason about Cardinal Directions, Revisited
Anthony G Cohn, Robert E Blackwell
We investigate the abilities of 28 Large language Models (LLMs) to reason about cardinal directions (CDs) using a benchmark generated from a set of templates, extensively testing a…
cs.CL2025
Towards Reproducible LLM Evaluation: Quantifying Uncertainty in LLM Benchmark Scores
Robert E. Blackwell, Jon Barry, Anthony G. Cohn
Large language models (LLMs) are stochastic, and not all models give deterministic answers, even when setting temperature to zero with a fixed random seed. However, few benchmark s…
cs.CL2024
Can Large Language Models Reason about the Region Connection Calculus?
Anthony G Cohn, Robert E Blackwell
Qualitative Spatial Reasoning is a well explored area of Knowledge Representation and Reasoning and has multiple applications ranging from Geographical Information Systems to Robot…