Assessing Logical Puzzle Solving in Large Language Models: Insights from a Minesweeper Case Study
arXiv:2311.07387 · doi:10.18653/v1/2024.naacl-long.4
Abstract
Large Language Models (LLMs) have shown remarkable proficiency in language understanding and have been successfully applied to a variety of real-world tasks through task-specific fine-tuning or prompt engineering. Despite these advancements, it remains an open question whether LLMs are fundamentally capable of reasoning and planning, or if they primarily rely on recalling and synthesizing information from their training data. In our research, we introduce a novel task -- Minesweeper -- specifically designed in a format unfamiliar to LLMs and absent from their training datasets. This task challenges LLMs to identify the locations of mines based on numerical clues provided by adjacent opened cells. Successfully completing this task requires an understanding of each cell's state, discerning spatial relationships between the clues and mines, and strategizing actions based on logical deductions drawn from the arrangement of the cells. Our experiments, including trials with the advanced GPT-4 model, indicate that while LLMs possess the foundational abilities required for this task, they struggle to integrate these into a coherent, multi-step logical reasoning process needed to solve Minesweeper. These findings highlight the need for further research to understand the nature of reasoning capabilities in LLMs under similar circumstances, and to explore pathways towards more sophisticated AI reasoning and planning models.
23 pages, 5 figures, 4 tables, in NAACL 2024
References in corpus (45)
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Sparks of Artificial General Intelligence: Early experiments with GPT-4
- Scaling Laws for Neural Language Models
- Emergent Abilities of Large Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models
- Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
- A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
- Voyager: An Open-Ended Embodied Agent with Large Language Models
- ChemCrow: Augmenting large-language models with chemistry tools
- Are Emergent Abilities of Large Language Models a Mirage?
- Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
- Evaluating the Logical Reasoning Ability of ChatGPT and GPT-4
- One Small Step for Generative AI, One Giant Leap for AGI: A Complete Survey on ChatGPT in AIGC Era
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
- Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents
- The False Promise of Imitating Proprietary LLMs
- A Constructive Prediction of the Generalization Error Across Scales
- PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change
- LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving
- Large Language Models Cannot Self-Correct Reasoning Yet
- LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
- How well do Large Language Models perform in Arithmetic tasks?
- Towards Generalist Biomedical AI
- Large Language Models Can Self-Improve
- The Chess Transformer: Mastering Play using Generative Language Models
- Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence
- Large Language Models are Better Reasoners with Self-Verification
- GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems
- Skywork: A More Open Bilingual Foundation Model
- Are Large Language Models Really Good Logical Reasoners? A Comprehensive Evaluation and Beyond
- Can Large Language Models Really Improve by Self-critiquing Their Own Plans?
- SPRING: Studying the Paper and Reasoning to Play Games
- ChessGPT: Bridging Policy Learning and Language Modeling
- Can Large Language Models Play Text Games Well? Current State-of-the-Art and Open Questions
- Investigating Data Contamination in Modern Benchmarks for Large Language Models
- Pretraining Data Mixtures Enable Narrow Model Selection Capabilities in Transformer Models
- Suspicion-Agent: Playing Imperfect Information Games with Theory of Mind Aware GPT-4
- BioPlanner: Automatic Evaluation of LLMs on Protocol Planning in Biology
- Are ChatGPT and GPT-4 Good Poker Players? -- A Pre-Flop Analysis
- Assessing and Enhancing the Robustness of Large Language Models with Task Structure Variations for Logical Reasoning
- Thespian: Multi-Character Text Role-Playing Game Agents
- Tabular Representation, Noisy Operators, and Impacts on Table Structure Understanding Tasks in LLMs
- A criterion for Artificial General Intelligence: hypothetic-deductive reasoning, tested on ChatGPT