5 papers
AI Learning and Conceptual Transfer in the Game of Hidden Rules
Christo Mathew, Wentian Wang, Jacob Feldman +4
This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback,…
Interpretable GOHR Agents via Sparse Autoencoders
Shiwei Tan, Yusong Zhao, Weiyi Qin +6
A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help explain their behavior. We rep…
The mechanistic origin of branching-driven nucleation in abrupt phase transitions
Leyang Xue, Shengling Gao, Bnaya Gross +5
Phase transitions are the macroscopic manifestation of microscopic processes that drive a system towards a new state. The detailed evolution of these processes, particularly in abr…
Toward a Metrology for Artificial Intelligence: Hidden-Rule Environments and Reinforcement Learning
Christo Mathew, Wentian Wang, Jacob Feldman +4
We investigate reinforcement learning in the Game Of Hidden Rules (GOHR) environment, a complex puzzle in which an agent must infer and execute hidden rules to clear a 66 b…
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
Wentian Wang, Sarthak Jain, Paul Kantor +3
We propose MMLU-SR, a novel dataset designed to measure the true comprehension abilities of Large Language Models (LLMs) by challenging their performance in question-answering task…