4 papers
AI Learning and Conceptual Transfer in the Game of Hidden Rules
Christo Mathew, Wentian Wang, Jacob Feldman +4
This report summarizes the work conducted on the Game of Hidden Rules (GOHR), focusing on reinforcement learning agents trained to infer hidden rules from trial-and-error feedback,…
Interpretable GOHR Agents via Sparse Autoencoders
Shiwei Tan, Yusong Zhao, Weiyi Qin +6
A central challenge in interpreting learned decision-making systems is to determine whether their internal representations contain concepts that help explain their behavior. We rep…
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
Liwei Che, Zhiyu Xue, Yihao Quan +7
Counting serves as a simple but powerful test of a Large Vision-Language Model's (LVLM's) reasoning; it forces the model to identify each individual object and then add them all up…
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models
Wentian Wang, Sarthak Jain, Paul Kantor +3
We propose MMLU-SR, a novel dataset designed to measure the true comprehension abilities of Large Language Models (LLMs) by challenging their performance in question-answering task…