3 papers
cs.AI2026
Overthinking: Amplifying Reasoning Weights to Extract Learned Secrets
Jack Hopkins, Dipika Khullar, Fabien Roger
Black box auditing of language models is an essential pre-deployment tool, but it may miss subtle forms of misalignment and hidden information. To better elicit hidden information…
cs.AI2026
Self-Attribution Bias: When AI Monitors Go Easy on Themselves
Dipika Khullar, Jack Hopkins, Rowan Wang +1
Agentic systems increasingly rely on language models to monitor their own behavior. For example, coding agents may self critique generated code for pull request approval or assess…
cs.MA2025
Factorio Learning Environment
Jack Hopkins, Mart Bakler, Akbir Khan
Large Language Models (LLMs) are rapidly saturating existing benchmarks, necessitating new open-ended evaluations. We introduce the Factorio Learning Environment (FLE), based on th…