4 papers
Mechanisms of Introspective Awareness
Uzay Macar, Li Yang, Atticus Wang +3
Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introsp…
Automatically Finding Reward Model Biases
Atticus Wang, Iván Arcuschin, Arthur Conmy
Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…
Simple Mechanistic Explanations for Out-Of-Context Reasoning
Atticus Wang, Joshua Engels, Oliver Clive-Griffin +2
Out-of-context reasoning (OOCR) is a phenomenon in which fine-tuned LLMs exhibit surprisingly deep out-of-distribution generalization. Rather than learning shallow heuristics, they…
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents
Kaivalya Hariharan, Uzay Girit, Atticus Wang +1
Benchmarks for large language models (LLMs) have predominantly assessed short-horizon, localized reasoning. Existing long-horizon suites (e.g. SWE-bench) rely on manually curated i…