Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Mechanisms of Introspective Awareness
Uzay Macar, Li Yang, Atticus Wang +3
Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introsp…
cs.LG2026
Automatically Finding Reward Model Biases
Atticus Wang, Iván Arcuschin, Arthur Conmy
Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…
cs.LG2025
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents
Kaivalya Hariharan, Uzay Girit, Atticus Wang +1
Benchmarks for large language models (LLMs) have predominantly assessed short-horizon, localized reasoning. Existing long-horizon suites (e.g. SWE-bench) rely on manually curated i…