◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Atticus Wang

4 papers hereh-index 212 citations4 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author2

Across the 4 of 4 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.CL1

identity via Semantic Scholar / OpenAlex

collaborators

4 papers

cs.LG2026

Mechanisms of Introspective Awareness

Uzay Macar, Li Yang, Atticus Wang +3

Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introsp…

cs.LG2026

Automatically Finding Reward Model Biases

Atticus Wang, Iván Arcuschin, Arthur Conmy

Reward models are central to large language model (LLM) post-training. However, past work has shown that they can reward spurious or undesirable attributes such as length, format,…

cs.CL2025

Simple Mechanistic Explanations for Out-Of-Context Reasoning

Atticus Wang, Joshua Engels, Oliver Clive-Griffin +2

Out-of-context reasoning (OOCR) is a phenomenon in which fine-tuned LLMs exhibit surprisingly deep out-of-distribution generalization. Rather than learning shallow heuristics, they…

cs.LG2025

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents

Kaivalya Hariharan, Uzay Girit, Atticus Wang +1

Benchmarks for large language models (LLMs) have predominantly assessed short-horizon, localized reasoning. Existing long-horizon suites (e.g. SWE-bench) rely on manually curated i…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.