◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Joe Benton

3 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • last author1

Across the 1 of 3 papers where every author was matched, so the position is known.

fields
  • cs.LG2
  • cs.CL1
ORCID 0000-0002-2103-6112
same name
  • Joe Benton — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20232025
most citedReasoning Models Don't Always Say What They Think

9 citations · 12 across the 3 of their papers we have counts for

collaborators

3 papers

cs.CL2025★ 9 cited

Reasoning Models Don't Always Say What They Think

Yanda Chen, Joe Benton, Ansh Radhakrishnan +12

Chain-of-thought (CoT) offers a potential boon for AI safety as it allows monitoring a model's CoT to try to understand its intentions and reasoning processes. However, the effecti…

cs.LG2024★ 2 cited

Sabotage Evaluations for Frontier Models

Joe Benton, Misha Wagner, Eric Christiansen +13

Sufficiently capable models could subvert human oversight and decision-making in important contexts. For example, in the context of AI development, models could covertly sabotage e…

cs.LG2023★ 1 cited

Measuring Feature Sparsity in Language Models

Mingyang Deng, Lucas Tao, Joe Benton

Recent works have proposed that activations in language models can be modelled as sparse linear combinations of vectors corresponding to features of input text. Under this assumpti…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.