◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Stephen Casper

5 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author5

Across the 5 of 5 papers where every author was matched, so the position is known.

fields
  • cs.LG3
  • cs.AI1
  • cs.CY1
same name
  • Stephen Casper — 4 papers
  • Stephen Casper — 2 papers, h 14
  • Stephen Casper — 1 paper

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

most citedInternational AI Safety Report

15 citations · 26 across the 5 of their papers we have counts for

collaborators
Showing cs.LGShow all

3 papers · 1 filter

cs.LG2025

Adversarial Alignment for LLMs Requires Simpler, Reproducible, and More Measurable Objectives

Leo Schwinn, Yan Scholten, Tom Wollschläger +4

Misaligned research objectives have considerably hindered progress in adversarial robustness research over the past decade. For instance, an extensive focus on optimizing target me…

cs.LG2025★ 9 cited

Open Problems in Mechanistic Interpretability

Lee Sharkey, Bilal Chughtai, Joshua Batson +26

Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goa…

cs.LG2025

Open Problems in Machine Unlearning for AI Safety

Fazl Barez, Tingchen Fu, Ameya Prabhu +16

As AI systems become more capable, widely deployed, and increasingly autonomous in critical areas such as cybersecurity, biological research, and healthcare, ensuring their safety…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.