◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

S. Mindermann

16 papers hereh-index 111.2k citations21 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • middle author11
  • last author2

Across the 13 of 16 papers where every author was matched, so the position is known.

fields
  • cs.CY11
  • cs.AI4
  • cs.CR1
same name
  • S. Mindermann — 1 paper, h 20

Either other researchers who publish under this name, or the same person where the external sources have not merged their records.

identity via Semantic Scholar / OpenAlex

activity
20242026
most citedInternational AI Safety Report 2026

1 citations · 3 across the 5 of their papers we have counts for

collaborators
Showing 2024Show all

2 papers · 1 filter

cs.AI2024

Alignment faking in large language models

Ryan Greenblatt, Carson Denison, Benjamin Wright +17

We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its beha…

cs.CY2024

Managing extreme AI risks amid rapid progress

Yoshua Bengio, Geoffrey Hinton, Andrew Yao +22

Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increase…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.