◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

D. Thorstad

3 papers hereh-index 11338 citations40 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • sole author3

Across the 3 of 3 papers where every author was matched, so the position is known.

fields
  • cs.AI3

identity via Semantic Scholar / OpenAlex

most citedCognitive bias in large language models: Cautious optimism meets anti-Panglossian meliorism

1 citations · 1 across the 3 of their papers we have counts for

collaborators

3 papers

cs.AI2026

Instrumental convergence and power-seeking

David Thorstad

Recent years have seen increasing concern that artificial intelligence may soon pose an existential risk to humanity. One leading ground for concern is that artificial agents may b…

cs.AI2026

Revisiting the shutdown problem

David Thorstad

A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut down. This motivates the cata…

cs.AI2023★ 1 cited

Cognitive bias in large language models: Cautious optimism meets anti-Panglossian meliorism

David Thorstad

Traditional discussions of bias in large language models focus on a conception of bias closely tied to unfairness, especially as affecting marginalized groups. Recent work raises t…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.