1 citations · 1 across the 3 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Instrumental convergence and power-seeking
David Thorstad
Recent years have seen increasing concern that artificial intelligence may soon pose an existential risk to humanity. One leading ground for concern is that artificial agents may b…
cs.AI2026
Revisiting the shutdown problem
David Thorstad
A key premise in leading arguments for existential risk from artificial intelligence is that malfunctioning artificial agents could not be easily shut down. This motivates the cata…
cs.AI2023★ 1 cited
Cognitive bias in large language models: Cautious optimism meets anti-Panglossian meliorism
David Thorstad
Traditional discussions of bias in large language models focus on a conception of bias closely tied to unfairness, especially as affecting marginalized groups. Recent work raises t…