From the 1 of 4 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
A dataset of rated conceptual arguments
Emery Cooper, Caspar Oesterheld, Linh Chi Nguyen +2
The paper introduces a dataset of 951 expert‑rated argumentative critiques on 442 position texts covering AI safety, decision theory, ethics, and politics, and uses it to benchmark…
cs.AI2025
Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety
Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38
AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…