2 papers
cs.CY2026
The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies
Benjamin Jensen, Ian Reynolds, Yasir Atalan +3
The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-a…
cs.CY2025
Critical Foreign Policy Decisions (CFPD)-Benchmark: Measuring Diplomatic Preferences in Large Language Models
Benjamin Jensen, Ian Reynolds, Yasir Atalan +4
As national security institutions increasingly integrate Artificial Intelligence (AI) into decision-making and content generation processes, understanding the inherent biases of la…