2 papers
cs.AI2026
Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets
Jillian Ross, Eric So, Zoe De Simone +2
Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individ…
cs.CY2026
The Foreign Policy AI Evaluation Gap
Charles Pozniak, Jeba Sania
We argue that AI systems used in conducting foreign policy tasks - broadly enacting 'statecraft' - should be a priority test case for technical AI governance research. In enacting…