3 papers
cs.AI2026
GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory
Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham +3
Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving mu…
cs.CL2026
Scheming Ability in LLM-to-LLM Strategic Interactions
Thao Pham
As large language model (LLM) agents are deployed autonomously in diverse contexts, evaluating their capacity for strategic deception becomes crucial. While recent research has exa…
cs.CL2025
Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow
Kyle Moore, Jesse Roberts, Thao Pham +1
Language models are known to absorb biases from their training data, leading to predictions driven by statistical regularities rather than semantic relevance. We investigate the im…