activity
20242026
collaborators

5 papers

cs.AI2026

GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory

Pepijn Cobben, Xuanqiang Angelo Huang, Thao Amelia Pham +3

Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving mu…

cs.CL2026

Scheming Ability in LLM-to-LLM Strategic Interactions

Thao Pham

As large language model (LLM) agents are deployed autonomously in diverse contexts, evaluating their capacity for strategic deception becomes crucial. While recent research has exa…

cs.CL2025

Chain of Thought Still Thinks Fast: APriCoT Helps with Thinking Slow

Kyle Moore, Jesse Roberts, Thao Pham +1

Language models are known to absorb biases from their training data, leading to predictions driven by statistical regularities rather than semantic relevance. We investigate the im…

cs.CL2024

The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance

Kyle Moore, Jesse Roberts, Thao Pham +2

Cloze testing is a common method for measuring the behavior of large language models on a number of benchmark tasks. Using the MMLU dataset, we show that the base-rate probability…

cs.CL2024

Large Language Model Recall Uncertainty is Modulated by the Fan Effect

Jesse Roberts, Kyle Moore, Thao Pham +2

This paper evaluates whether large language models (LLMs) exhibit cognitive fan effects, similar to those discovered by Anderson in humans, after being pre-trained on human textual…