4 papers
Watson & Holmes: A Naturalistic Benchmark for Comparing Human and LLM Reasoning
Thatchawin Leelawat, Lewis D Griffin
Existing benchmarks for AI reasoning provide limited insight into how closely these capabilities resemble human reasoning in naturalistic contexts. We present an adaptation of the…
Frequency and Generalisation of Periodic Activation Functions in Reinforcement Learning
Augustine N. Mavor-Parker, Matthew J. Sargent, Caswell Barry +2
Periodic activation functions, often referred to as learned Fourier features have been widely demonstrated to improve sample efficiency and stability in a variety of deep RL algori…
Evaluating Cultural Adaptability of a Large Language Model via Simulation of Synthetic Personas
Louis Kwok, Michal Bravansky, Lewis D. Griffin
The success of Large Language Models (LLMs) in multicultural environments hinges on their ability to understand users' diverse cultural backgrounds. We measure this capability by h…
Transcript of GPT-4 playing a rogue AGI in a Matrix Game
Lewis D Griffin, Nicholas Riggs
Matrix Games are a type of unconstrained wargame used by planners to explore scenarios. Players propose actions, and give arguments and counterarguments for their success. An umpir…