5 papers
The Behavioural Reflection Test: A time-efficient measure of reflective reasoning in morally and epistemically charged decisions
Sion Weatherhead, Flora Salim, Aaron Belbasis +1
How readily people override intuitive conclusions through reflection shapes how they navigate dense information environments with reliable and misleading sources; yet the effective…
SOCIA-EVO: Automated Simulator Construction via Dual-Anchored Bi-Level Optimization
Yuncheng Hua, Sion Weatherhead, Mehdi Jafari +2
Automated simulator construction requires distributional fidelity, distinguishing it from generic code generation. We identify two failure modes in long-horizon LLM agents: context…
SOCIA-: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
Yuncheng Hua, Sion Weatherhead, Mehdi Jafari +2
In this paper, we present SOCIA-, an end-to-end, agentic framework that treats simulator construction asinstance optimization over code within a textual computation graph.…
SOCIA-Nabla: Textual Gradient Meets Multi-Agent Orchestration for Automated Simulator Generation
Yuncheng Hua, Sion Weatherhead, Mehdi Jafari +2
In this paper, we present SOCIA-Nabla, an end-to-end, agentic framework that treats simulator construction asinstance optimization over code within a textual computation graph. Spe…
Illusions of reflection: open-ended task reveals systematic failures in Large Language Models' reflective reasoning
Sion Weatherhead, Flora Salim, Aaron Belbasis
Humans do not just find mistakes after the fact -- we often catch them mid-stream because 'reflection' is tied to the goal and its constraints. Today's large language models produc…