8 papers · 1 filter
Time, Identity and Consciousness in Language Model Agents
Elija Perrier, Michael Timothy Bennett
Machine consciousness evaluations mostly see behavior. For language model agents that behavior is language and tool use. That lets an agent say the right things about itself even w…
Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning
Elija Perrier
While Chain-of-Thought (CoT) prompting enhances the reasoning capabilities of large language models, the faithfulness of the generated rationales remains an open problem for model…
Agent Identity Evals: Measuring Agentic Identity
Elija Perrier, Michael Timothy Bennett
Central to agentic capability and trustworthiness of language model agents (LMAs) is the extent they maintain stable, reliable, identity over time. However, LMAs inherit pathologie…
Towards Measurement Theory for Artificial Intelligence
Elija Perrier
We motivate and outline a programme for a formal theory of measurement of artificial intelligence. We argue that formalising measurement for AI will allow researchers, practitioner…
Out of Control -- Why Alignment Needs Formal Control Theory (and an Alignment Control Stack)
Elija Perrier
This position paper argues that formal optimal control theory should be central to AI alignment research, offering a distinct perspective from prevailing AI safety and security app…
Statistical Scenario Modelling and Lookalike Distributions for Multi-Variate AI Risk
Elija Perrier
Evaluating AI safety requires statistically rigorous methods and risk metrics for understanding how the use of AI affects aggregated risk. However, much AI safety literature focuse…