3 papers
cs.AI2025
Testing the Machine Consciousness Hypothesis
Stephen Fitz
The Machine Consciousness Hypothesis states that consciousness is a substrate-free functional property of computational systems capable of second-order perception. I propose a rese…
cs.AI2025
Psychometric Personality Shaping Modulates Capabilities and Safety in Language Models
Stephen Fitz, Peter Romero, Steven Basart +2
Large Language Models increasingly mediate high-stakes interactions, intensifying research on their capabilities and safety. While recent work has shown that LLMs exhibit consisten…
cs.LG2024
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
Richard Ren, Steven Basart, Adam Khoja +9
As artificial intelligence systems grow more powerful, there has been increasing interest in "AI safety" research to address emerging and future risks. However, the field of AI saf…