activity
20222025
most citedSolutions to preference manipulation in recommender systems require knowledge of meta-preferences

22 citations · 34 across the 6 of their papers we have counts for

collaborators
Showing cs.AIShow all

5 papers · 1 filter

cs.AI2025

Model-Free RL Agents Demonstrate System 1-Like Intentionality

Hal Ashton, Matija Franklin

This paper argues that model-free reinforcement learning (RL) agents, while lacking explicit planning mechanisms, exhibit behaviours that can be analogised to System 1 ("thinking f…

cs.AI202431 cited

Beyond Preferences in AI Alignment

Tan Zhi-Xuan, Micah Carroll, Matija Franklin +1

The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizi…

cs.AI20232 cited

Strengthening the EU AI Act: Defining Key Terms on AI Manipulation

Matija Franklin, Philip Moreira Tomei, Rebecca Gorman

The European Union's Artificial Intelligence Act aims to regulate manipulative and harmful uses of AI, but lacks precise definitions for key concepts. This paper provides technical…

cs.AI2023

Concept Extrapolation: A Conceptual Primer

Matija Franklin, Rebecca Gorman, Hal Ashton +1

This article is a primer on concept extrapolation - the ability to take a concept, a feature, or a goal that is defined in one context and extrapolate it safely to a more general c…

cs.AI20229 cited

Recognising the importance of preference change: A call for a coordinated multidisciplinary research effort in the age of AI

Matija Franklin, Hal Ashton, Rebecca Gorman +1

As artificial intelligence becomes more powerful and a ubiquitous presence in daily life, it is imperative to understand and manage the impact of AI systems on our lives and decisi…