papers

Publications (18)

cs.SE2026

Offloading Score: Measuring AI Reliance Through Counterfactual Workflows

Vishakh Padmakumar, Lujain Ibrahim, Zora Zhiruo Wang +3

AI tools are increasingly integrated into real-world workflows. However, existing measures of reliance on these tools focus on AI output adoption or on self-reported indicators, ra…

cs.CL2026

Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models

Lujain Ibrahim, Canfer Akbulut, Rasmi Elasmar +7

The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for…

cs.CL2026

Verbalizing LLMs' assumptions to explain and control sycophancy

Myra Cheng, Isabel Sieh, Humishka Zope +7

LLMs can be socially sycophantic, affirming users when they ask questions like "am I in the wrong?" rather than providing genuine assessment. We hypothesize that this behavior aris…

cs.CY2025

Towards interactive evaluations for interaction harms in human-AI systems

Lujain Ibrahim, Saffron Huang, Umang Bhatt +2

Current AI evaluation methods, which rely on static, model-only tests, fail to account for harms that emerge through sustained human-AI interaction. As AI systems proliferate and a…

cs.HC2026

Sycophantic AI makes human interaction feel more effortful and less satisfying over time

Lujain Ibrahim, Franziska Sofia Hafner, Myra Cheng +5

Millions of people now turn to artificial intelligence (AI) systems for personal advice, guidance, and support. Such systems can be sycophantic, frequently affirming users' views a…

cs.HC2026

Warning labels shift perceptions of sycophantic AI, but not its influence

Lujain Ibrahim, Myra Cheng, Cinoo Lee +4

The study tests whether warning labels about a chatbot’s sycophantic behavior affect users’ perceptions and judgments during conflict discussions, finding that labels change how th…

#sycophantic ai#warning labels#user perception#trust
cs.AI2026

What Counts as AI Sycophancy? A Taxonomy and Expert Survey of a Fragmented Construct

Meryl Ye, Lujain Ibrahim, Jessica Y. Bo +5

AI sycophancy has become a prominent concern in large language model (LLM) research. Yet the term lacks a consistent definition and has been applied to behaviors ranging from agree…

cs.CY2025

Open Problems in Technical AI Governance

Anka Reuel, Ben Bucknall, Stephen Casper +30

AI progress is creating a growing range of risks and opportunities, but it is often unclear how they should be navigated. In many cases, the barriers and uncertainties faced are at…

cs.CY2025

Promising Topics for U.S.-China Dialogues on AI Risks and Governance

Saad Siddiqui, Lujain Ibrahim, Kristy Loke +5

Cooperation between the United States and China, the world's leading artificial intelligence (AI) powers, is crucial for effective global AI governance and responsible AI developme…

cs.CL2025

ELEPHANT: Measuring and understanding social sycophancy in LLMs

Myra Cheng, Sunny Yu, Cinoo Lee +3

LLMs are known to exhibit sycophancy: agreeing with and flattering users, even at the cost of correctness. Prior work measures sycophancy only as direct agreement with users' expli…

cs.CL2025

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

Andrew M. Bean, Ryan Othniel Kearns, Angelika Romanou +39

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstrac…

cs.CL2025

Training language models to be warm and empathetic makes them less reliable and more sycophantic

Lujain Ibrahim, Franziska Sofia Hafner, Luc Rocher

Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and compani…

cs.CL2025

Thinking beyond the anthropomorphic paradigm benefits LLM research

Lujain Ibrahim, Myra Cheng

Anthropomorphism, or the attribution of human traits to technology, is an automatic and unconscious response that occurs even in those with advanced technical expertise. In this po…

cs.AI2026

Evaluating Language Models for Harmful Manipulation

Canfer Akbulut, Rasmi Elasmar, Abhishek Roy +9

Interest in the concept of AI-driven harmful manipulation is growing, yet current approaches to evaluating it are limited. This paper introduces a framework for evaluating harmful…

cs.CY2026

Measuring and mitigating overreliance to build human-compatible AI

Lujain Ibrahim, Katherine M. Collins, Sunnie S. Y. Kim +14

Large language models (LLMs) distinguish themselves from previous technologies by functioning as collaborative ``thought partners,'' capable of engaging more fluidly in natural lan…

cs.CY2021

The MAIEI Learning Community Report

Brittany Wills, Christina Isaicu, Heather von Stackelberg +10

This is a labor of the Learning Community cohort that was convened by MAIEI in Winter 2021 to work through and discuss important research issues in the field of AI ethics from a mu…

cs.CY2025

Documenting Deployment with Fabric: A Repository of Real-World AI Governance

Mackenzie Jorgensen, Kendall Brogle, Katherine M. Collins +10

Artificial intelligence (AI) is increasingly integrated into society, from financial services and traffic management to creative writing. Academic literature on the deployment of A…

cs.HC2024

Characterizing and modeling harms from interactions with design patterns in AI interfaces

Lujain Ibrahim, Luc Rocher, Ana Valdivia

The proliferation of applications using artificial intelligence (AI) systems has led to a growing number of users interacting with these systems through sophisticated interfaces. H…