activity
20242026
most citedFlattering to Deceive: The Impact of Sycophantic Behavior on User Trust in Large Language Model

6 citations · 6 across the 8 of their papers we have counts for

collaborators

8 papers

cs.AI2026

Human Attribution of Causality to AI Across Agency, Misuse, and Misalignment

Maria Victoria Carro, David Lagnado

AI-related incidents are becoming increasingly frequent and severe, ranging from safety failures to misuse by malicious actors. In such complex situations, identifying which elemen…

cs.LG2026

Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering

Eitan Sprejer, Oscar Agustín Stanchi, María Victoria Carro +2

Feature steering has emerged as a promising approach for controlling LLM behavior through direct manipulation of internal representations, offering advantages over prompt engineeri…

cs.CL2025

AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs

María Victoria Carro, Denise Alejandra Mester, Facundo Nieto +9

The core premise of AI debate as a scalable oversight technique is that it is harder to lie convincingly than to refute a lie, enabling the judge to identify the correct position.…

cs.AI2025

Do Large Language Models Show Biases in Causal Learning? Insights from Contingency Judgment

María Victoria Carro, Denise Alejandra Mester, Francisca Gauna Selasco +4

Causal learning is the cognitive process of developing the capability of making causal inferences based on available information, often guided by normative principles. This process…

cs.AI2025

A Conceptual Framework for AI Capability Evaluations

María Victoria Carro, Denise Alejandra Mester, Francisca Gauna Selasco +7

As AI systems advance and integrate into society, well-designed and transparent evaluations are becoming essential tools in AI governance, informing decisions by providing evidence…

cs.AI2024

Do Large Language Models Show Biases in Causal Learning?

Maria Victoria Carro, Francisca Gauna Selasco, Denise Alejandra Mester +4

Causal learning is the cognitive process of developing the capability of making causal inferences based on available information, often guided by normative principles. This process…