collaborators

7 papers

cs.CL2025

AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs

María Victoria Carro, Denise Alejandra Mester, Facundo Nieto +9

The core premise of AI debate as a scalable oversight technique is that it is harder to lie convincingly than to refute a lie, enabling the judge to identify the correct position.…

cs.AI2025

Do Large Language Models Show Biases in Causal Learning? Insights from Contingency Judgment

María Victoria Carro, Denise Alejandra Mester, Francisca Gauna Selasco +4

Causal learning is the cognitive process of developing the capability of making causal inferences based on available information, often guided by normative principles. This process…

cs.AI2025

Error Detection and Correction for Interpretable Mathematics in Large Language Models

Yijin Yang, Cristina Cornelio, Mario Leiva +1

Recent large language models (LLMs) have demonstrated the ability to perform explicit multi-step reasoning such as chain-of-thought prompting. However, their intermediate steps oft…

cs.AI2025

A Conceptual Framework for AI Capability Evaluations

María Victoria Carro, Denise Alejandra Mester, Francisca Gauna Selasco +7

As AI systems advance and integrate into society, well-designed and transparent evaluations are becoming essential tools in AI governance, informing decisions by providing evidence…

cs.LG2025

Machine Learning Model Integration with Open World Temporal Logic for Process Automation

Dyuman Aditya, Colton Payne, Mario Leiva +1

Recent advances in Machine Learning (ML) have produced models that extract structured information from complex data. However, a significant challenge lies in translating these perc…

cs.LG2025

Multiple Distribution Shift -- Aerial (MDS-A): A Dataset for Test-Time Error Detection and Model Adaptation

Noel Ngu, Aditya Taparia, Gerardo I. Simari +5

Machine learning models assume that training and test samples are drawn from the same distribution. As such, significant differences between training and test distributions often l…