collaborators

6 papers

cs.LG2026

Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering

Eitan Sprejer, Oscar Agustín Stanchi, María Victoria Carro +2

Feature steering has emerged as a promising approach for controlling LLM behavior through direct manipulation of internal representations, offering advantages over prompt engineeri…

cs.LG2025

Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity

Austin Meek, Eitan Sprejer, Iván Arcuschin +2

Chain-of-thought (CoT) outputs let us read a model's step-by-step reasoning. Since any long, serial reasoning process must pass through this textual trace, the quality of the CoT i…

cs.CL2025

AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs

María Victoria Carro, Denise Alejandra Mester, Facundo Nieto +9

The core premise of AI debate as a scalable oversight technique is that it is harder to lie convincingly than to refute a lie, enabling the judge to identify the correct position.…

cs.AI2025

Approximating Human Preferences Using a Multi-Judge Learned System

Eitán Sprejer, Fernando Avalos, Augusto Bernardi +3

Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Ove…

hep-ph2025

Automatizing the search for mass resonances using BumpNet

Jean-François Arguin, Georges Azuelos, Émile Baril +15

Physics Beyond the Standard Model (BSM) has yet to be observed at the Large Hadron Collider (LHC), motivating the development of model-agnostic, machine learning-based strategies t…

physics.data-an2025

Automatizing the search for mass resonances using BumpNet

Jean-Francois Arguin, Georges Azuelos, Émile Baril +15

The search for resonant mass bumps in invariant-mass distributions remains a cornerstone strategy for uncovering Beyond the Standard Model (BSM) physics at the Large Hadron Collide…