6 papers
Mind the Performance Gap: Capability-Behavior Trade-offs in Feature Steering
Eitan Sprejer, Oscar AgustÃn Stanchi, MarÃa Victoria Carro +2
Feature steering has emerged as a promising approach for controlling LLM behavior through direct manipulation of internal representations, offering advantages over prompt engineeri…
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
Austin Meek, Eitan Sprejer, Iván Arcuschin +2
Chain-of-thought (CoT) outputs let us read a model's step-by-step reasoning. Since any long, serial reasoning process must pass through this textual trace, the quality of the CoT i…
AI Debaters are More Persuasive when Arguing in Alignment with Their Own Beliefs
MarÃa Victoria Carro, Denise Alejandra Mester, Facundo Nieto +9
The core premise of AI debate as a scalable oversight technique is that it is harder to lie convincingly than to refute a lie, enabling the judge to identify the correct position.…
Approximating Human Preferences Using a Multi-Judge Learned System
Eitán Sprejer, Fernando Avalos, Augusto Bernardi +3
Aligning LLM-based judges with human preferences is a significant challenge, as they are difficult to calibrate and often suffer from rubric sensitivity, bias, and instability. Ove…
Automatizing the search for mass resonances using BumpNet
Jean-François Arguin, Georges Azuelos, Ãmile Baril +15
Physics Beyond the Standard Model (BSM) has yet to be observed at the Large Hadron Collider (LHC), motivating the development of model-agnostic, machine learning-based strategies t…
Automatizing the search for mass resonances using BumpNet
Jean-Francois Arguin, Georges Azuelos, Ãmile Baril +15
The search for resonant mass bumps in invariant-mass distributions remains a cornerstone strategy for uncovering Beyond the Standard Model (BSM) physics at the Large Hadron Collide…