From the 1 of 4 linked papers with an AI index.
4 papers
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Lena Libon, Ben Rank, Jehyeok Yeon +5
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such…
What Does It Mean to Break a Distillation Defense?
Lena Libon, Pura Peetathawatchai, Michael Aerni +2
The paper examines how to evaluate defenses that add noise to large language model outputs to thwart distillation attacks, proposing a three‑dimensional threat model (query budget,…
RewardUQ: A Unified Framework for Uncertainty-Aware Reward Models
Daniel Yang, Samuel Stante, Florian Redhardt +5
Reward models are central to aligning large language models (LLMs) with human preferences. Yet most approaches rely on pointwise reward estimates that overlook the epistemic uncert…
Conscious Data Contribution via Community-Driven Chain-of-Thought Distillation
Lena Libon, Meghana Bhange, Rushabh Solanki +2
The current era of AI development places a heavy emphasis on training large models on increasingly scaled-up datasets. This paradigm has catalyzed entirely new product categories,…