4 papers · 1 filter
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov +4
AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untru…
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
Akshit Sinha, Arvindh Arun, Shashwat Goel +2
Does continued scaling of large language models (LLMs) yield diminishing returns? In this work, we show that short-task benchmarks may give an illusion of slowing progress, as even…
Capability-Based Scaling Trends for LLM-Based Red-Teaming
Alexander Panfilov, Paul Kassianik, Maksym Andriushchenko +1
As large language models grow in capability and agency, identifying vulnerabilities through red-teaming becomes vital for safe deployment. However, traditional prompt-engineering a…
Fine, I'll Merge It Myself: A Multi-Fidelity Framework for Automated Model Merging
Guinan Su, Jonas Geiping
Reasoning capabilities represent a critical frontier for large language models (LLMs), but developing them requires extensive proprietary datasets and computational resources. One…