1 citations · 1 across the 13 of their papers we have counts for
5 papers · 1 filter
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
Lena Libon, Ben Rank, Jehyeok Yeon +5
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such…
InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents
Jehyeok Yeon, Ben Rank, Maksym Andriushchenko
AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces. Even no…
Instrumental Choices: Measuring the Propensity of LLM Agents to Pursue Instrumental Behaviors
Jonas Wiedermann-Möller, Leonard Dung, Maksym Andriushchenko
AI systems have become increasingly capable of dangerous behaviours in many domains. This raises the question: Do models sometimes choose to violate human instructions in order to…
Characterizing the Consistency of the Emergent Misalignment Persona
Anietta Weckauff, Yuchen Zhang, Maksym Andriushchenko
Fine-tuning large language models (LLMs) on narrowly misaligned data generalizes to broadly misaligned behavior, a phenomenon termed emergent misalignment (EM). While prior work ha…
HalluHard: A Hard Multi-Turn Hallucination Benchmark
Dongyang Fan, Sebastien Delsad, Nicolas Flammarion +1
Large language models (LLMs) still produce plausible-sounding but ungrounded factual claims, a problem that worsens in multi-turn dialogue as context grows and early errors cascade…