Showing cs.AIShow all
2 papers · 1 filter
cs.AI2025
Establishing Best Practices for Building Rigorous Agentic Benchmarks
Yuxuan Zhu, Tengjun Jin, Yada Pruksachatkun +22
Benchmarks are essential for quantitatively tracking progress in AI. As AI agents become increasingly capable, researchers and practitioners have introduced agentic benchmarks to e…
cs.AI2025
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language
Christopher Summerfield, Lennart Luettgau, Magda Dubois +9
We examine recent research that asks whether current AI systems may be developing a capacity for "scheming" (covertly and strategically pursuing misaligned goals). We compare curre…