6 citations · 14 across the 10 of their papers we have counts for
3 papers · 1 filter
Scalable AI Safety via Doubly-Efficient Debate
Jonah Brown-Cohen, Geoffrey Irving, Georgios Piliouras
The emergence of pre-trained AI systems with powerful capabilities across a diverse and ever-increasing set of complex domains has raised a critical challenge for AI safety as task…
Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models
Dingli Yu, Simran Kaur, Arushi Gupta +3
With LLMs shifting their role from statistical modeling of language to serving as general-purpose AI agents, how should LLM evaluations change? Arguably, a key ability of an AI age…
Detecting Adversarial Directions in Deep Reinforcement Learning to Make Robust Decisions
Ezgi Korkmaz, Jonah Brown-Cohen
Learning in MDPs with highly complex state representations is currently possible due to multiple advancements in reinforcement learning algorithm design. However, this incline in c…