113 citations · 361 across the 6 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024★ 11 cited
Evaluating Frontier Models for Dangerous Capabilities
Mary Phuong, Matthew Aitchison, Elliot Catt +24
To understand the risks posed by a new AI system, we must understand what it can and cannot do. Building on prior work, we introduce a programme of new "dangerous capability" evalu…
cs.LG2023★ 7 cited
Simfluence: Modeling the Influence of Individual Training Examples by Simulating Training Runs
Kelvin Guu, Albert Webson, Ellie Pavlick +3
Training data attribution (TDA) methods offer to trace a model's prediction on any given example back to specific influential training examples. Existing approaches do so by assign…