219 citations · 235 across the 2 of their papers we have counts for
4 papers
Toward Trustworthy AI Development: Mechanisms for Supporting Verifiable Claims
Miles Brundage, Shahar Avin, Jasmine Wang +56
With the recent wave of progress in artificial intelligence (AI) has come a growing awareness of the large-scale impacts of AI systems, and recognition that existing regulations an…
Explainable Machine Learning in Deployment
Umang Bhatt, Alice Xiang, Shubham Sharma +7
Explainable machine learning offers the potential to provide stakeholders with insights into model behavior by using various methods such as feature importance scores, counterfactu…
Theories of Parenting and their Application to Artificial Intelligence
Sky Croeser, Peter Eckersley
As machine learning (ML) systems have advanced, they have acquired more power over humans' lives, and questions about what values are embedded in them have become more complex and…
Impossibility and Uncertainty Theorems in AI Value Alignment (or why your AGI should not have a utility function)
Peter Eckersley
Utility functions or their equivalents (value functions, objective functions, loss functions, reward functions, preference orderings) are a central tool in most current machine lea…