2 citations · 4 across the 3 of their papers we have counts for
3 papers
cs.CY2024★ 1 cited
Reasons to Doubt the Impact of AI Risk Evaluations
Gabriel Mukobi
AI safety practitioners invest considerable resources in AI system evaluations, but these investments may be wasted if evaluations fail to realize their impact. This paper question…
cs.CL2023★ 2 cited
SuperHF: Supervised Iterative Learning from Human Feedback
Gabriel Mukobi, Peter Chatain, Su Fong +4
While large language models demonstrate remarkable capabilities, they often present challenges in terms of safety, alignment with human values, and stability during training. Here,…
cs.MA2023★ 1 cited
Welfare Diplomacy: Benchmarking Language Model Cooperation
Gabriel Mukobi, Hannah Erlebach, Niklas Lauffer +3
The growing capabilities and increasingly widespread deployment of AI systems necessitate robust benchmarks for measuring their cooperative capabilities. Unfortunately, most multi-…