14 citations · 14 across the 1 of their papers we have counts for
4 papers
Measuring AI Ability to Complete Long Software Tasks
Thomas Kwa, Ben West, Joel Becker +23
Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities,…
Automated Detection of Malignant Lesions in the Ovary Using Deep Learning Models and XAI
Md. Hasin Sarwar Ifty, Nisharga Nirjan, Labib Islam +4
The unrestrained proliferation of cells that are malignant in nature is cancer. In recent times, medical professionals are constantly acquiring enhanced diagnostic and treatment ab…
The science and practice of proportionality in AI risk evaluations
Carlos Mougan, Lauritz Morlock, Jair Aguirre +19
A global challenge in artificial intelligence (AI) regulation lies in achieving effective risk management without compromising innovation and technical progress. The European Union…
HCAST: Human-Calibrated Autonomy Software Tasks
David Rein, Joel Becker, Amy Deng +19
To understand and predict the societal impacts of highly autonomous AI systems, we need benchmarks with grounding, i.e., metrics that directly connect AI performance to real-world…