24 citations · 34 across the 2 of their papers we have counts for
3 papers
cs.CL2025★ 24 cited
HealthBench: Evaluating Large Language Models Towards Improved Human Health
Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks +9
We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations…
cs.CL2024
GPT-4o System Card
OpenAI, :, Aaron Hurst +416
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's…
cs.LG2021★ 10 cited
Fairness On The Ground: Applying Algorithmic Fairness Approaches to Production Systems
Chloé Bakalar, Renata Barreto, Stevie Bergman +13
Many technical approaches have been proposed for ensuring that decisions made by machine learning systems are fair, but few of these proposals have been stress-tested in real-world…