24 citations · 24 across the 2 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Ji Soo Lee, Xilun Chen, Pierce Chuang +5
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over…
cs.CL2025★ 24 cited
HealthBench: Evaluating Large Language Models Towards Improved Human Health
Rahul K. Arora, Jason Wei, Rebecca Soskin Hicks +9
We present HealthBench, an open-source benchmark measuring the performance and safety of large language models in healthcare. HealthBench consists of 5,000 multi-turn conversations…
cs.CL2024
GPT-4o System Card
OpenAI, :, Aaron Hurst +416
GPT-4o is an autoregressive omni model that accepts as input any combination of text, audio, image, and video, and generates any combination of text, audio, and image outputs. It's…