34 papers
HEARTS: Benchmarking LLM Reasoning on Health Time Series
Sirui Li, Shuhan Xiao, Mihir Joshi +4
The rise of large language models (LLMs) has shifted time series analysis from narrow analytics to general-purpose reasoning. Yet, existing benchmarks cover only a small set of hea…
An AI Co-Data-Scientist for Prioritizing Candidate Biomarkers from Wearable Sensor Data
Yubin Kim, Salman Rahman, Samuel Schmidgall +33
Wearable devices generate continuous physiological and behavioral data, but converting these signals into clinically reviewable biomarker hypotheses remains labor-intensive. We int…
RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills
Weizhi Zhang, Zechen Li, Hamid Palangi +16
The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities in healthcare access. However, large-sca…
WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning
Yuwei Zhang, Tong Xia, Bianca Emmerich +5
Language models are remarkably capable at medical question answering, in some cases surpassing the accuracy of general physicians. However, answering questions about wearable healt…
Evaluating the Utility of Personal Health Records in Personalized Health AI
Rory Sayres, Kejia Chen, Ayush Jain +19
Patient-managed Personal Health Records (PHRs) promises to empower patients to better understand their health; but information in the record is complex, potentially hindering insig…
Towards a General Intelligence and Interface for Wearable Health Data
Girish Narayanswamy, Maxwell A. Xu, A. Ali Heydari +37
While ubiquitous wearable sensors capture a wealth of behavioral and physiological information, effectively transforming these signals into personalized health insights is challeng…