2 papers
cs.DB2025
RADAR: Benchmarking Language Models on Imperfect Tabular Data
Ken Gu, Zhihan Zhang, Kate Lin +18
Language models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness -- the ability to recognize, reason over, and appropriately…
cs.AI2025
The Anatomy of a Personal Health Agent
A. Ali Heydari, Ken Gu, Vidya Srinivas +35
Health is a fundamental pillar of human wellness, and the rapid advancements in large language models (LLMs) have driven the development of a new generation of health agents. Howev…