8 papers
RADAR: Benchmarking Language Models on Imperfect Tabular Data
Ken Gu, Zhihan Zhang, Kate Lin +18
Language models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness -- the ability to recognize, reason over, and appropriately…
Learn Globally, Speak Locally: Bridging the Gaps in Multilingual Reasoning
Jaedong Hwang, Kumar Tanmay, Seok-Jin Lee +5
Large Language Models (LLMs) have achieved strong performance in domains like mathematics, factual question answering, and code generation, yet their ability to reason on these tas…
The Anatomy of a Personal Health Agent
A. Ali Heydari, Ken Gu, Vidya Srinivas +35
Health is a fundamental pillar of human wellness, and the rapid advancements in large language models (LLMs) have driven the development of a new generation of health agents. Howev…
Transforming Wearable Data into Personal Health Insights using Large Language Model Agents
Mike A. Merrill, Akshay Paruchuri, Naghmeh Rezaei +17
Deriving personalized insights from popular wearable trackers requires complex numerical reasoning that challenges standard LLMs, necessitating tool-based approaches like code gene…
Substance over Style: Evaluating Proactive Conversational Coaching Agents
Vidya Srinivas, Xuhai Xu, Xin Liu +5
While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coachi…
SensorLM: Learning the Language of Wearable Sensors
Yuwei Zhang, Kumar Ayush, Siyuan Qiao +17
We present SensorLM, a family of sensor-language foundation models that enable wearable sensor data understanding with natural language. Despite its pervasive nature, aligning and…