2 citations · 4 across the 10 of their papers we have counts for
5 papers · 1 filter
RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills
Weizhi Zhang, Zechen Li, Hamid Palangi +16
The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities in healthcare access. However, large-sca…
SynthWorlds: Controlled Parallel Worlds for Disentangling Reasoning and Knowledge in Language Models
Ken Gu, Advait Bhat, Mike A Merrill +4
Evaluating the reasoning ability of language models (LMs) is complicated by their extensive parametric world knowledge, where benchmark performance often reflects factual recall ra…
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
Yubin Kim, Hyewon Jeong, Shan Chen +24
Hallucinations in foundation models arise from autoregressive training objectives that prioritize token-likelihood optimization over epistemic accuracy, fostering overconfidence an…
Substance over Style: Evaluating Proactive Conversational Coaching Agents
Vidya Srinivas, Xuhai Xu, Xin Liu +5
While NLP research has made strides in conversational tasks, many approaches focus on single-turn responses with well-defined objectives or evaluation criteria. In contrast, coachi…
BehaviorSFT: Behavioral Token Conditioning for Clinical Agents Across the Proactivity Spectrum
Yubin Kim, Zhiyuan Hu, Hyewon Jeong +11
Large Language Models (LLMs) as clinical agents require careful behavioral adaptation. While adept at reactive tasks (e.g., diagnosis reasoning), LLMs often struggle with proactive…