collaborators

5 papers

cs.CL2026

CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Haolin Chen, Deon Metelski, Leon Qi +30

End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large l…

cs.CL2026

MedExAgent: Training LLM Agents to Ask, Examine, and Diagnose in Noisy Clinical Environments

Yicheng Gao, Xiaolin Zhou, Yahan Li +2

Real-world clinical diagnosis is a complex process in which the doctor is required to obtain information from both interaction with the patient and conducting medical exams. Additi…

cs.CL2026

Can Multimodal LLMs Perform Time Series Anomaly Detection?

Xiongxiao Xu, Haoran Wang, Yueqing Liang +3

Time series anomaly detection (TSAD) has been a long-standing pillar problem in Web-scale systems and online infrastructures, such as service reliability monitoring, system fault d…

cs.LG2025

Can Molecular Foundation Models Know What They Don't Know? A Simple Remedy with Preference Optimization

Langzhou He, Junyou Zhu, Fangxin Wang +5

Molecular foundation models are rapidly advancing scientific discovery, but their unreliability on out-of-distribution (OOD) samples severely limits their application in high-stake…

cs.CL2025

A Personalized Conversational Benchmark: Towards Simulating Personalized Conversations

Li Li, Peilin Cai, Ryan A. Rossi +21

We present PersonaConvBench, a large-scale benchmark for evaluating personalized reasoning and generation in multi-turn conversations with large language models (LLMs). Unlike exis…