2 papers
cs.CL2026
Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
So Hasegawa, Shailaja Keyur Sampat, Lei Liu +1
Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables…
cs.AI2026
EHR-Complex: Benchmarking Medical Agents for Complex Clinical Reasoning
Yitong Qiao, Lei Liu, Yue Shen +4
Clinical agents promise to democratize access to electronic health records (EHRs), yet existing benchmarks fail to reflect the complexity of practical EHR analysis, e.g., often ope…