Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Benchmarking AI Agents for Addressing Scientific Challenges Across Scales
Tianyu Liu, Allen Xin Wang, Antonia Panescu +30
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchma…
cs.AI2026
MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models
Zeyu Zhang, Ziqing Wang, Kaize Ding
MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existing medical QA benchmarks largely test shor…