3 papers
cs.AI2026
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
Yihao Wang, Haoran Xu, Renjie Gu +10
The large-scale deployment of personalized healthcare agents demands memory mechanisms that are exceptionally precise, safe, and capable of long-term clinical tracking. However, ex…
cs.IR2026
Benchmarking Real-Time Question Answering via Executable Code Workflows
Wenjie Zhou, Yuan Gao, Xin Zhou +5
Retrieving real-time information is a fundamental capability for search-integrated agents in real-world applications. However, existing benchmarks are predominantly static and ther…
cs.AI2025
Bohrium + SciMaster: Building the Infrastructure and Ecosystem for Agentic Science at Scale
Linfeng Zhang, Siheng Chen, Yuzhu Cai +46
AI agents are emerging as a practical way to run multi-step scientific workflows that interleave reasoning with tool use and verification, pointing to a shift from isolated AI-assi…