4 papers
When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
Xiaoqing Wu, Xingyu Fan, Feifei Li +1
Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain structurally valid and semanticall…
WikiKV: Schema-Evolving Path-Indexed Storage for Hierarchical Knowledge Navigation
Feifei Li, Haoliang Ming, Zihan Li +5
LLM-curated hierarchical knowledge bases, namely a tree-structured wiki whose nodes summarize an underlying corpus, have become a dominant substrate for retrieval-augmented applica…
EpiPlanAgent: Agentic Automated Epidemic Response Planning
Kangkun Mao, Fang Xu, Jinru Ding +8
Epidemic response planning is essential yet traditionally reliant on labor-intensive manual methods. This study aimed to design and evaluate EpiPlanAgent, an agent-based system usi…
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
Jinru Ding, Lu Lu, Chao Ding +15
Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We…