6 papers
When History Lies: Evaluating and Improving Tool Use under Misleading Multi-Turn Histories
Xiaoqing Wu, Xingyu Fan, Feifei Li +1
Tool-calling agents infer task state from accumulated dialogue and tool traces. In persistent interactions, however, historical traces may remain structurally valid and semanticall…
WikiKV: Schema-Evolving Path-Indexed Storage for Hierarchical Knowledge Navigation
Feifei Li, Haoliang Ming, Zihan Li +5
LLM-curated hierarchical knowledge bases, namely a tree-structured wiki whose nodes summarize an underlying corpus, have become a dominant substrate for retrieval-augmented applica…
Retrieval as Reasoning: Self-Evolving Agent-Native Retrieval via LLM-Wiki
Haoliang Ming, Feifei Li, Xiaoqing Wu +1
LLM agents require retrieval to behave less like one-shot context fetching and more like reasoning: searching, reading, traversing, and deciding when evidence is sufficient. Yet cu…
AuthTrace: Diagnosing Evidence Construction in Thematically Dense Single-Author Corpora
Xiaoqing Wu, Feifei Li, Haoliang Ming +1
Evidence construction--the stage that determines which passages reach the language model before generation begins--is evaluated paradigm by paradigm, leaving practitioners with no…
EpiPlanAgent: Agentic Automated Epidemic Response Planning
Kangkun Mao, Fang Xu, Jinru Ding +8
Epidemic response planning is essential yet traditionally reliant on labor-intensive manual methods. This study aimed to design and evaluate EpiPlanAgent, an agent-based system usi…
MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents
Jinru Ding, Lu Lu, Chao Ding +15
Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We…