3 papers
cs.LG2026
DRACO: a Cross-Domain Benchmark for Deep Research Accuracy, Completeness, and Objectivity
Joey Zhong, Hao Zhang, Clare Southern +7
We present DRACO (Deep Research Accuracy, Completeness, and Objectivity), a benchmark of complex deep research tasks. These tasks, which span 10 domains and draw on information sou…
cs.AI2025
Assessing LLMs' Performance: Insights from the Chinese Pharmacist Exam
Xinran Wang, Boran Zhu, Shujuan Zhou +3
Background: As large language models (LLMs) become increasingly integrated into digital health education and assessment workflows, their capabilities in supporting high-stakes, dom…
cs.IR2025
Research on Evaluation Methods for Patent Novelty Search Systems and Empirical Analysis
Shu Zhang, LiSha Zhang, Kai Duan +1
Patent novelty search systems are critical to IP protection and innovation assessment; their retrieval accuracy directly impacts patent quality. We propose a comprehensive evaluati…