6 papers
MultiDx: A Multi-Source Knowledge Integration Framework towards Diagnostic Reasoning
Yimin Deng, Zhenxi Lin, Yejing Wang +9
Diagnostic prediction and clinical reasoning are critical tasks in healthcare applications. While Large Language Models (LLMs) have shown strong capabilities in commonsense reasoni…
AdapTime: Enabling Adaptive Temporal Reasoning in Large Language Models
Yimin Deng, Yejing Wang, Zhenxi Lin +8
Large language models have demonstrated strong reasoning capabilities in general knowledge question answering. However, their ability to handle temporal information remains limited…
HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
Yiqian Yang, Tian Lan, Qianghuai Jia +6
Effective deep search agents must not only access open-domain and domain-specific knowledge but also apply complex rules-such as legal clauses, medical manuals and tariff rules. Th…
ADC: Enhancing Function Calling Via Adversarial Datasets and Code Line-Level Feedback
Wei Zhang, Yi Zhang, Li Zhu +5
Large Language Models (LLMs) have made significant strides in Natural Language Processing and coding, yet they struggle with robustness and accuracy in complex function calls. To t…
ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World
Weixiang Yan, Haitian Liu, Tengxiao Wu +9
LLMs have achieved significant performance progress in various NLP applications. However, LLMs still struggle to meet the strict requirements for accuracy and reliability in the me…
CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
Weixiang Yan, Haitian Liu, Yunkun Wang +8
Large Language Models (LLMs) have demonstrated remarkable performance on assisting humans in programming and facilitating programming automation. However, existing benchmarks for e…