6 papers
TCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine
Zihao Cheng, Yuheng Lu, Huaiqian Ye +10
Large Language Models (LLMs) have demonstrated remarkable capabilities in modern medicine, yet their application in Traditional Chinese Medicine (TCM) remains severely limited by t…
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
Jingjing Liu, Zeming Liu, Zihao Cheng +7
Large Language Models (LLMs) have exhibited significant proficiency in code debugging, especially in automatic program repair, which may substantially reduce the time consumption o…
TiKMiX: Take Data Influence into Dynamic Mixture for Language Model Pre-training
Yifan Wang, Binbin Liu, Fengze Liu +6
The data mixture used in the pre-training of a language model is a cornerstone of its final performance. However, a static mixing strategy is suboptimal, as the model's learning pr…
RETAIL: Towards Real-world Travel Planning for Large Language Models
Bin Deng, Yizhe Feng, Zeming Liu +5
Although large language models have enhanced automated travel planning abilities, current systems remain misaligned with real-world scenarios. First, they assume users provide expl…
ToolSpectrum : Towards Personalized Tool Utilization for Large Language Models
Zihao Cheng, Hongru Wang, Zeming Liu +4
While integrating external tools into large language models (LLMs) enhances their ability to access real-time information and domain-specific services, existing approaches focus na…
Deepfake Detection via Knowledge Injection
Tonghui Li, Yuanfang Guo, Zeming Liu +2
Deepfake detection technologies become vital because current generative AI models can generate realistic deepfakes, which may be utilized in malicious purposes. Existing deepfake d…