9 papers
TCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine
Zihao Cheng, Yuheng Lu, Huaiqian Ye +10
Large Language Models (LLMs) have demonstrated remarkable capabilities in modern medicine, yet their application in Traditional Chinese Medicine (TCM) remains severely limited by t…
Learn More, Forget Less: A Gradient-Aware Data Selection Approach for LLM
Yibai Liu, Shihang Wang, Zeming Liu +5
Despite large language models (LLMs) have achieved impressive achievements across numerous tasks, supervised fine-tuning (SFT) remains essential for adapting these models to specia…
RETAIL: Towards Real-world Travel Planning for Large Language Models
Bin Deng, Yizhe Feng, Zeming Liu +5
Although large language models have enhanced automated travel planning abilities, current systems remain misaligned with real-world scenarios. First, they assume users provide expl…
ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations
Yiming Lei, Zhizheng Yang, Zeming Liu +5
Multi-modal large language models have demonstrated remarkable zero-shot abilities and powerful image-understanding capabilities. However, the existing open-source multi-modal mode…
GODBench: A Benchmark for Multimodal Large Language Models in Video Comment Art
Yiming Lei, Chenkai Zhang, Zeming Liu +5
Video Comment Art enhances user engagement by providing creative content that conveys humor, satire, or emotional resonance, requiring a nuanced and comprehensive grasp of cultural…
SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
Chenkai Zhang, Yiming Lei, Zeming Liu +5
With the rapid development of Multi-modal Large Language Models (MLLMs), an increasing number of benchmarks have been established to evaluate the video understanding capabilities o…