3 papers
cs.AI2025
Erkang-Diagnosis-1.1 Technical Report
Jianbing Ma, Ao Feng, Zhenjie Gao +5
This report provides a detailed introduction to Erkang-Diagnosis-1.1 model, our AI healthcare consulting assistant developed using Alibaba Qwen-3 model. The Erkang model integrates…
cs.CL2025
Can Large Language Models Master Complex Card Games?
Wei Wang, Fuqing Bie, Junzhe Chen +4
Complex games have long been an important benchmark for testing the progress of artificial intelligence algorithms. AlphaGo, AlphaZero, and MuZero have defeated top human players i…
cs.CL2025
DataSciBench: An LLM Agent Benchmark for Data Science
Dan Zhang, Sining Zhoubian, Min Cai +7
This paper presents DataSciBench, a comprehensive benchmark for evaluating Large Language Model (LLM) capabilities in data science. Recent related benchmarks have primarily focused…