2 papers
cs.CL2026
Measuring Competency, Not Performance: Item-Aware Evaluation Across Medical Benchmarks
Zhimeng Luo, Lixin Wu, Adam Frisch +1
Accuracy-based evaluation of Large Language Models (LLMs) measures benchmark-specific performance rather than underlying medical competency: it treats all questions as equally info…
cs.LG2025
Confucius3-Math: A Lightweight High-Performance Reasoning LLM for Chinese K-12 Mathematics Learning
Lixin Wu, Na Cai, Qiao Cheng +2
We introduce Confucius3-Math, an open-source large language model with 14B parameters that (1) runs efficiently on a single consumer-grade GPU; (2) achieves SOTA performances on a…