3 papers
cs.CL2026
QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models
Yao Wu, Kangping Yin, Liang Dong +13
While Large Language Models (LLMs) excel on standardized medical exams, high scores often fail to translate to high-quality responses for real-world medical queries. Current evalua…
cs.CL2025
LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs
Da-Chen Lian, Ri-Sheng Huang, Pin-Er Chen +7
We propose LingBench++, a linguistically-informed benchmark and reasoning framework designed to evaluate large language models (LLMs) on complex linguistic tasks inspired by the In…
cs.CL2025
Spontaneous Speech Variables for Evaluating LLMs Cognitive Plausibility
Sheng-Fu Wang, Laurent Prevot, Jou-an Chi +2
The achievements of Large Language Models in Natural Language Processing, especially for high-resource languages, call for a better understanding of their characteristics from a co…