3 papers
cs.CL2026
Evaluating Scoring Bias in LLM-as-a-Judge
Qingquan Li, Shaoyu Dou, Kailai Shao +2
The "LLM-as-a-Judge" paradigm, using Large Language Models (LLMs) as automated evaluators, is pivotal to LLM development, offering scalable feedback for complex tasks. However, the…
cs.CL2025
FinEval-KR: A Financial Domain Evaluation Framework for Large Language Models' Knowledge and Reasoning
Shaoyu Dou, Yutian Shen, Mofan Chen +9
Large Language Models (LLMs) demonstrate significant potential but face challenges in complex financial reasoning tasks requiring both domain knowledge and sophisticated reasoning.…
cs.CL2025
INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
Shisong Chen, Qian Zhu, Wenyan Yang +15
Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capa…