3 papers
cs.AI2026
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
Tianjun Pan, Xuan Lin, Wenyan Yang +7
Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these…
cs.CL2025
INSEva: A Comprehensive Chinese Benchmark for Large Language Models in Insurance
Shisong Chen, Qian Zhu, Wenyan Yang +15
Insurance, as a critical component of the global financial system, demands high standards of accuracy and reliability in AI applications. While existing benchmarks evaluate AI capa…
cs.SE2025
AdaptiveLog: An Adaptive Log Analysis Framework with the Collaboration of Large and Small Language Model
Lipeng Ma, Weidong Yang, Yixuan Li +6
Automated log analysis is crucial to ensure high availability and reliability of complex systems. The advent of LLMs in NLP has ushered in a new era of language model-driven automa…