works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.CL2026

Can LLMs Write Reliable Rubrics? A Meta-Evaluation for Experiment Reproduction

Hanhua Hong, Yizhi Li, Jiaoyan Chen +4

The paper conducts a meta‑evaluation of rubrics generated by large language models for assessing the reproducibility of research papers, comparing intrinsic semantic similarity and…

cs.SD2026

The Achilles' Heel of Angular Margins: A Chebyshev Polynomial Fix for Speaker Verification

Yang Wang, Yiqi Liu, Chenghao Xiao +1

Angular margin losses, such as AAM-Softmax, have become the de facto in speaker and face verification. Their success hinges on directly manipulating the angle between features and…

cs.CL2026

RIGOURATE: Quantifying Scientific Exaggeration with Evidence-Aligned Claim Evaluation

Joseph James, Chenghao Xiao, Yucheng Li +2

Scientific rigour tends to be sidelined in favour of bold statements, leading authors to overstate claims beyond what their results support. We present RIGOURATE, a two-stage multi…

cs.CL2025

ContrastScore: Towards Higher Quality, Less Biased, More Efficient Evaluation Metrics with Contrastive Evaluation

Xiao Wang, Daniil Larionov, Siwei Wu +4

Evaluating the quality of generated text automatically remains a significant challenge. Conventional reference-based metrics have been shown to exhibit relatively weak correlation…

cs.CL2025

Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts

Hanhua Hong, Chenghao Xiao, Yang Wang +3

Evaluating natural language generation systems is challenging due to the diversity of valid outputs. While human evaluation is the gold standard, it suffers from inconsistencies, l…