2 papers
cs.AI2026
RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
Tianjun Pan, Xuan Lin, Wenyan Yang +7
Rubric-based evaluation has become a prevailing paradigm for evaluating instruction following in large language models (LLMs). Despite its widespread use, the reliability of these…
cs.SE2026
Detecting Underspecification in Software Requirements via k-NN Coverage Geometry
Wenyan Yang, Tomáš Janovec, Samantha Bavautdin
We propose \geogap{}, a geometric method for detecting missing requirement types in software specifications. The method represents each requirement as a unit vector via a pretraine…