Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Evidence Lock Before Commitment: A Frozen Interface Degrades LLM-as-Judge Evaluation
Divyansh Singh
LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate record preserves the information…
cs.CL2026
RADAR: Rubric-Aware Dependency and Redundancy Analysis for LLM-as-Judge Evaluation
Divyansh Singh, Reza Davari, Afra Mashhadi
Rubric-based LLM-as-judge pipelines often assume that evaluation criteria provide independent signals. In practice, however, criteria can be behaviorally coupled: improving one cri…