Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Reference-Free Rating of LLM Responses via Latent Information
Leander Girrbach, Chi-Ping Su, Tankred Saanum +3
How reliable are single-response LLM-as-a-judge ratings without references, and can we obtain fine-grained, deterministic scores in this setting? We study the common practice of as…
cs.CL2024
Granular Change Accuracy: A More Accurate Performance Metric for Dialogue State Tracking
Taha Aksu, Nancy F. Chen
Current metrics for evaluating Dialogue State Tracking (DST) systems exhibit three primary limitations. They: i) erroneously presume a uniform distribution of slots throughout the…