1 paper
Yu Wang, Craig Erickson, Kevin Small
LLM-based factuality judges provide scalable evaluation signals, but their metrics are often systematically biased relative to human judgments. We study human-anchored factuality e…