2 papers
cs.CL2026
No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding
Michael Krumdick, Charles Lovering, Varshini Reddy +2
Reliable evaluation of large language models (LLMs) is critical as their deployment rapidly expands, particularly in high-stakes domains such as business and finance. The LLM-as-a-…
cs.AI2025
Language Model Probabilities are Not Calibrated in Numeric Contexts
Charles Lovering, Michael Krumdick, Viet Dac Lai +5
Some statements have one well-defined continuation (e.g., "the Eiffel Tower is in [Paris]"), whereas others have a natural distribution over multiple options (e.g., "the weighted c…