Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Reference-Free Rating of LLM Responses via Latent Information
Leander Girrbach, Chi-Ping Su, Tankred Saanum +3
How reliable are single-response LLM-as-a-judge ratings without references, and can we obtain fine-grained, deterministic scores in this setting? We study the common practice of as…
cs.CL2024
Inducing anxiety in large language models can induce bias
Julian Coda-Forno, Kristin Witte, Akshay K. Jagadish +3
Large language models (LLMs) are transforming research on machine learning while galvanizing public debates. Understanding not only when these models work well and succeed but also…