1 paper · 1 filter
Davood Wadi, Mohsen Ghodrat, Matthew Philp
As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evalua…