1 paper
Cheng Wu, Vishal Anand, Jaya Krishna Mandivarapu +2
Generative AI systems are increasingly producing real-world artifacts, however their efficacy and validity are often evaluated via context-free LLM-scoring. These judges can be mis…