1 paper
Jennifer Healey, Laurie Byrum, Md Nadeem Akhtar +2
LLM evaluation is challenging even the case of base models. In real world deployments, evaluation is further complicated by the interplay of task specific prompts and experiential…