1 paper
Alexandra Abbas, Celia Waggoner, Justin Olive
AI evaluations have become critical tools for assessing large language model capabilities and safety. This paper presents practical insights from eight months of maintaining $inspe…