1 paper
Xiting Wang, Liming Jiang, Jose Hernandez-Orallo +4
Comprehensive and accurate evaluation of general-purpose AI systems such as large language models allows for effective mitigation of their risks and deepened understanding of their…