1 paper
Dustin Arendt, Zhuanyi Huang, Prasha Shrestha +3
Evaluation beyond aggregate performance metrics, e.g. F1-score, is crucial to both establish an appropriate level of trust in machine learning models and identify future model impr…