Showing stat.MLShow all
3 papers · 1 filter
stat.ML2024
How to Choose a Threshold for an Evaluation Metric for Large Language Models
Bhaskarjit Sarmah, Mingshu Li, Jingrao Lyu +4
To ensure and monitor large language models (LLMs) reliably, various evaluation metrics have been proposed in the literature. However, there is little research on prescribing a met…
stat.ML2024
Can an unsupervised clustering algorithm reproduce a categorization system?
Nathalia Castellanos, Dhruv Desai, Sebastian Frank +2
Peer analysis is a critical component of investment management, often relying on expert-provided categorization systems. These systems' consistency is questioned when they do not a…
stat.ML2024
Enhanced Local Explainability and Trust Scores with Random Forest Proximities
Joshua Rosaler, Dhruv Desai, Bhaskarjit Sarmah +4
We initiate a novel approach to explain the predictions and out of sample performance of random forest (RF) regression and classification models by exploiting the fact that any RF…