3 papers
cs.AI2026
DiG-bench: Discovery in Games
Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few…
stat.ML2024
How to Choose a Threshold for an Evaluation Metric for Large Language Models
Bhaskarjit Sarmah, Mingshu Li, Jingrao Lyu +4
To ensure and monitor large language models (LLMs) reliably, various evaluation metrics have been proposed in the literature. However, there is little research on prescribing a met…
stat.ML2024
Can an unsupervised clustering algorithm reproduce a categorization system?
Nathalia Castellanos, Dhruv Desai, Sebastian Frank +2
Peer analysis is a critical component of investment management, often relying on expert-provided categorization systems. These systems' consistency is questioned when they do not a…