4 papers
DiG-bench: Discovery in Games
Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few…
STRAPSim: A Portfolio Similarity Metric for ETF Alignment and Portfolio Trades
Mingshu Li, Dhruv Desai, Jerinsh Jeyapaulraj +4
Accurately measuring portfolio similarity is critical for a wide range of financial applications, including Exchange-traded Fund (ETF) recommendation, portfolio trading, and risk a…
Explainable Unsupervised Anomaly Detection with Random Forest
Joshua S. Harvey, Joshua Rosaler, Mingshu Li +2
We describe the use of an unsupervised Random Forest for similarity learning and improved unsupervised anomaly detection. By training a Random Forest to discriminate between real d…
How to Choose a Threshold for an Evaluation Metric for Large Language Models
Bhaskarjit Sarmah, Mingshu Li, Jingrao Lyu +4
To ensure and monitor large language models (LLMs) reliably, various evaluation metrics have been proposed in the literature. However, there is little research on prescribing a met…