3 papers
cs.AI2026
DiG-bench: Discovery in Games
Ruairidh M. Battleday, Kai Sandbrink, Jimi Cullen-Drohan +13
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few…
q-fin.ST2025
AlphaAgents: Large Language Model based Multi-Agents for Equity Portfolio Constructions
Tianjiao Zhao, Jingrao Lyu, Stokes Jones +3
The field of artificial intelligence (AI) agents is evolving rapidly, driven by the capabilities of Large Language Models (LLMs) to autonomously perform and refine tasks with human…
stat.ML2024
How to Choose a Threshold for an Evaluation Metric for Large Language Models
Bhaskarjit Sarmah, Mingshu Li, Jingrao Lyu +4
To ensure and monitor large language models (LLMs) reliably, various evaluation metrics have been proposed in the literature. However, there is little research on prescribing a met…