benchmark 1layout analysis 1multimodal reasoning 1pdf document understanding 1table and chart comprehension 1
From the 1 of 12 linked papers with an AI index.
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Machine Learning as a Tool (MLAT): A Framework for Integrating Statistical ML Models as Callable Tools within LLM Agent Workflows
Edwin Chen, Zulekha Bibi
We introduce Machine Learning as a Tool (MLAT), a design pattern in which pre-trained statistical machine learning models are exposed as callable tools within large language model…
cs.LG2025
FrontierCS: Evolving Challenges for Evolving Intelligence
Qiuyang Mang, Wenhao Chai, Zhifei Li +48
We introduce FrontierCS, a benchmark of 156 open-ended problems across diverse areas of computer science, designed and reviewed by experts, including CS PhDs and top-tier competiti…