2 papers
cs.AI2026
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Thomson Yen, Julian Poeltl, Harshith Srinivas Gear +10
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs h…
cs.LG2024
AiSciVision: A Framework for Specializing Large Multimodal Models in Scientific Image Classification
Brendan Hogan, Anmol Kabra, Felipe Siqueira Pacheco +10
Trust and interpretability are crucial for the use of Artificial Intelligence (AI) in scientific research, but current models often operate as black boxes offering limited transpar…