2 papers
cs.AI2026
FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance
Wolfgang M. Pauli, Sarah Panda, Kidus Admassu +3
Recent advances in large language models have accelerated deployment of agentic systems in operational finance. Existing benchmarks emphasize measuring general capabilities, instru…
cs.CL2024
Evaluation Methodology for Large Language Models for Multilingual Document Question and Answer
Adar Kahana, Jaya Susan Mathew, Said Bleik +2
With the widespread adoption of Large Language Models (LLMs), in this paper we investigate the multilingual capability of these models. Our preliminary results show that, translati…