1 paper · 1 filter
Anton Kolonin, Alexey Glushchenko, Evgeny Bochkov +1
Evaluating the reasoning capabilities of Large Language Models (LLMs) for complex, quantitative financial tasks is a critical and unsolved challenge. Standard benchmarks often fail…