2 papers
cs.CL2026
GRAFITE: Generative Regression Analysis Framework for Issue Tracking and Evaluation
Ja Young Lee, MÃrian Silva, Mohamed Nasr +6
Large language models (LLMs) are largely motivated by their performance on popular topics and benchmarks at the time of their release. However, over time, contamination occurs due…
cs.AI2026
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
Xiyuan Zhang, Huihang Wu, Jiayu Guo +8
We introduce FIRE, a comprehensive benchmark designed to evaluate both the theoretical financial knowledge of LLMs and their ability to handle practical business scenarios. For the…