FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance
arXiv:2508.05201 · doi:10.1145/3768292.3770433
Abstract
Hallucination remains a critical challenge for deploying Large Language Models (LLMs) in finance. Accurate extraction and precise calculation from tabular data are essential for reliable financial analysis, since even minor numerical errors can undermine decision-making and regulatory compliance. Financial applications have unique requirements, often relying on context-dependent, numerical, and proprietary tabular data that existing hallucination benchmarks rarely capture. In this study, we develop a rigorous and scalable framework for evaluating intrinsic hallucinations in financial LLMs, conceptualized as a context-aware masked span prediction task over real-world financial documents. Our main contributions are: (1) a novel, automated dataset creation paradigm using a masking strategy; (2) a new hallucination evaluation dataset derived from S&P 500 annual reports; and (3) a comprehensive evaluation of intrinsic hallucination patterns in state-of-the-art LLMs on financial tabular data. Our work provides a robust methodology for in-house LLM evaluation and serves as a critical step toward building more trustworthy and reliable financial Generative AI systems.
9 pages, AMC ICAIF'25
References in corpus (9)
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Deficiency of Large Language Models in Finance: An Empirical Examination of Hallucination
- HalluDial: A Large-Scale Benchmark for Automatic Dialogue-Level Hallucination Evaluation
- FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
- RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
- Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge
- ERBench: An Entity-Relationship based Automatically Verifiable Hallucination Benchmark for Large Language Models
- ANAH-v2: Scaling Analytical Hallucination Annotation of Large Language Models
- FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information