2 papers
cs.AI2026
Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis
Jiayu Fu, Mourad Heddaya, Chenhao Tan
Numerous math benchmarks exist to evaluate LLMs' mathematical capabilities. However, most involve extensive manual effort and are difficult to scale. Consequently, they cannot keep…
cs.LG2025
FAITH: A Framework for Assessing Intrinsic Tabular Hallucinations in Finance
Mengao Zhang, Jiayu Fu, Tanya Warrier +3
Hallucination remains a critical challenge for deploying Large Language Models (LLMs) in finance. Accurate extraction and precise calculation from tabular data are essential for re…