7 papers
Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering
Zhuohan Xie, Xueqing Peng, Georgi Georgiev +18
FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence. Each final-test item pairs an English question with financial statements and n…
Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering
Zhuohan Xie, Yuyang Dai, Rania Elbadry +18
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the corr…
FinReporting: An Agentic Workflow for Localized Reporting of Cross-Jurisdiction Financial Disclosures
Fan Zhang, Mingzi Song, Rania Elbadry +20
Financial reporting systems increasingly leverage Large Language Models (LLMs) to extract and summarize corporate disclosures. However, most existing approaches assume a single-mar…
When Alpha Disappears: A One-Switch Benchmark for Decision-Time Leakage in Financial Backtests
Fan Zhang, Zhen Li, Sijia Peng +1
We introduce When Alpha Disappears, a paired evaluation benchmark for diagnosing decision-time leakage in financial machine-learning backtests. Rather than treating leakage as a bi…
FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering
Yixi Zhou, Fan Zhang, Yu Chen +3
Financial question answering (QA) over long corporate filings requires evidence to satisfy strict constraints on entities, financial metrics, fiscal periods, and numeric values. Ho…
SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation
Yixi Zhou, Fan Zhang, Zhiqiao Guo +4
Despite strong performance on Text-to-SQL benchmarks, it remains unclear whether LLM-generated SQL programs are structurally reliable. In this work, we investigate the structural b…