activity
20242026
most citedFinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information

1 citations · 1 across the 4 of their papers we have counts for

collaborators

8 papers

cs.CL2026

NTDH: Complex Reasoning for Comprehensive Affective Analysis

Tianlei Zhu, Zhiwei Liu, Yuyan Wang +2

Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is…

cs.AI2026

Herculean: An Agentic Benchmark for Financial Intelligence

Xueqing Peng, Zhuohan Xie, Yupeng Cao +60

As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…

cs.CL2026

FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs

Yan Wang, Keyi Wang, Shanshan Yang +12

Going beyond simple text processing, financial auditing requires detecting semantic, structural, and numerical inconsistencies across large-scale disclosures. As financial reports…

cs.CL20261 cited

FinTagging: Benchmarking LLMs for Extracting and Structuring Financial Information

Yan Wang, Lingfei Qian, Xueqing Peng +18

Accurate interpretation of numerical data in financial reports is critical for markets and regulators. Although XBRL (eXtensible Business Reporting Language) provides a standard fo…

cs.CL2026

CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance

Haochen Liu, Weien Li, Rui Song +6

Large language model (LLM) systems are increasingly used to support high-stakes decision-making, but they typically perform worse when the available evidence is internally inconsis…

cs.CL2025

MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application

Xueqing Peng, Lingfei Qian, Yan Wang +44

Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings. Yet most existing eval…