20 papers
NTDH: Complex Reasoning for Comprehensive Affective Analysis
Tianlei Zhu, Zhiwei Liu, Yuyan Wang +2
Comprehensive affective analysis is challenging for two reasons: it spans heterogeneous prediction tasks with continuous, ordinal, and multi-label outputs, and affective meaning is…
ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
Jie Gong, Maowei Jiang, Zhiwei Liu +14
Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily asses…
AuditFraudBench: Benchmarking Audit Judgment in Detecting Fraudulent Misstatements
Zhiwei Liu, Yueru He, Qing Ou +4
Large language models (LLMs) have shown strong performance in financial analysis and surface-level factual error detection, yet their ability to identify fraudulent financial misin…
Herculean: An Agentic Benchmark for Financial Intelligence
Xueqing Peng, Zhuohan Xie, Yupeng Cao +60
As AI agents improve, the central question is no longer whether they can solve isolated well-defined financial tasks, but whether they can reliably carry out financial professional…
AutoRedTrader: Autonomous Red Teaming of Trading Agents through Synthetic Misinformation Injection
Zhiwei Liu, Yangyang Yu, Yupeng Cao +8
LLM-based financial agents increasingly rely on both numerical market data and textual signals for sequential trading and stock prediction. However, financial misinformation often…
MFMDQwen: Multilingual Financial Misinformation Detection Based on Large Language Model
Zhiwei Liu, Yuyan Wang, Yuechen Jiang +8
Financial misinformation poses significant threats to financial market stability and individuals' investment decisions. The multilingual environment and the inherent complexity of…