9 papers
All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection
Yuechen Jiang, Zhiwei Liu, Yupeng Cao +10
We introduce RFC Bench, a benchmark for evaluating large language models on financial misinformation under realistic news. RFC Bench operates at the paragraph level and captures th…
When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
Lingfei Qian, Xueqing Peng, Yan Wang +14
Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies t…
FinAgentBench: A Benchmark Dataset for Agentic Retrieval in Financial Question Answering
Chanyeol Choi, Jihoon Kwon, Alejandro Lopez-Lira +8
Accurate information retrieval (IR) is critical in the financial domain, where investors must identify relevant information from large collections of documents. Traditional IR meth…
Your AI, Not Your View: The Bias of LLMs in Investment Analysis
Hoyoung Lee, Junhyuk Seo, Suhwan Park +5
In finance, Large Language Models (LLMs) face frequent knowledge conflicts arising from discrepancies between their pre-trained parametric knowledge and real-time market data. Thes…
Structuring the Unstructured: A Multi-Agent System for Extracting and Querying Financial KPIs and Guidance
Chanyeol Choi, Alejandro Lopez-Lira, Yongjae Lee +8
Extracting structured and quantitative insights from unstructured financial filings is essential in investment research, yet remains time-consuming and resource-intensive. Conventi…
MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application
Xueqing Peng, Lingfei Qian, Yan Wang +44
Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings. Yet most existing eval…