activity
20242026
collaborators

9 papers

cs.CL2026

All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection

Yuechen Jiang, Zhiwei Liu, Yupeng Cao +10

We introduce RFC Bench, a benchmark for evaluating large language models on financial misinformation under realistic news. RFC Bench operates at the paragraph level and captures th…

cs.MA2025

Orchestration Framework for Financial Agents: From Algorithmic Trading to Agentic Trading

Jifeng Li, Arnav Grover, Abraham Alpuerto +2

The financial market is a mission-critical playground for AI agents due to its temporal dynamics and low signal-to-noise ratio. Building an effective algorithmic trading system may…

cs.CL2025

When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents

Lingfei Qian, Xueqing Peng, Yan Wang +14

Although Large Language Model (LLM)-based agents are increasingly used in financial trading, it remains unclear whether they can reason and adapt in live markets, as most studies t…

cs.CL2025

Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim Evidence Reasoning

Shashidhar Reddy Javaji, Yupeng Cao, Haohang Li +3

Large language models (LLMs) are increasingly being used for complex research tasks such as literature review, idea generation, and scientific paper analysis, yet their ability to…

cs.CL2025

MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial Application

Xueqing Peng, Lingfei Qian, Yan Wang +44

Real-world financial analysis involves information across multiple languages and modalities, from reports and news to scanned filings and meeting recordings. Yet most existing eval…

cs.CL2025

Truth Neurons

Haohang Li, Yupeng Cao, Yangyang Yu +2

Despite their remarkable success and deployment across diverse workflows, language models sometimes produce untruthful responses. Our limited understanding of how truthfulness is m…