activity
20242026
collaborators

6 papers

cs.CL2026

FinSafetyBench: Evaluating LLM Safety in Real-World Financial Scenarios

Yutao Hou, Yihan Jiang, Yuhan Xie +5

Large language models (LLMs) are increasingly applied in financial scenarios. However, they may produce harmful outputs, including facilitating illegal activities or unethical beha…

cs.CL2026

Fin-R1: A Large Language Model for Financial Reasoning through Reinforcement Learning

Zhaowei Liu, Xin Guo, Zhi Yang +14

In recent years, general-purpose large language models (LLMs) such as GPT, Gemini, Claude, and DeepSeek have advanced at an unprecedented pace. Despite these achievements, their ap…

cs.CL2025

FinEval-KR: A Financial Domain Evaluation Framework for Large Language Models' Knowledge and Reasoning

Shaoyu Dou, Yutian Shen, Mofan Chen +9

Large Language Models (LLMs) demonstrate significant potential but face challenges in complex financial reasoning tasks requiring both domain knowledge and sophisticated reasoning.…

cs.CE2025

VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding

Zhaowei Liu, Xin Guo, Haotian Xia +11

Multimodal large language models (MLLMs) hold great promise for automating complex financial analysis. To comprehensively evaluate their capabilities, we introduce VisFinEval, the…

cs.CL2025

FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain

Lingfeng Zeng, Fangqi Lou, Zixuan Wang +18

The booming development of AI agents presents unprecedented opportunities for automating complex tasks across various domains. However, their multi-step, multi-tool collaboration c…

cs.CL2024

FinEval: A Chinese Financial Domain Knowledge Evaluation Benchmark for Large Language Models

Xin Guo, Haotian Xia, Zhaowei Liu +14

Large language models have demonstrated outstanding performance in various natural language processing tasks, but their security capabilities in the financial domain have not been…