7 papers
FIRE: A Comprehensive Benchmark for Financial Intelligence and Reasoning Evaluation
Xiyuan Zhang, Huihang Wu, Jiayu Guo +8
We introduce FIRE, a comprehensive benchmark designed to evaluate both the theoretical financial knowledge of LLMs and their ability to handle practical business scenarios. For the…
MMFCTUB: Multi-Modal Financial Credit Table Understanding Benchmark
Cui Yakun, Yanting Zhang, Zhu Lei +5
The advent of multi-modal language models (MLLMs) has spurred research into their application across various table understanding tasks. However, their performance in credit table u…
Deliberation on Priors: Trustworthy Reasoning of Large Language Models on Knowledge Graphs
Jie Ma, Ning Qu, Zhitao Gao +8
Knowledge graph-based retrieval-augmented generation seeks to mitigate hallucinations in Large Language Models (LLMs) caused by insufficient or outdated knowledge. However, existin…
Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks
Jiannan Guan, Qiguang Chen, Libo Qin +5
Large Language Models (LLMs) excel in reasoning tasks requiring a single correct answer, but they perform poorly in multi-solution tasks that require generating comprehensive and d…
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
Chenglin Zhu, Tao Zhang, Chong Li +3
Multimodal large language models (MLLMs) still perform poorly on scientific tasks, particularly those requiring multi-step and interpretable reasoning. Their limitations include in…
Efficient Medical VIE via Reinforcement Learning
Lijun Liu, Ruiyang Li, Zhaocheng Liu +5
Visual Information Extraction (VIE) converts unstructured document images into structured formats like JSON, critical for medical applications such as report analysis and online co…