3 papers
cs.CV2026
FinMTM: A Multi-Turn Multimodal Benchmark for Financial Reasoning and Agent Evaluation
Chenxi Zhang, Ziliang Gan, Liyun Zhu +3
The financial domain poses substantial challenges for vision-language models (VLMs) due to specialized chart formats and knowledge-intensive reasoning requirements. However, existi…
cs.MM2025
Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision
Che Liu, Yingji Zhang, Dong Zhang +13
This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limit…
cs.CV2024
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning
Ziliang Gan, Yu Lu, Dong Zhang +9
In recent years, multimodal benchmarks for general domains have guided the rapid development of multimodal models on general tasks. However, the financial field has its peculiariti…