2 papers
cs.CV2024
mChartQA: A universal benchmark for multimodal Chart Question Answer based on Vision-Language Alignment and Reasoning
Jingxuan Wei, Nan Xu, Guiyong Chang +3
In the fields of computer vision and natural language processing, multimodal chart question-answering, especially involving color, structure, and textless charts, poses significant…
cs.CL2023
A Survey on Image-text Multimodal Models
Ruifeng Guo, Jingxuan Wei, Linzhuang Sun +7
With the significant advancements of Large Language Models (LLMs) in the field of Natural Language Processing (NLP), the development of image-text multimodal models has garnered wi…