4 papers
SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images
Jialu Shen, Han Lyu, Suyang Zhong +5
Spectra are a prevalent yet highly information-dense form of scientific imagery, presenting substantial challenges to multimodal large language models (MLLMs) due to their unstruct…
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
Haoyi Tao, Chaozheng Huang, Nan Wang +4
Multimodal Large Language Models demonstrate strong performance on natural image understanding, yet exhibit limited capability in interpreting scientific images, including but not…
Innovator-VL: A Multimodal Large Language Model for Scientific Discovery
Zichen Wen, Boxue Yang, Shuang Chen +30
We present Innovator-VL, a scientific multimodal large language model designed to advance understanding and reasoning across diverse scientific domains while maintaining excellent…
Uni-Parser Technical Report
Xi Fang, Haoyi Tao, Shuwen Yang +8
This technical report introduces Uni-Parser, an industrial-grade document parsing engine tailored for scientific literature and patents, delivering high throughput, robust accuracy…