5 papers
Improving Multimodal Reasoning via Worst Dimension Optimization
Haocheng Lv, Huaping Zhang, Qiuchi Li +2
Multimodal reasoning requires a path that retains integrity over a wide range of constraints, from visual grounding to logic consistency. However, the current Process Reward Models…
RefiningGPT: Specialized language Models for Automated Refinery Unit-level Process Diagram Synthesis
Dongxiao Liu, Yuwen Ding, Xinghai Wei +4
Applying LLMs to complex industrial processes remains challenging due to the semantic gap between natural language design intents and the rigorous physical logic of engineering. In…
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
Zhiyu Xu, Lean Wang, Yuanxin Liu +5
Vision-Language Models (VLMs) have demonstrated remarkable proficiency in general multi-modal understanding; yet they struggle to efficiently acquire continually evolving domain-sp…
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
Botian Jiang, Lei Li, Xiaonan Li +5
The rapid advancement of Multimodal Large Language Models (MLLMs) has been accompanied by the development of various benchmarks to evaluate their capabilities. However, the true na…
ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom
Jingqi Zhou, Sheng Wang, Jingwei Dong +6
Large vision-language models (LVLMs) have witnessed significant progress on visual understanding tasks. However, they often prioritize language knowledge over image information on…