5 papers
Improving Multimodal Reasoning via Worst Dimension Optimization
Haocheng Lv, Huaping Zhang, Qiuchi Li +2
Multimodal reasoning requires a path that retains integrity over a wide range of constraints, from visual grounding to logic consistency. However, the current Process Reward Models…
RefiningGPT: Specialized language Models for Automated Refinery Unit-level Process Diagram Synthesis
Dongxiao Liu, Yuwen Ding, Xinghai Wei +4
Applying LLMs to complex industrial processes remains challenging due to the semantic gap between natural language design intents and the rigorous physical logic of engineering. In…
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
Zhiyu Xu, Lean Wang, Yuanxin Liu +5
Vision-Language Models (VLMs) have demonstrated remarkable proficiency in general multi-modal understanding; yet they struggle to efficiently acquire continually evolving domain-sp…
ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom
Jingqi Zhou, Sheng Wang, Jingwei Dong +6
Large vision-language models (LVLMs) have witnessed significant progress on visual understanding tasks. However, they often prioritize language knowledge over image information on…
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
Botian Jiang, Lei Li, Xiaonan Li +5
The rapid advancement of Multimodal Large Language Models (MLLMs) has been accompanied by the development of various benchmarks to evaluate their capabilities. However, the true na…