4 papers
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
Jiahe Wu, Bing Cao, Qilong Wang +3
Multimodal Large Language Models (MLLM) are primarily pre-trained on the RGB modality, thereby limiting their performance on other modalities, such as infrared, depth, and event da…
Reversible Efficient Diffusion for Image Fusion
Xingxin Xu, Bing Cao, DongDong Li +2
Multi-modal image fusion aims to consolidate complementary information from diverse source images into a unified representation. The fused image is expected to preserve fine detail…
Intelligent Power Grid Design Review via Active Perception-Enabled Multimodal Large Language Models
Taoliang Tan, Chengwei Ma, Zhen Tian +3
The intelligent review of power grid engineering design drawings is crucial for power system safety. However, current automated systems struggle with ultra-high-resolution drawings…
X-Intelligence 3.0: Training and Evaluating Reasoning LLM for Semiconductor Display
Xiaolin Yan, Yangxing Liu, Jiazhang Zheng +53
Large language models (LLMs) have recently achieved significant advances in reasoning and demonstrated their advantages in solving challenging problems. Yet, their effectiveness in…