Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
Jiahe Wu, Bing Cao, Qilong Wang +3
Multimodal Large Language Models (MLLM) are primarily pre-trained on the RGB modality, thereby limiting their performance on other modalities, such as infrared, depth, and event da…
cs.CV2026
Reversible Efficient Diffusion for Image Fusion
Xingxin Xu, Bing Cao, DongDong Li +2
Multi-modal image fusion aims to consolidate complementary information from diverse source images into a unified representation. The fused image is expected to preserve fine detail…
cs.CV2026
Intelligent Power Grid Design Review via Active Perception-Enabled Multimodal Large Language Models
Taoliang Tan, Chengwei Ma, Zhen Tian +3
The intelligent review of power grid engineering design drawings is crucial for power system safety. However, current automated systems struggle with ultra-high-resolution drawings…