6 papers
Pest-Thinker: Learning to Think and Reason like Entomologists via Reinforcement Learning
Xueheng Li, Yu Wang, Tao Hu +6
Pest-induced crop losses pose a major threat to global food security and sustainable agricultural development. While recent advances in Multimodal Large Language Models (MLLMs) hav…
PestVL-Net: Enabling Multimodal Pest Learning via Fine-grained Vision-Language Interaction
Xueheng Li, Tao Hu, Ke Cao +5
Effective pest recognition and management are crucial for sustainable agricultural development. However, collecting pest data in real scenarios is often challenging. Compared to ot…
Shuffle Mamba: State Space Models with Random Shuffle for Multi-Modal Image Fusion
Ke Cao, Xuanhua He, Tao Hu +3
Multi-modal image fusion integrates complementary information from different modalities to produce enhanced and informative images. Although State-Space Models, such as Mamba, are…
RelaCtrl: Relevance-Guided Efficient Control for Diffusion Transformers
Ke Cao, Jing Wang, Ao Ma +11
The Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled di…
Rethinking Pan-sharpening: A New Training Process for Full-Resolution Generalization
Ran Zhang, Xuanhua He, Li Xueheng +6
The field of pan-sharpening has recently seen a trend towards increasingly large and complex models, often trained on single, specific satellite datasets. This one-dataset, one-mod…
Distilling Textual Priors from LLM to Efficient Image Fusion
Ran Zhang, Xuanhua He, Ke Cao +4
Multi-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but strugg…