works on

From the 1 of 9 linked papers with an AI index.

collaborators

9 papers

cs.AI2026

Dual-Domain Manifold Modeling for Hyperspectral Image Fusion

Chengxin Xie, Qiya Song, Yangbangyan Jiang +2

The paper proposes a dual-domain manifold modeling framework that combines a topology-aware transformer with frequency-decoupled fusion to better preserve spatial structures and sp…

cs.CV2026

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

Qiwei Ma, Chunping Qiu, Xinjun Cheng +5

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language…

cs.CV2026

Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training

Qiwei Ma, Bin Deng, Junjie Zhu +5

Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can b…

cs.CV2026

SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery

Qiwei Ma, Xukun Lu, Wang Liu +3

Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in self-supervised learning and masked image m…

cs.CV2026

Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding

Chang Liu, Henghui Ding, Nikhila Ravi +40

This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challenge, hosted at CVPR 2026, whi…

cs.CV2026

2nd of the 5th PVUW MeViS-Audio Track: ASR-SaSaSa2VA

Zhiyu Wang, Xudong Kang, Shutao Li

Audio-based video object segmentation aims to locate and segment objects in videos conditioned on audio cues, requiring precise understanding of both appearance and motion. Recent…