From the 1 of 7 linked papers with an AI index.
7 papers
MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling
Rong Fu, Chunlei Meng, Yangchen Zeng +9
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming…
ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction
Kai Li, Yupeng Deng, Ligao Deng +6
The paper presents ObliCity, a large-scale benchmark for extracting roof-to-footprint offset vectors in oblique urban remote sensing images, and introduces DragRoof, an ODE-based m…
Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
Chenhao Wang, Yingrui Ji, Yu Meng +1
Open-vocabulary segmentation models often struggle to generalize to unseen combinations of object categories and attributes, because fine-grained descriptions are typically encoded…
RS-OOD: A Vision-Language Augmented Framework for Out-of-Distribution Detection in Remote Sensing
Chenhao Wang, Yingrui Ji, Yu Meng +2
Out-of-distribution (OOD) detection represents a critical challenge in remote sensing applications, where reliable identification of novel or anomalous patterns is essential for au…
SOPSeg: Prompt-based Small Object Instance Segmentation in Remote Sensing Imagery
Chenhao Wang, Yingrui Ji, Yu Meng +2
Extracting small objects from remote sensing imagery plays a vital role in various applications, including urban planning, environmental monitoring, and disaster management. While…
CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization
Yingrui Ji, Xi Xiao, Gaofei Chen +5
Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success in cross-modal tasks such as zero-shot image classification and text-image retrieval by effectively al…