works on

From the 1 of 7 linked papers with an AI index.

collaborators

7 papers

cs.CV2026

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

Rong Fu, Chunlei Meng, Yangchen Zeng +9

Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming…

cs.CV2026

ObliCity: A Benchmark and Baseline for Roof-to-Ground Projection Displacement Correction

Kai Li, Yupeng Deng, Ligao Deng +6

The paper presents ObliCity, a large-scale benchmark for extracting roof-to-footprint offset vectors in oblique urban remote sensing images, and introduces DragRoof, an ODE-based m…

cs.CV2026

Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation

Chenhao Wang, Yingrui Ji, Yu Meng +1

Open-vocabulary segmentation models often struggle to generalize to unseen combinations of object categories and attributes, because fine-grained descriptions are typically encoded…

cs.CV2025

RS-OOD: A Vision-Language Augmented Framework for Out-of-Distribution Detection in Remote Sensing

Chenhao Wang, Yingrui Ji, Yu Meng +2

Out-of-distribution (OOD) detection represents a critical challenge in remote sensing applications, where reliable identification of novel or anomalous patterns is essential for au…

cs.CV2025

SOPSeg: Prompt-based Small Object Instance Segmentation in Remote Sensing Imagery

Chenhao Wang, Yingrui Ji, Yu Meng +2

Extracting small objects from remote sensing imagery plays a vital role in various applications, including urban planning, environmental monitoring, and disaster management. While…

cs.CV2025

CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization

Yingrui Ji, Xi Xiao, Gaofei Chen +5

Contrastive Language-Image Pretraining (CLIP) has achieved remarkable success in cross-modal tasks such as zero-shot image classification and text-image retrieval by effectively al…