activity
20202026
collaborators
Showing cs.CVShow all

12 papers · 1 filter

cs.CV2026

Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation

Kejia Liu, Haoyang Zhou, Ruoyu Xu +3

Recent advances in cross-view geo-localization (CVGL) methods have shown strong potential for supporting unmanned aerial vehicle (UAV) navigation in GNSS-denied environments. Howev…

cs.CV2026

Rethinking Token Reduction for Large Vision-Language Models

Yi Wang, Haofei Zhang, Qihan Huang +7

Large Vision-Language Models (LVLMs) excel in visual understanding and reasoning, but the excessive visual tokens lead to high inference costs. Although recent token reduction meth…

cs.CV2026

Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation

Chenggong Hu, Yi Wang, Mengqi Xue +3

Textile pattern generation (TPG) aims to synthesize fine-grained textile pattern images based on given clothing images. Although previous studies have not explicitly investigated T…

cs.CV2026

-RSMDE: 40 Faster and High-Fidelity Remote Sensing Monocular Depth Estimation

Ruizhi Wang, Weihan Li, Zunlei Feng +5

Real-time, high-fidelity monocular depth estimation from remote sensing imagery is crucial for numerous applications, yet existing methods face a stark trade-off between accuracy a…

cs.CV2025

Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning

Qihan Huang, Haofei Zhang, Rong Wei +4

RL (reinforcement learning) methods (e.g., GRPO) for MLLM (Multimodal LLM) perception ability has attracted wide research interest owing to its remarkable generalization ability. N…

cs.CV2025

RS3DBench: A Comprehensive Benchmark for 3D Spatial Perception in Remote Sensing

Jiayu Wang, Ruizhi Wang, Jie Song +4

In this paper, we introduce a novel benchmark designed to propel the advancement of general-purpose, large-scale 3D vision models for remote sensing imagery. While several datasets…