activity
20242026
most citedLike Humans to Few-Shot Learning through Knowledge Permeation of Vision and Text

2 citations · 3 across the 10 of their papers we have counts for

collaborators
Showing cs.CVShow all

9 papers · 1 filter

cs.CV2026

STAR-IOD: Scale-decoupled Topology Alignment with Pseudo-label Refinement for Remote Sensing Incremental Object Detection

Yaoteng Zhang, Qing Zhou, Junyu Gao +1

Remote sensing imagery typically arrives in the form of continuous data streams. Traditional detectors often forget previously learned categories when learning new ones; therefore,…

cs.CV2026

Efficient Reasoning via Thought Compression for Language Segmentation

Qing Zhou, Shiyu Zhang, Yuyu Jia +4

Chain-of-thought (CoT) reasoning has significantly improved the performance of large multimodal models in language-guided segmentation, yet its prohibitive computational cost, stem…

cs.CV2026

Discriminative Perception via Anchored Description for Reasoning Segmentation

Tao Yang, Qing Zhou, Yanliang Li +1

Reasoning segmentation increasingly employs reinforcement learning to generate explanatory reasoning chains that guide Multimodal Large Language Models. While these geometric rewar…

cs.CV2026

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis

Bingxuan Zhao, Qing Zhou, Chuang Yang +2

Text-to-image synthesis for remote sensing (RS) lacks an accessible, high-performance generative foundation, as directly training diffusion models at large, high resolutions is com…

cs.CV2025★ 1 cited

Scale Efficient Training for Large Datasets

Qing Zhou, Junyu Gao, Qi Wang

The rapid growth of dataset scales has been a key driver in advancing deep learning research. However, as dataset scale increases, the training process becomes increasingly ineffic…

cs.CV2025

A Benchmark for Multi-Lingual Vision-Language Learning in Remote Sensing Image Captioning

Qing Zhou, Tao Yang, Junyu Gao +3

Remote Sensing Image Captioning (RSIC) is a cross-modal field bridging vision and language, aimed at automatically generating natural language descriptions of features and scenes i…