activity
20242026
collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV2026

One-Shot Crowd Counting With Density Guidance For Scene Adaptation

Jiwei Chen, Qi Wang, Junyu Gao +3

Crowd scenes captured by cameras at different locations vary greatly, and existing crowd models have limited generalization for unseen surveillance scenes. To improve the generaliz…

cs.CV2026

Efficient Reasoning via Thought Compression for Language Segmentation

Qing Zhou, Shiyu Zhang, Yuyu Jia +4

Chain-of-thought (CoT) reasoning has significantly improved the performance of large multimodal models in language-guided segmentation, yet its prohibitive computational cost, stem…

cs.CV2025

Scale Efficient Training for Large Datasets

Qing Zhou, Junyu Gao, Qi Wang

The rapid growth of dataset scales has been a key driver in advancing deep learning research. However, as dataset scale increases, the training process becomes increasingly ineffic…

cs.CV2025

A Benchmark for Multi-Lingual Vision-Language Learning in Remote Sensing Image Captioning

Qing Zhou, Tao Yang, Junyu Gao +3

Remote Sensing Image Captioning (RSIC) is a cross-modal field bridging vision and language, aimed at automatically generating natural language descriptions of features and scenes i…

cs.CV2024

A Training-Free Framework for Video License Plate Tracking and Recognition with Only One-Shot

Haoxuan Ding, Qi Wang, Junyu Gao +1

Traditional license plate detection and recognition models are often trained on closed datasets, limiting their ability to handle the diverse license plate formats across different…

cs.CV2024

Text-only Synthesis for Image Captioning

Qing Zhou, Junlin Huang, Qiang Li +2

From paired image-text training to text-only training for image captioning, the pursuit of relaxing the requirements for high-cost and large-scale annotation of good quality data r…