3 papers
cs.CV2026
When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
Congyang Ou, Ruike Song, Yang Zhou +3
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…
cs.CV2025
Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
Yang Zhou, Junjie Li, CongYang Ou +3
Due to its extensive applications, aerial image object detection has long been a hot topic in computer vision. In recent years, advancements in Unmanned Aerial Vehicles (UAV) techn…
cs.IT2024
Cosine Annealing Optimized Denoising Diffusion Error Correction Codes
Congyang Ou, Xiaojing Chen, Wan Jiang
To address the issue of increased bit error rates during the later stages of linear search in denoising diffusion error correction codes, we propose a novel method that optimizes d…