Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
Congyang Ou, Ruike Song, Yang Zhou +3
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…
cs.CV2025
Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
Yang Zhou, Junjie Li, CongYang Ou +3
Due to its extensive applications, aerial image object detection has long been a hot topic in computer vision. In recent years, advancements in Unmanned Aerial Vehicles (UAV) techn…