4 papers
When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
Congyang Ou, Ruike Song, Yang Zhou +3
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…
TRIO: Token Reduction via Inference-Objective Guidance for Efficient Vision-Language Models
Haokui Zhang, Congyang Ou, Dawei Yan +5
Recently, reducing redundant visual tokens in vision-language models (VLMs) to accelerate VLM inference has emerged as a hot topic. However, most existing methods rely on heuristic…
CC-Pan: Channel-wise Compression based Diffusion for Efficient Pan-Sharpening
Junjie Li, Congyang Ou, Haokui Zhang +3
Recently, diffusion models have brought novel insights to pan-sharpening and notably boosted fusion precision. However, most existing models perform diffusion in the pixel space an…
Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
Yang Zhou, Junjie Li, CongYang Ou +3
Due to its extensive applications, aerial image object detection has long been a hot topic in computer vision. In recent years, advancements in Unmanned Aerial Vehicles (UAV) techn…