4 papers
When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
Congyang Ou, Ruike Song, Yang Zhou +3
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
Guoting Wei, Xia Yuan, Yang Zhou +6
Open-Vocabulary Aerial Detection (OVAD) and Remote Sensing Visual Grounding (RSVG) have emerged as two key paradigms for aerial scene understanding. However, each paradigm suffers…
UVLM: Benchmarking Video Language Model for Underwater World Understanding
Xizhe Xue, Yang Zhou, Dawei Yan +5
Recently, the remarkable success of large language models (LLMs) has achieved a profound impact on the field of artificial intelligence. Numerous advanced works based on LLMs have…
Open-Vocabulary Object Detection in UAV Imagery: A Review and Future Perspectives
Yang Zhou, Junjie Li, CongYang Ou +3
Due to its extensive applications, aerial image object detection has long been a hot topic in computer vision. In recent years, advancements in Unmanned Aerial Vehicles (UAV) techn…