12 papers
When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
Congyang Ou, Ruike Song, Yang Zhou +3
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…
Rethinking Practical and Efficient Quantization Calibration for Vision-Language Models
Zhenhao Shang, Haizhao Jing, Guoting Wei +4
Post-training quantization (PTQ) is a primary approach for deploying large language models without fine-tuning, and the quantized performance is often strongly affected by the cali…
Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection
Guoting Wei, Xia Yuan, Yang Zhou +6
Open-Vocabulary Aerial Detection (OVAD) and Remote Sensing Visual Grounding (RSVG) have emerged as two key paradigms for aerial scene understanding. However, each paradigm suffers…
Unlocking Prototype Potential: An Efficient Tuning Framework for Few-Shot Class-Incremental Learning
Shengqin Jiang, Xiaoran Feng, Yuankai Qi +6
Few-shot class-incremental learning (FSCIL) seeks to continuously learn new classes from very limited samples while preserving previously acquired knowledge. Traditional methods of…
UVLM: Benchmarking Video Language Model for Underwater World Understanding
Xizhe Xue, Yang Zhou, Dawei Yan +5
Recently, the remarkable success of large language models (LLMs) has achieved a profound impact on the field of artificial intelligence. Numerous advanced works based on LLMs have…
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
Feng Lin, Marco Chen, Haokui Zhang +3
This paper investigates the role of attention heads in CLIP's image encoder. Building on interpretability studies, we conduct an exhaustive analysis and find that certain heads, di…