activity
20242026
collaborators

12 papers

cs.CV2026

When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs

Congyang Ou, Ruike Song, Yang Zhou +3

Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods rely on simil…

cs.CV2026

Rethinking Practical and Efficient Quantization Calibration for Vision-Language Models

Zhenhao Shang, Haizhao Jing, Guoting Wei +4

Post-training quantization (PTQ) is a primary approach for deploying large language models without fine-tuning, and the quantized performance is often strongly affected by the cali…

cs.CV2026

Open-Text Aerial Detection: A Unified Framework For Aerial Visual Grounding And Detection

Guoting Wei, Xia Yuan, Yang Zhou +6

Open-Vocabulary Aerial Detection (OVAD) and Remote Sensing Visual Grounding (RSVG) have emerged as two key paradigms for aerial scene understanding. However, each paradigm suffers…

cs.CV2026

Unlocking Prototype Potential: An Efficient Tuning Framework for Few-Shot Class-Incremental Learning

Shengqin Jiang, Xiaoran Feng, Yuankai Qi +6

Few-shot class-incremental learning (FSCIL) seeks to continuously learn new classes from very limited samples while preserving previously acquired knowledge. Traditional methods of…

cs.CV2025

UVLM: Benchmarking Video Language Model for Underwater World Understanding

Xizhe Xue, Yang Zhou, Dawei Yan +5

Recently, the remarkable success of large language models (LLMs) has achieved a profound impact on the field of artificial intelligence. Numerous advanced works based on LLMs have…

cs.CV2025

Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation

Feng Lin, Marco Chen, Haokui Zhang +3

This paper investigates the role of attention heads in CLIP's image encoder. Building on interpretability studies, we conduct an exhaustive analysis and find that certain heads, di…