collaborators

7 papers

cs.CV2026

Towards Joint Quantization and Token Pruning of Vision-Language Models

Xinqing Li, Xin He, Xindong Zhang +3

Deploying Vision-Language Models (VLMs) under aggressive low-bit inference remains challenging because inference cost is dominated by the long visual-token prefix during prefill an…

cs.CV2026

Certainty Is Redundant: Token Sparsification for Efficient Camouflaged Object Detection with Vision Foundation Models

Yuhan Gao, Shuhao Kang, Xin He +4

Camouflaged object detection (COD) aims to segment objects that closely resemble their surrounding environments. Vision foundation models (VFMs) provide strong transferable represe…

cs.CV2026

The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results

Xingyu Qiu, Yuqian Fu, Jiawei Geng +70

Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across…

cs.CV2026

Amped: Adaptive Multi-stage Non-edge Pruning for Edge Detection

Yuhan Gao, Xinqing Li, Xin He +4

Edge detection is a fundamental image analysis task that underpins numerous high-level vision applications. Recent advances in Transformer architectures have significantly improved…

cs.CV2026

Make It Up: Fake Images, Real Gains in Generalized Few-shot Semantic Segmentation

Guohuan Xie, Xin He, Dingying Fan +3

Generalized few-shot semantic segmentation (GFSS) is fundamentally limited by the coverage of novel-class appearances under scarce annotations. While diffusion models can synthesiz…

cs.CV2025

A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects

Guohuan Xie, Syed Ariff Syed Hesham, Wenya Guo +4

Video Scene Parsing (VSP) studies dense video understanding, where every pixel in each frame must be segmented, each region must be named, and each object identity must remain cohe…