7 papers
Towards Joint Quantization and Token Pruning of Vision-Language Models
Xinqing Li, Xin He, Xindong Zhang +3
Deploying Vision-Language Models (VLMs) under aggressive low-bit inference remains challenging because inference cost is dominated by the long visual-token prefix during prefill an…
Certainty Is Redundant: Token Sparsification for Efficient Camouflaged Object Detection with Vision Foundation Models
Yuhan Gao, Shuhao Kang, Xin He +4
Camouflaged object detection (COD) aims to segment objects that closely resemble their surrounding environments. Vision foundation models (VFMs) provide strong transferable represe…
The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results
Xingyu Qiu, Yuqian Fu, Jiawei Geng +70
Cross-domain few-shot object detection (CD-FSOD) remains a challenging problem for existing object detectors and few-shot learning approaches, particularly when generalizing across…
Amped: Adaptive Multi-stage Non-edge Pruning for Edge Detection
Yuhan Gao, Xinqing Li, Xin He +4
Edge detection is a fundamental image analysis task that underpins numerous high-level vision applications. Recent advances in Transformer architectures have significantly improved…
Make It Up: Fake Images, Real Gains in Generalized Few-shot Semantic Segmentation
Guohuan Xie, Xin He, Dingying Fan +3
Generalized few-shot semantic segmentation (GFSS) is fundamentally limited by the coverage of novel-class appearances under scarce annotations. While diffusion models can synthesiz…
A Comprehensive Survey on Video Scene Parsing:Advances, Challenges, and Prospects
Guohuan Xie, Syed Ariff Syed Hesham, Wenya Guo +4
Video Scene Parsing (VSP) studies dense video understanding, where every pixel in each frame must be segmented, each region must be named, and each object identity must remain cohe…