78 citations · 111 across the 6 of their papers we have counts for
13 papers · 1 filter
P4Q: Learning to Prompt for Quantization in Visual-language Models
Huixin Sun, Runqi Wang, Yanjing Li +4
Large-scale pre-trained Vision-Language Models (VLMs) have gained prominence in various visual and multimodal tasks, yet the deployment of VLMs on downstream application platforms…
Controllable Mind Visual Diffusion Model
Bohan Zeng, Shanglin Li, Xuhui Liu +6
Brain signal visualization has emerged as an active research area, serving as a critical interface between the human visual system and computer vision models. Although diffusion mo…
MVP-SEG: Multi-View Prompt Learning for Open-Vocabulary Semantic Segmentation
Jie Guo, Qimeng Wang, Yan Gao +4
CLIP (Contrastive Language-Image Pretraining) is well-developed for open-vocabulary zero-shot image-level recognition, while its applications in pixel-level tasks are less investig…
PiClick: Picking the desired mask from multiple candidates in click-based interactive segmentation
Cilin Yan, Haochen Wang, Jie Liu +5
Click-based interactive segmentation aims to generate target masks via human clicking, which facilitates efficient pixel-level annotation and image editing. In such a task, target…
Towards Open-Vocabulary Video Instance Segmentation
Haochen Wang, Cilin Yan, Shuai Wang +5
Video Instance Segmentation (VIS) aims at segmenting and categorizing objects in videos from a closed set of training categories, lacking the generalization ability to handle novel…
SwiftNet: Real-time Video Object Segmentation
Haochen Wang, Xiaolong Jiang, Haibing Ren +2
In this work we present SwiftNet for real-time semisupervised video object segmentation (one-shot VOS), which reports 77.8% J &F and 70 FPS on DAVIS 2017 validation dataset, leadin…