2 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.CV2024
MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer
Jianjian Cao, Peng Ye, Shengze Li +4
Vision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large…
cs.CV2024★ 2 cited
ClipSAM: CLIP and SAM Collaboration for Zero-Shot Anomaly Segmentation
Shengze Li, Jianjian Cao, Peng Ye +3
Recently, foundational models such as CLIP and SAM have shown promising performance for the task of Zero-Shot Anomaly Segmentation (ZSAS). However, either CLIP-based or SAM-based Z…