1 paper
Pouya Sadeghi, Shawn He, Pedro Pablo Guerrero Vela +3
Modern referring image segmentation pipelines couple a vision-language model (VLM) for grounding with a promptable segmenter such as the Segment Anything Model (SAM) for mask gener…