1 paper
Seonghoon Yu, Ilchae Jung, Byeongju Han +4
Referring image segmentation (RIS) requires dense vision-language interactions between visual pixels and textual words to segment objects based on a given description. However, com…