1 paper
Tianming Liang, Kun-Yu Lin, Chaolei Tan +3
Referring video object segmentation (RVOS) aims to segment target objects throughout a video based on a text description. This is challenging as it involves deep vision-language un…