1 paper
Zesen Cheng, Kehan Li, Hao Li +5
Temporally locating objects with arbitrary class texts is the primary pursuit of open-vocabulary Video Instance Segmentation (VIS). Because of the insufficient vocabulary of video…