1 paper
Deshui Miao, Yameng Gu, Chao Yang +3
This report presents an Audio-aware Referring Video Object Segmentation (Ref-VOS) pipeline tailored to the MEVIS\_Audio setting, where the referring expression is provided in spoke…