1 paper
Yiwen Ren, Jianing Liu, Yingxin Wang +4
The MeViS-Audio track asks a system to segment the objects described by a spoken motion expression throughout a video and to return empty masks when the described target is absent.…