1 paper
Zhixuan Wu, Quanxing Zha, Teng Wang +6
Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively…