1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Ye Sun, Hao Zhang, Henghui Ding +3
Achieving fine-grained spatio-temporal understanding in videos remains a major challenge for current Video Large Multimodal Models (Video LMMs). Addressing this challenge requires…