11 citations · 28 across the 9 of their papers we have counts for
1 paper · 2 filters
Yucheng Suo, Fan Ma, Linchao Zhu +3
Multi-modal Large language models (MLLMs) show remarkable ability in video understanding. Nevertheless, understanding long videos remains challenging as the models can only process…