5 citations · 5 across the 1 of their papers we have counts for
1 paper
Munan Ning, Bin Zhu, Yujia Xie +5
Video-based large language models (Video-LLMs) have been recently introduced, targeting both fundamental improvements in perception and comprehension, and a diverse range of user i…