1 paper
Zhende Song, Chenchen Wang, Jiamu Sheng +4
Recent large vision-language models (LVLMs) for video understanding are primarily fine-tuned with various videos scraped from online platforms. Existing datasets, such as ActivityN…