1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Yuefei Chen, Jiang Liu, Xiaodong Lin +1
Vision Language Models (VLMs) have recently shown significant advancements in video understanding, especially in feature alignment, event reasoning, and instruction-following tasks…