1 citations · 1 across the 1 of their papers we have counts for
1 paper
Enxin Song, Wenhao Chai, Tian Ye +3
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet…