4 citations · 4 across the 1 of their papers we have counts for
1 paper
Enxin Song, Wenhao Chai, Guanhong Wang +10
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet…