4 citations · 8 across the 2 of their papers we have counts for
1 paper · 1 filter
Enxin Song, Wenhao Chai, Guanhong Wang +10
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet…