4 citations · 4 across the 1 of their papers we have counts for
1 paper
Mingze Xu, Mingfei Gao, Zhe Gan +5
We propose SlowFast-LLaVA (or SF-LLaVA for short), a training-free video large language model (LLM) that can jointly capture detailed spatial semantics and long-range temporal cont…