4 citations · 4 across the 3 of their papers we have counts for
2 papers
cs.CV2024★ 4 cited
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Mingze Xu, Mingfei Gao, Zhe Gan +5
We propose SlowFast-LLaVA (or SF-LLaVA for short), a training-free video large language model (LLM) that can jointly capture detailed spatial semantics and long-range temporal cont…
cs.CV2023
SkeleTR: Towrads Skeleton-based Action Recognition in the Wild
Haodong Duan, Mingze Xu, Bing Shuai +4
We present SkeleTR, a new framework for skeleton-based action recognition. In contrast to prior work, which focuses mainly on controlled environments, we target more general scenar…