2 citations · 4 across the 2 of their papers we have counts for
2 papers
cs.CV2024★ 2 cited
VideoLLM-MoD: Efficient Video-Language Streaming with Mixture-of-Depths Vision Computation
Shiwei Wu, Joya Chen, Kevin Qinghong Lin +7
A well-known dilemma in large vision-language models (e.g., GPT-4, LLaVA) is that while increasing the number of vision tokens generally enhances visual understanding, it also sign…
cs.CV2023★ 2 cited
AU-aware graph convolutional network for Macro- and Micro-expression spotting
Shukang Yin, Shiwei Wu, Tong Xu +3
Automatic Micro-Expression (ME) spotting in long videos is a crucial step in ME analysis but also a challenging task due to the short duration and low intensity of MEs. When solvin…