15 citations · 15 across the 3 of their papers we have counts for
3 papers
cs.CV2023
Memory-and-Anticipation Transformer for Online Action Understanding
Jiahao Wang, Guo Chen, Yifei Huang +2
Most existing forecasting systems are memory-based methods, which attempt to mimic human forecasting ability by employing various memory mechanisms and have progressed in temporal…
cs.CV2023★ 15 cited
VideoLLM: Modeling Video Sequence with Large Language Models
Guo Chen, Yin-Dong Zheng, Jiahao Wang +8
With the exponential growth of video data, there is an urgent need for automated technology to analyze and comprehend video content. However, existing video understanding models ar…
cs.CV2023
RIFormer: Keep Your Vision Backbone Effective While Removing Token Mixer
Jiahao Wang, Songyang Zhang, Yong Liu +6
This paper studies how to keep a vision backbone effective while removing token mixers in its basic building blocks. Token mixers, as self-attention for vision transformers (ViTs),…