2 papers
cs.CV2023
SVT: Supertoken Video Transformer for Efficient Video Understanding
Chenbin Pan, Rui Hou, Hanchao Yu +3
Whether by processing videos with fixed resolution from start to end or incorporating pooling and down-scaling strategies, existing video transformers process the whole video conte…
cs.CV2023
EgoViT: Pyramid Video Transformer for Egocentric Action Recognition
Chenbin Pan, Zhiqi Zhang, Senem Velipasalar +1
Capturing interaction of hands with objects is important to autonomously detect human actions from egocentric videos. In this work, we present a pyramid video transformer with a dy…