4 papers
Uni-AdaFocus: Spatial-temporal Dynamic Computation for Video Recognition
Yulin Wang, Haoji Zhang, Yang Yue +4
This paper presents a comprehensive exploration of the phenomenon of data redundancy in video understanding, with the aim to improve computational efficiency. Our investigation com…
Bridging the Divide: Reconsidering Softmax and Linear Attention
Dongchen Han, Yifan Pu, Zhuofan Xia +6
Widely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when d…
Advancing Generalization in PINNs through Latent-Space Representations
Honghui Wang, Yifan Pu, Shiji Song +1
Physics-informed neural networks (PINNs) have made significant strides in modeling dynamical systems governed by partial differential equations (PDEs). However, their generalizatio…
End-to-End Streaming Video Temporal Action Segmentation with Reinforce Learning
Jinrong Zhang, Wujun Wen, Shenglan Liu +3
The streaming temporal action segmentation (STAS) task, a supplementary task of temporal action segmentation (TAS), has not received adequate attention in the field of video unders…