56 citations · 58 across the 3 of their papers we have counts for
3 papers
cs.CV2021
Searching for Two-Stream Models in Multivariate Space for Video Recognition
Xinyu Gong, Heng Wang, Zheng Shou +3
Conventional video models rely on a single stream to capture the complex spatial-temporal features. Recent work on two-stream video models, such as SlowFast network and AssembleNet…
cs.CV2021★ 56 cited
Multiscale Vision Transformers
Haoqi Fan, Bo Xiong, Karttikeya Mangalam +4
We present Multiscale Vision Transformers (MViT) for video and image recognition, by connecting the seminal idea of multiscale feature hierarchies with transformer models. Multisca…
cs.CV2020★ 2 cited
FP-NAS: Fast Probabilistic Neural Architecture Search
Zhicheng Yan, Xiaoliang Dai, Peizhao Zhang +3
Differential Neural Architecture Search (NAS) requires all layer choices to be held in memory simultaneously; this limits the size of both search space and final architecture. In c…